A high-quality oQ5e quantization of Qwen3.8-Flash-Next, targeting improved reasoning robustness and generation stability over oQ4e while still fitting comfortably on a 128 GB Apple Silicon system with SSD N-gram offload. This model was quantized directly from the official Qwen/Qwen3.8-Flash-Next weights using oMLX 0.6.4. The goal of this quantization is to retain more of the original model’s reasoning quality and generation stability than lower-bit oQ4e variants, while remaining practical for local inference on a 128 GB Apple Silicon Mac. The oQ5e weights were quantized directly from the original BF16 source model. Jundot/Qwen3.8-Flash-Next-oQ4e-mtp was used only as the sensitivity model for determining the mixed-precision allocation. Its quantized weights are not the source weights of this model. The included oqimatrixreport.json contains the quantization/imatrix report generated during the process. > Note: oQ mixed-precision quantization can use different effective settings for selected tensors/modules. The nominal group size used for this quantization was 64. — MacBook Pro — Apple M5 Max — 128 GB Unified Memory — oMLX — SSD N-gram Offload enabled — Lightning MTP enabled During…
Модальности:
Генерация текста Компьютерное зрение
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: GBP-DE
Теги: mlx, qwen4_exp, omlx, qwen, qwen3.8, oq, oq5, oq5e
Лайков: 4 | Загрузок: 14
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.