avlp12/GLM-5.3-Flash-Alis-MLX-8bit - Каталог нейросетей
Генерация текста

avlp12/GLM-5.3-Flash-Alis-MLX-8bit

Добавлено:
avlp12/GLM-5.3-Flash-Alis-MLX-8bit

MLX (mlx-vlm tree) 8-bit build of GLM-5.3-Flash (320B-A18B, glm5next), converted by streaming dequant of the official FP8 release (84c6a6aa`). This replaces the withdrawn earlier build (which quantized the MoE router). This build keeps the router (mlp.gate) and correction bias unquantized/fp32, the mHC arrays and KDA Alog/dtbias in fp32 as stored, and the vision tower in bf16; the MTP layer (45) is dropped (standalone drafter: avlp12/GLM-5.3-Flash-Alis-MTP-Drafter). Conversion receipts (state.json, finalizereceipt.json`) are in-repo. 334.1 GB on disk. 8-bit g64 affine on experts, attention, dense/shared MLPs, embeddings and head; router / mHC / KDA decay params / norms / convs as stored; vision bf16. Per-module map in config.json. This is the highest-fidelity tier of the family and the teacher/reference class used to score the 4-bit builds (held-out paired KL): 4-bit QUASAR avlp12/GLM-5.3-Flash-Alis-MLX-4bit measures −8.5% KL vs its 4-bit RTN baseline against an 8-bit teacher of this layout family (6-bit measures ≈0.0226 mean KL on the same panel). Expect ≈24-25 tok/s decode @ p512 on an M3 Ultra 512GB (8-bit class), vs ≈29-33 tok/s for the 4-bit build; add MLXMAXMBPERBUFFER=2048…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: avlp12
Теги: mlx, glm5_next, apple-silicon, mixture-of-experts, 8-bit, conversational
Лайков: 4  |  Загрузок: 1,270

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.