— Model architecture: full GLM-5.3 (GlmMoeDsaForCausalLM) — Input: text — Output: text — Source checkpoint: zai-org/GLM-5.3-BF16, revision 304b8051cfb2b260b61ce0cbe330e02a98e73639 — Validated hardware: 4× AMD Instinct MI350 GPUs (gfx950) — Validated runtime: stock InferenceX/SGLang ROCm path — SGLang image tag lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260908 — TP4/EP4, EAGLE MTP, TileLang DSA, AITER MXFP4 MoE, FP8 E4M3 KV cache, and HiCache — AMD Quark 0.12.post1+1b229f7 c The promoted checkpoint’s internal candidate name is Strong7. It was quantized from the BF16 checkpoint, not from the published FP8 checkpoint. The 282 model shards contain 438,001,945,864 bytes (407.92 GiB) of indexed model weights. This is 42.04% smaller than the official GLM-5.3 FP8 checkpoint and 70.93% smaller than the BF16 source. AMD Quark applies OCP MXFP4 E2M1 quantization to the routed MoE expert weights. Weights use static 1×32 block scaling with E8M0 scales; expert activations are quantized dynamically with the same 1×32 layout. No calibration dataset is required for the initial MXFP4 conversion. -…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: OneNexus
Теги: glm_moe_dsa, quark, mxfp4, rocm, sglang, conversational, en, zh
Лайков: 4 | Загрузок: 933
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.