brandonmusic/GLM-5.3-Flash-TrellisMX-MXFP8 - Каталог нейросетей
Генерация текста

brandonmusic/GLM-5.3-Flash-TrellisMX-MXFP8

Добавлено:
brandonmusic/GLM-5.3-Flash-TrellisMX-MXFP8

> Use this image: September 9 reference stack > verdictai/trellismx:glm53-flash-p8-r27-reference-20260909 > This is the selected serving image for this checkpoint and supersedes the earlier serving-image recommendations on this card. Use the current Docker Compose configuration and serving script; Compose pins the verified image digest. Older images and experimental candidates are historical comparison artifacts, not the recommended launch stack. > Docker Hub image tag · Reference stack details A compressed GLM-5.3-Flash checkpoint whose routed experts reconstruct directly into FP8 Tensor Core operands. The September 9 reference stack is the selected runtime: 222.100 tokens/s C1 decode at 8K and 8,457 tokens/s prefill at 32K in separate llmdecodebench runs on four RTX PRO 6000 Blackwell GPUs at 300W each. Use the supplied Docker Compose configuration and serving script. The reference runtime recipe and source manifest record the inference overlays. The checkpoint has not been re-encoded for this update. FP8 math and NVFP4 KV cache describe different parts of the model. Compressed K4/K5 expert weights decode to FP8 operands for multiplication. NVFP4 stores the MLA attention cache.…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: brandonmusic
Теги: trellismx, glm, mixture-of-experts, quantization, vllm, sm120, custom-runtime, shapleymcg
Лайков: 4  |  Загрузок: 0

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.