GLM-5.2-504B-W4A16
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
This repository now bundles glmmoedsa.py (declared via modelfile in config.json`), a fixed runtime for this architecture, and needs...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing GLM-4.7-Flash-REAP-23B-A3B, a memory-efficient compressed variant of...
The Model Wwayu/GLM-4.7-PRISM-mlx-2Bit was converted to MLX format from Ex0bit/GLM-4.7-PRISM using mlx-lm version 0.28.3. Модальности:Генерация текста Области применения:Диалог...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing GLM-4.6-REAP-252B-A32B, a memory-efficient compressed variant of...
Модель генерации текста Модальности:Генерация текста Задача: Генерация текста Автор: SicariusSicariiStuff Теги: glm4_moe, glm, MOE, endpoints_compatible, compressed-tensorsЛайков: 4 | ...
Original model: https://huggingface.co/byroneverson/LongWriter-glm4-9b-abliterated Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings...
Исходная модель: https://huggingface.co/THUDM/glm-4-9b-chat-1m Некоторые из этих квантов (Q3KXL, Q4KL и т. д.) представляют собой стандартный метод квантования, в котором...
Отдельный проект MTP (многотокеновое предсказание) предназначен для спекулятивного декодирования с помощью CosmicRaisins/GLM-5.2-AWQ-INT4-15pct. cyankiwi/GLM-5.2-AWQ-INT4 удаляет собственный уровень MTP GLM-5.2,...
Macaron-V1-Preview-744B-Merged — это объединенный вариант Macaron-V1-Preview от MindLab Research с полной контрольной точкой, прошедший пост-обучение из GLM-5.1 с...