Qwen3.8-27B-DFlash2-Q3_K_M-GGUF
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
This repository contains a DFlash draft model for moonshotai/Kimi-K3 trained only on a generic data mix (no tool...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
GGUF conversion of z-lab/Qwen3.6-35B-A3B-DFlash for llama.cpp. > This is a DFlash draft model, not a standalone language model....
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
Bielik-Minitron-7B-v3.0-DFlash is a DFlash draft model designed for use with Bielik-Minitron-7B-v3.0-Instruct. Its development and training were supported by...
I’ve been trying to get as much performance as possible out of a single DGX Spark with Gemma...
First public extraction of Google Gemma 4’s Multi-Token Prediction (MTP) drafter weights from LiteRT format into standard PyTorch...
Автономный черновой вариант («поддерживаемая») модели спекулятивного декодирования DSpark, упакованный как один GGUF емкостью 5,6 ГиБ для движка ds4....
Отдельный проект MTP (многотокеновое предсказание) предназначен для спекулятивного декодирования с помощью CosmicRaisins/GLM-5.2-AWQ-INT4-15pct. cyankiwi/GLM-5.2-AWQ-INT4 удаляет собственный уровень MTP GLM-5.2,...