williamliao/Qwen3.6-35B-A3B-DFlash-GGUF - Каталог нейросетей
Генерация текста

williamliao/Qwen3.6-35B-A3B-DFlash-GGUF

Добавлено:
williamliao/Qwen3.6-35B-A3B-DFlash-GGUF

GGUF conversion of z-lab/Qwen3.6-35B-A3B-DFlash for llama.cpp. > This is a DFlash draft model, not a standalone language model. > It must be used together with a compatible Qwen3.6-35B-A3B target model. Base model: z-lab/Qwen3.6-35B-A3B-DFlash Target model: Qwen/Qwen3.6-35B-A3B Format: GGUF Quantization: Q4KM Converted from the original Hugging Face model using the latest converthftogguf.py`. nmax = 2 provides the highest acceptance rate. nmax = 3 provides the best balance between throughput and acceptance rate. nmax = 4/5 may improve peak throughput slightly in some low-entropy tasks such as code completion, JSON, and repetitive patterns, but the overall wall time improves only marginally while acceptance rate drops noticeably. A compatible Qwen3.6-35B-A3B GGUF target model is required for speculative decoding. z-lab — Original DFlash model Qwen Team — Qwen3.6-35B-A3B * ggml-org/llama.cpp — GGUF format and DFlash inference implementation This repository contains a converted GGUF version of the original DFlash draft model. All original licenses, usage restrictions, and intellectual property remain with the upstream authors. Please refer to the original repositories for…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: williamliao
Теги: llama.cpp, gguf, dflash, speculative-decoding, speculator, draft-model, qwen3.6, qwen
Лайков: 4  |  Загрузок: 925

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.