pugant/Qwen3.8-27B-MTP-Q4_0_ROCMFP4_STRIX_LEAN - Каталог нейросетей
Генерация текста

pugant/Qwen3.8-27B-MTP-Q4_0_ROCMFP4_STRIX_LEAN

Добавлено:
pugant/Qwen3.8-27B-MTP-Q4_0_ROCMFP4_STRIX_LEAN

Qwen3.8-27B — the dense hybrid-attention Qwen release (48 gated-delta-net layers + 16 full-attention layers, fullattentioninterval = 4, native 262K context, qwen35 GGUF arch) — quantized to Q40ROCMFP4STRIXLEAN (4.34 BPW effective, 13.82 GiB). MTP layer included (blk.64 with nextn. tensors, nextnpredictlayers = 1): serve with —spec-type draft-mtp to enable speculative decoding. Tuned for AMD Strix Halo (gfx1151) on the ROCmFPX fork family — we serve and benchmark these files on our lab runtime (full source: pugant/strix-nebulosa; upstream: charlie12345/ROCmFPX). > ⚠️ This GGUF is for the ROCmFPX fork of llama.cpp. It will not load in stock llama.cpp > (invalid ggml type). — Fork-specific tensor types (q40rocmfp4, K/V protection, Q5K token embeddings). Requires a ROCmFPX fork build — see Runtime. Upstream charlie12345/ROCmFPX loads these files too (any build with the custom GGML types — HIP or the Vulkan-only build, both ship the deltanet kernels: gateddeltanet.comp, ssmscan.comp, ssmconv.comp). — AMD RDNA 3.5 (gfx1151 / Strix Halo) target. Tested on Radeon 8060S iGPU, not elsewhere. — FP4 is a memory-bandwidth play** on this class of hardware (RDNA 3.5 has no FP4 silicon): smaller…

Модальности:
Генерация текста

Области применения:
Диалог / чат Мультиязычность


Задача: Генерация текста
Автор: pugant
Теги: llama.cpp, gguf, rocmfpx, gfx1151, strix-halo, qwen35, dense, mtp
Лайков: 4  |  Загрузок: 1,003

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.