Ornith1.5-Ciru-Halo-Agent-vllm-strix-halo
A local AI team, built around AMD Strix Halo. Ornith1.5 Ciru Halo Agent combines a custom quantization of...
A local AI team, built around AMD Strix Halo. Ornith1.5 Ciru Halo Agent combines a custom quantization of...
This repository contains the quantized AMD Q4NX weights, AIE-ML firmware kernels, and runtime configuration for running openbmb/MiniCPM5-2B natively...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
Qwen3.8-27B — the dense hybrid-attention Qwen release (48 gated-delta-net layers + 16 full-attention layers, fullattentioninterval = 4, native...
Text-only ROCmFPX/ROCmFP4 GGUF builds of nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451, with the checkpoint’s embedded one-layer MTP draft model preserved. These are experimental...
A 4-bit ROCmFP4STRIX` quant of poolside/Laguna-S-2.1 (118B-A8B), built and measured on an AMD Ryzen AI Max+ 395 (Strix...
> ## ⚠️ Эти файлы НЕ загружаются в стандартный llama.cpp > Они используют собственные тензорные типы AMD RCMFPX...
MiniMax-M2.7-AWQ-G32-STRIX-2H — это AWQ-квантование смешанной точности amd/MiniMax-M2.7-BF16, созданное для двухузлового вывода AMD Strix Halo (gfx1151) с тензорным параллелизмом...
> [!TIP] > Поддержите эту работу → · X · GitHub · Документ REAP · Cerebras REAP GGUF-квантование...