Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v2
— bits: 4 — dynamic: null — groupsize: 32 — descact: true — staticgroups: false — sym: true...
— bits: 4 — dynamic: null — groupsize: 32 — descact: true — staticgroups: false — sym: true...
If you run into errors on a multi GPU machine, I’ve found that setting CUDAVISIBLEDEVICES=0 helps. Llama-3.1-Nemotron-70B-Instruct is...
4-bit GPTQ (W4A16) quantization of deepseek-ai/DeepSeek-V4-Flash-0731, produced with GPTQModel. This is the V1 release (first calibration run). A...
An 8.9M-parameter question-answering model that runs entirely offline on an ESP32-S3 microcontroller, answering espresso questions a piece at...
Canonical one-package weights for 4× DGX Spark. Repo id: drowzeys/keys-latest-GLM-5.2-Quantrio-INT4-INT8-Mixed-Abliterated-DFlash > Gated access. Safety refusals have been removed....
Pre-converted weights for colibrì — a pure-C engine that runs huge MoE models on consumer hardware by keeping...
Массы квантуют до INT4 (размер группы 128); активации выполняют в BF16. Результатом является контрольная точка объемом 388 ГБ...
MobileMoE — это семейство языковых моделей Mixture-of-Experts (MoE) на устройстве с субмиллиардными активными параметрами, предназначенных для продвижения границы...
Qwen 3 Embedding 8B in INT4. About 5 GB on disk. Runs on an 8 GB consumer GPU....
Partial GPTQ Int4 quant of Qwen/Qwen3.6-27B, produced with the verbatim recipe from Qwen’s own Qwen/Qwen3.5-27B-GPTQ-Int4 — only MLP/FFN...