gemma-4-E4B-it-fp8
FP8 quantized version of google/gemma-4-E4B-it (8B params, edge model). Produced and maintained by vrfai. This model was quantized...
FP8 quantized version of google/gemma-4-E4B-it (8B params, edge model). Produced and maintained by vrfai. This model was quantized...
The following layers are not quantized to preserve model quality: > Note: The DFlash drafter was trained on...
> Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no...
> [!IMPORTANT] > Уведомление о наименовании (2026-04-10). Метод «HLWQ», используемый в этой модели, переименовывается в HLWQ (Hadamard-Lloyd Weight...
— Единый вывод ~106 токенов/с @ 1000 токенов (измеряется в режиме отладки) — Использование памяти: ~2,35 гиБ 2,25bpw...
> [!IMPORTANT] > Уведомление о наименовании (2026-04-10). Метод «HLWQ», используемый в этой модели, переименовывается в HLWQ (Hadamard-Lloyd Weight...
Квантование MLX nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 для Apple Silicon. — Apple Silicon Mac с унифицированной памятью 128 ГБ — mlx-lm >=...
Block-wise FP8 (E4M3) quantization of DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED, optimized for vLLM inference. — Full 262K context on a single RTX...
This is an FP8-quantized version of dnotitia/DNA-2.0-14B, optimized for efficient inference by DLM (Data Science Lab., Ltd.). FP8...
This repository hosts the instruction‑tuned Qwen2.5-7B-Instruct model. Compared to earlier versions, Qwen2.5 offers stronger performance in areas such...