Mixtral-8x7B-Instruct-v0.1-AutoFP8
— Model Architecture: Mixtral-8x7B-Instruct-v0.1 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
— Model Architecture: Mixtral-8x7B-Instruct-v0.1 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
A local AI team, built around AMD Strix Halo. Ornith1.5 Ciru Halo Agent combines a custom quantization of...
> Use this image: September 9 reference stack > verdictai/trellismx:glm53-flash-p8-r27-reference-20260909 > This is the selected serving image for...
NVFP4 (W4A16) quantization of an abliterated (refusal-direction removed) build of ornith-ai/Ornith-1.5-35B-A3B. The per-layer quantization recipe is matched exactly...
A second quantization of inclusionAI/Ling-3.0-flash, alongside Ling-3.0-flash-HybridQuant-NVFP4-W4A16. Identical bit placement and identical serving contract; it differs in how...
A 13 GB download: 12,982,426,694 bytes = 12.98 GB = 12.09 GiB. Weights and codebooks are 12,958,885,485 B...
4-bit GPTQ (W4A16) quantization of deepseek-ai/DeepSeek-V4-Flash-0731, produced with GPTQModel. This is the V1 release (first calibration run). A...
A.X K2 — это крупномасштабная языковая модель Mixure-of-Experts (MoE), обученная с нуля как высокопроизводительная агентная базовая модель и...
> Неудавшиеся ворота продвижения — исследуйте только артефакт. Эта контрольная точка > опубликована для воспроизводимости и диагностики. Он...
A reproducible deployment and benchmark package for the official deepseek-ai/DeepSeek-V4-Flash-0731 checkpoint on two NVIDIA GB10-class nodes using vLLM...