NVFP4-quantized build of prefeitura-rio/Rio-3.5-Open-397B — a 397B-parameter (17B active) Qwen3.5-MoE vision-language model (512 experts, hybrid softmax + linear/DeltaNet attention), itself a finetune of Qwen/Qwen3.5-397B-A17B. edit: ir daid ir was rhese thinfs. but it appears to be. nex n2 with a system prompt laugh out loud Quantized with NVIDIA TensorRT Model Optimizer 0.44 using the per-expert streaming calibration pipeline from local-inference-lab/quant-toolkit (Luke Alonso). Produced on 8x B200. Size on disk: ~251 GB (46 shards). Only the routed MoE expert MLPs (gate/up/down) are NVFP4 (4-bit, blockwise FP8 scales, group size 16), calibrated per-expert. Left in BF16: shared-expert MLPs (active every token), attention (softmax + DeltaNet), router/gates, vision tower, MTP, embeddings, lm_head. KV cache is FP8 (e4m3). This mirrors lukealonso/Qwen3.5-397B-A17B-NVFP4. — quantalgo: NVFP4 | quantmethod: modelopt — kvcachescheme: {‘dynamic’: False, ‘numbits’: 8, ‘type’: ‘float’}` Per-expert max-calibration over a finetune-appropriate subset (Rio tracks the Qwen3.5 base, so the full base-model corpus is unnecessary): deep-reasoning + diverse-instruction + agentic-coding corpora…
Модальности:
Генерация текста Мультимодальность
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: brandonmusic
Теги: qwen3_5_moe, qwen3.5, moe, quantized, nvfp4, fp4, multimodal, conversational
Лайков: 4 | Загрузок: 11
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.