This is a tiny version of Qwen/Qwen3.8-2.4T-A95B created for testing and development. — Base Model: Qwen/Qwen3.8-2.4T-A95B — Architecture: qwen35moetext (Qwen35MoeForCausalLM) — Total Parameters: 0.97B — Activated Parameters: ~0.6B The following parameters were reduced from the original model: Single safetensors file with packed expert format (experts.gateupproj, experts.downproj`) matching the original checkpoint structure. Layer types follow the [linear, linear, linear, full] x 2 pattern from the original. This model was created using the llm-compressor create-tiny-model claude skill. 1. Inspected the original 2.4T parameter MoE model configuration 2. Reduced all dimensions to create a ~1B parameter model while preserving the hybrid linear/full attention architecture 3. Fine-tuned on a toy dataset to achieve perplexity ~1.0 4. Converted checkpoint format to match the original (packed experts, correct tensor naming) 5. Validated model loading and generation — The model preserves the hybrid attention pattern: 6 linear attention layers and 2 full attention layers in a [linear, linear, linear, full] x 2 pattern — MTP (multi-token prediction) layers are removed (mtpnumhiddenlayers=0)…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: inference-optimization
Теги: qwen3_5_moe_text, conversational, endpoints_compatible
Лайков: 4 | Загрузок: 5,330
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.