huihui-ai/Huihui-MoE-1.2B-A0.6B - Каталог нейросетей
Генерация текста

huihui-ai/Huihui-MoE-1.2B-A0.6B

Добавлено:
huihui-ai/Huihui-MoE-1.2B-A0.6B

Huihui-MoE-1.2B-A0.6B is a Mixture of Experts (MoE) language model developed by huihui.ai, built upon the Qwen/Qwen3-0.6B base model. It enhances the standard Transformer architecture by replacing MLP layers with MoE layers, each containing 3 experts, to achieve high performance with efficient inference. The model is designed for natural language processing tasks, including text generation, question answering, and conversational applications. huihui-ai/Huihui-MoE-1B-A0.6B Because tiewordembeddings=True, the parameters for the lm_head were not saved, which causes ollama to be unable to use it. Therefore, this version supports ollama. — Architecture: Qwen3MoeForCausalLM model with 3 experts per layer (numexperts=3), activating 1 expert per token (numexpertspertok=1). — Total Parameters: ~1.2 billion (1.2B) — Activated Parameters: ~0.62 billion (0.6B) during inference, comparable to Qwen3-0.6B — Developer: huihui.ai — Release Date: June 2025 — License: Inherits the license of the Qwen3 base model (apache-2.0) This model was fully fine-tuned with BF16 on first 20k rows of nvidia/OpenCodeReasoning dataset for 1 epoch. This model was fully fine-tuned with BF16 on entire…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: huihui-ai
Теги: qwen3_moe, moe, conversational, endpoints_compatible
Лайков: 4  |  Загрузок: 133

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.