aayanmishra-ml/Athena-3-14B - Каталог нейросетей
Генерация текста

aayanmishra-ml/Athena-3-14B

Добавлено:
aayanmishra-ml/Athena-3-14B

Athena-3 🚀 Faster, Sharper, Smarter than Athena 1 and Athena 2🌟 Athena-3-14B is a 14.0-billion-parameter causal language model fine-tuned from Qwen2.5-14B-Instruct. This model is designed to provide highly fluent, contextually aware, and logically sound outputs across a broad range of NLP and reasoning tasks. It balances instruction-following with generative flexibility. — Model Developer: Aayan Mishra — Model Type: Causal Language Model — Architecture: Transformer with Rotary Position Embeddings (RoPE), SwiGLU activation, RMSNorm, Attention QKV bias, and tied word embeddings — Parameters: 14.0 billion total (12.84 billion non-embedding) — Layers: 40 — Attention Heads: 40 for query and 4 for key-value (Grouped Query Attention) — Vocabulary Size: Approximately 151,646 tokens — Context Length: Supports up to 131,072 tokens — Languages Supported The fine-tuning process spanned approximately 90 minutes over 60 epochs, utilizing a curated instruction-tuned dataset. It is tailored for…

Модальности:
Генерация текста

Области применения:
Генерация кода Диалог / чат Биология Химия


Задача: Генерация текста
Автор: aayanmishra-ml
Теги: qwen2, chemistry, biology, code, text-generation-inference, STEM, unsloth, trl
Лайков: 3  |  Загрузок: 38

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.