StarpowerTechnology/arXiv-WVY-43M - Каталог нейросетей
Генерация текста

StarpowerTechnology/arXiv-WVY-43M

Добавлено:
StarpowerTechnology/arXiv-WVY-43M

arXiv-WVY-43M is a 43.5M parameter tiny language model developed by Starpower Technology intended for autonomous research. This is the prototype experiment. We are still creating finetuning dataset Kaggle [https://www.kaggle.com/code/starpowertechnology/arxiv-wvy-43m-demo] The model uses a compact DeepSeek-V3-style architecture and was trained from scratch on arXiv titles and abstracts. This is a base language model, not an instruction-tuned or chat-tuned model. WVY-43M uses a compact DeepSeek-V3-style causal language model architecture containing: — Multi-head latent attention — Mixture-of-Experts layers — 8 routed experts with Top-2 routing — 1 shared expert — Q/KV low-rank projections — Rotary positional embeddings — YaRN RoPE scaling — Multi-Token Prediction layer The architecture is intentionally kept small for research into compact language models, training from scratch, experimentation, and low-compute deployment. The training corpus gives the model significant exposure to scientific and technical language, particularly terminology appearing in academic research. Enable Internet access for the Kaggle notebook so the model can be downloaded from Hugging Face. Because…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: StarpowerTechnology
Теги: deepseek_v3, causal-lm, tiny-language-model, en, endpoints_compatible
Лайков: 4  |  Загрузок: 838

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.