SupraLabs/Supra-Mini-v3-0.5M - Каталог нейросетей
Генерация текста

SupraLabs/Supra-Mini-v3-0.5M

Добавлено:
SupraLabs/Supra-Mini-v3-0.5M

Supra Mini v3 0.5M is a very tiny base model trained on 1 billion tokens of Fineweb-Edu for 2 epochs as the third version of our Supra Mini series. — Parameters: 467,648 (0.5M) — Architecture: Llama — Vocab size with custom BPE tokenizer: 4096 — Hidden Size: 64 — Intermediate Size: 128 — Hidden Layers: 5 — Attention Heads: 8 — Max Position Embeddings: 512 — Learning rate: 5e-4 — Weight Decay: 0.01 — Trained in bfloat16 This model reached a final train loss after 2 epochs of 4.872. Prompt: «Artificial intelligence is » Output:: «Artificial intelligence is ignored by the Sunnette-Caiti. Prequent and Burkin, A.S. Cougalia (Mean) Researchers in the Northern Museum of Treaty (Figure 1). The first thing that has a drop or an ambitions to the timeline, which is used for more than a large amount of cosmic moisture (Bargeus), the vital bones are not a few weeks ago, but there was no longer in the same time where he had a hour. It’s also known as the “the person” because they were unhab» Prompt: «The main concept of physics is » Output:: «The main concept of physics is escape, and that the other important to be unable to be used. «It’s a new way for how it can be used in this article on…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: SupraLabs
Теги: llama, small, cpu, supra, v3, tiny, mini, open
Лайков: 4  |  Загрузок: 61

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.