Quazim0t0/Escarda-86M-Base - Каталог нейросетей
Генерация текста

Quazim0t0/Escarda-86M-Base

Добавлено:
Quazim0t0/Escarda-86M-Base

Escarda-86M-Base is a ~86M-parameter, from-scratch decoder-only language model — the base sibling of Quazim0t0/Escarda-86M (the chat-tuned model). It shares the same SpikeWhaleLM architecture (Multi-head Latent Attention, an n-gram «engram» memory, hash-lookup layers, hyper-connections, an HRM refinement step, and JEPA / multi-token-prediction training objectives) and the same custom ChatML-aware tokenizer. This checkpoint is a JEPA-distilled base. It is best used as a starting point for continued pretraining / fine-tuning rather than as a chat assistant. > Related models: SFT / chat model → Quazim0t0/Escarda-86M > · live demo → Escarda-86M-Chat Space Trained using Modal’s credits during the Small Models, Big Adventures Hackathon. For the full architecture description see the chat model’s card. > Note: as a distilled base, this checkpoint has the lowest byte-perplexity of the > Escarda family but trades off downstream task accuracy — a good reminder that perplexity > alone is not a reliable capability ranking. For the strongest chat behaviour use > Escarda-86M; use this model when you > want a low-loss base to continue pretraining or fine-tune. Custom architecture — load with…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: Quazim0t0
Теги: spike_whale, feature-extraction, small-models, base-model, mla, jepa, experimental, custom_code
Лайков: 4  |  Загрузок: 367

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.