Zhayr1/BitMamba-2-0.25B - Каталог нейросетей
Генерация текста

Zhayr1/BitMamba-2-0.25B

Добавлено:
Zhayr1/BitMamba-2-0.25B

BitMamba-2-255M is the ultra-efficient baseline model of the BitMamba-2 family. It integrates 1.58-bit ternary quantization (BitNet) into the Mamba-2 architecture. Despite its small size, it demonstrates stable convergence and surprising reasoning capabilities, serving as the proof-of-concept for scaling ternary State Space Models. — Architecture: Mamba-2 SSM + BitNet b1.58 (Ternary Weights). — Parameters: 255M. — Precision: 1.58-bit (weights {-1, 0, 1}). — Training Tokens: Trained on high-quality data (FineWeb-Edu, Cosmopedia, Stack-Dedup). — Hardware: Trained on Google Cloud TPU v6e. This model serves as the baseline for our scaling laws analysis. As shown in the scaling analysis below, the 255M model (blue line) establishes a stable learning trajectory, which is significantly improved upon by the 1B model (red line). This model is optimized for extreme edge deployment (IoT, Mobile, Legacy Hardware) using our custom C++ inference engine. The bitmamba255m.msgpack contains the raw JAX weights for research purposes. You can load them using the source code provided in src/` on GitHub. If you use this model or our architecture, please cite our paper:

Модальности:
Генерация текста


Задача: Генерация текста
Автор: Zhayr1
Теги: jax, bitmamba, bitnet, mamba, ssm, 1.58-bit, ternary, efficient-inference
Лайков: 4  |  Загрузок: 27

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.