> [!WARNING] > This model is a beta preview and a research artifact. It is not an open release. This checkpoint is an intermediate result, currently in a closed beta with already selected partners to work on an improved version. It is under development and subject to change based on partner feedback. Outputs, weights, and behavior may differ between checkpoints. The final model will be released openly under a permissive license, without gated access. We will share access details as soon as it is ready. Lineage: This checkpoint is based on Soofi-Project/Soofi-S-Base. Please refer to the base model card for architectural and training details. > Sizes scale with the total 30B parameters (not the 3.5B active). Q4KM is > an estimate; the others are measured. > > No Q6K: this architecture’s tensor columns (2688/1856/3712) are not > divisible by 256, so every K-quant tensor falls back to a non-K type. For > Q6K that fallback is q80, making it ~as large as Q80 for no gain; Q5KM > and Q4KM fall back to q51/q41 and still shrink. Directly from this repo (select the quant level via the tag): > For a private repo, Ollama needs to know your HF token; otherwise use the > local Modelfile route.…
Модальности:
Генерация текста
Области применения:
Логика и рассуждение Диалог / чат
Задача: Генерация текста
Автор: Soofi-Project
Теги: gguf, mamba-2, moe, reasoning, thinking, soofi, llama.cpp, ollama
Лайков: 4 | Загрузок: 8
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.