younissk/nanoBeard-frigate-360M-GGUF - Каталог нейросетей
Генерация текста

younissk/nanoBeard-frigate-360M-GGUF

Добавлено:
younissk/nanoBeard-frigate-360M-GGUF

GGUF quantizations of a nanoBeard Frigate pirate chat model (358.3M params), for on-device inference with llama.cpp and the NanoBeard mobile app. SFT val loss ≈ 3.036. Frigate’s architecture (RoPE + SwiGLU + RMSNorm + per-head QK-norm, no biases, tied embeddings) is Qwen3-equivalent, so these load with the upstream qwen3 GGUF arch — no custom runtime needed. Q4KM is the default for phones (smallest + fastest). Q80` is a near-lossless fallback when you have the storage and want max quality. Turns are separated by a single newline; stop generation at the token. Example with llama.cpp: Custom 16,384-token byte-level BPE (GPT-2-style pre-tokenizer), embedded in the GGUF. No external tokenizer file required.

Модальности:
Генерация текста


Задача: Генерация текста
Автор: younissk
Теги: gguf, llama.cpp, nanobeard, pirate, on-device, mobile, en, endpoints_compatible
Лайков: 4  |  Загрузок: 64

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.