A tiny yet powerful instruction-tuned language model optimized for CPU inference. With only 135 million parameters and a file size of 138 MB, this model delivers impressive performance even on modest hardware. — Tiny Footprint: Only 138 MB in size — CPU-Friendly: Runs efficiently without a GPU — Low Resource Requirements: Works on systems with just 1-2 GB RAM — Fast Inference: Responsive even on older CPUs — Instruction-Tuned: Optimized for chat and instruction-following tasks — Long Context: Supports up to 8,192 tokens — Architecture: LLaMA-like transformer — Parameters: 135M — Format: GGUF (compatible with llama.cpp ecosystem) — Quantization: Q80 (8-bit linear quantization) — Typ Use the following message format for best results: — llama.cpp — text-generation-webui — LM Studio — KoboldCPP — llama-cpp-python — ✨ Runs Offline: No internet connection needed — 📱 Tiny Footprint:…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: HackNetAyush
Теги: gguf, llama, q8_0, quantized, llama.cpp, smollm2, embedded-ai, lightweight
Лайков: 4 | Загрузок: 268
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.