A Hindi instruction-tuned fine-tune of Gemma 4 E4B, quantized to GGUF for local / CPU / edge use via llama.cpp, LM Studio, and llama-cpp-python. > The smallest quant is ~5.3 GB and runs on an 8 GB laptop — CPU or GPU, fully offline. No API, no cloud. > Part of my 🇮🇳 Hindi LLM Series — small, openly-documented Indic models that actually follow instructions in Hindi. ▶️ Try it live (no install, runs on free CPU): pankajpandey-dev/gemma-4-e4b-hindi-demo This is the GGUF build. The 16-bit model and LoRA adapter are in separate repos. 🔍 Where the base still wins — being honest: vanilla gemma-4-E4B is more detailed and broader in general knowledge. This is not a «smarter» model — it’s a focused, Hindi-native, edge-friendly one. If you want maximum detail and don’t mind Hindi-English mixing, the base may actually suit you better. 📄 The full 25-prompt comparison is written up in the announcement post. > 📐 Size note: Gemma 4 E4B runs at ~4B speed but carries ~8B total parameters (Matformer / per-layer embeddings), so the files are sized like an 8B model. Q4KM still fits comfortably on a laptop with 8 GB+ RAM, CPU or GPU. > 💡 Apple Silicon / integrated GPUs with unified memory count too —…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: pankajpandey-dev
Теги: gguf, gemma4, hindi, indic, india, gemma, llama.cpp, lm-studio
Лайков: 4 | Загрузок: 854
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.