— Developed by: klei aliaj — Model type: Bleta-Meditor 27B fine-tuned with GRPO for Albanian reasoning tasks — License: apache-2.0 — Finetuned from model: Bleta-Meditor 27B (based on Gemma 3 architecture) — Language: Albanian — Framework: Hugging Face Transformers This model is a fine-tuned version of the Bleta-Meditor 27B model, specifically optimized for the Albanian language using Generative Rejection Policy Optimization (GRPO) to improve its reasoning capabilities. Bleta is an Albanian adaptation based on Google’s Gemma 3 architecture. This Albanian language model was fine-tuned using GRPO (Generative Rejection Policy Optimization), a reinforcement learning technique that trains models to optimize for specific reward functions. The model was trained to: 1. Follow a specific reasoning format with dedicated sections for workings and solutions 2. Produce correct mathematical solutions in Albanian 3. Show clear step-by-step reasoning processes The model has been trained to follow a specific reasoning format: — Working out/reasoning sections are enclosed within and tags — Final solutions are provided between and tags — Framework: Hugging Face’s TRL library — Optimization: LoRA…
Модальности:
Генерация текста
Области применения:
Математика Логика и рассуждение Диалог / чат
Задача: Генерация текста
Автор: klei1
Теги: gguf, gemma3_text, text-generation-inference, albanian, gemma3, reasoning, mathematics, grpo
Лайков: 3 | Загрузок: 64
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.