OmniGEC-Minimal-8B extends the open-weight AYA-Expanse-8B with instruction tuning and supervised fine-tuning on OmniGEC, a silver-standard GEC corpus that includes MultiGEC-25, Wikipedia, Reddit edits for 11 low-/mid-resource European languages. The result is a single model capable of paragraph-level correction across all covered languages achieving State-Of-The-Art (SOTA) results for paragraph-based editing in minimal and fluency tracks. Silver corrections were created with a three-step prompt → generate 3 candidates → aggregate pipeline using o1-preview and GPT-4o-mini. — Metric: GLEU via the official MultiGEC-25 CodaLab evaluator (minimal & fluency tracks). — Both OmniGEC-tuned models surpass the paragraph-based baseline LLaMA-3-8B by +9–10 GLEU on the minimal track and deliver the current best open scores for Estonian and Latvian. — Reddit and UberText corrections are machine-generated; noise remains, esp. in slang. — Sequences > 1,600 tokens are truncated unless you raise maxnewtokens. — For details on use, please refer to our GitHub — We strongly recommend you to follow the inference code we used in notebooks for both gemma and aya, as there’s additional parameters, like…
Модальности:
Генерация текста
Области применения:
Диалог / чат Мультиязычность
Задача: Генерация текста
Автор: lang-uk
Теги: cohere, GEC, Multilingual, Education, GrammaticalErrorCorrection, text2text-generation, conversational, en
Лайков: 3 | Загрузок: 33
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.