GEC model for Estonian based on tartuNLP/Llammas-base and fine-tuned on 1) correcting 1M synthetic errors produced by our Llama-based error generation model 2) human GEC data. For training and inference code used in our paper see our repository https://github.com/TartuNLP/gec-llm. Simple example (we provide the templating in tokenizer.chattemplate`) For Estonian, we used a detokenization script (detokenize.py) that also did whitespace and quote normalization, so you might also want to apply those regex rules.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: tartuNLP
Теги: llama, GEC, conversational, et, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 34
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.