Recently, due to the success of ChatGPT, numerous large language models have emerged in an attempt to catch up with ChatGPT’s capabilities. However, when it comes to Korean language performance, it has been observed that many models still struggle to provide accurate answers or generate Korean text effectively. This study addresses these challenges by introducing a multi-task instruction technique that leverages supervised datasets from various tasks to create training data for Large Language Models (LLMs). Model Developers : davidkim(changyeon kim) Repository : https://github.com/davidkim205/komt Lora target modules : qproj, oproj, vproj, gateproj, downproj, kproj, upproj Model Size : 120MB Model Architecture : komt-llama-2-13b-v1-lora is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning by multi-task instruction License: This model is under a Non-commercial** Bespoke License and governed by the Meta license. For objective model evaluation, we initially used EleutherAI’s lm-evaluation-harness but obtained unsatisfactory results. Consequently, we conducted evaluations using ChatGPT, a widely used…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: davidkim205
Теги: peft, facebook, meta, llama, llama-2, llama-2-chat, en, ko
Лайков: 3 | Загрузок: 5
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.