Original model elyza/ELYZA-japanese-Llama-2-7b-fast-instruct which is based on Meta’s «Llama 2» and has undergone additional pre-training in Japanese, and thier original post-training and speed up tuning. This model is a quantized(miniaturized to 4.11GB) version of the original model(13.69GB). Quantization reduces the amount of memory required and improves execution speed, but unfortunately performance deteriorates. In particular, the original model is tuned for the purpose of strengthening the ability to follow Japanese instructions, not as a benchmark. Although the ability to follow instructions cannot be measured using existing automated benchmarks, we have confirmed that quantized model significantly deteriorates the ability to follow instructions. At least one GPU is currently required due to a limitation of the Accelerate library. So this model cannot be run with the huggingface space free version. You need autoGPTQ library to use this model. dahara1/ELYZA-japanese-Llama-2-7b-instruct-AWQ is newly published. The awq model has improved ability to follow instructions, so please try it. There are another two llama.cpp version quantized model. If you want to run it in a CPU-only…
Модальности:
Генерация текста
Области применения:
Следование инструкциям
Задача: Генерация текста
Автор: dahara1
Теги: llama, ja, en, text-generation-inference
Лайков: 3 | Загрузок: 378
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.