Each branch contains an individual bits per weight, with the main one containing only the meaurement.json for further conversions. VRAM requirements listed for both 4k context and 16k context since without GQA the differences are massive (6.2 GB) To download the main (only useful if you only care about measurement.json) branch to a folder called FuseLLM-7B-exl2: To download from a different branch, add the —revision parameter:
Модальности:
Генерация текста
Задача: Генерация текста
Автор: bartowski
Теги: llama, open-llama, mpt, model-fusion, en, endpoints_compatible
Лайков: 3 | Загрузок: 5
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.