bartowski/rednote-hilab_dots.llm1.inst-GGUF - Каталог нейросетей
Генерация текста

bartowski/rednote-hilab_dots.llm1.inst-GGUF

Добавлено:
bartowski/rednote-hilab_dots.llm1.inst-GGUF

Используя llama.cpp с высвобождением b5669 с моим PR отсюда для квантования. Оригинальная модель: https://huggingface.co/rednote-hilab/dots.llm1.inst Запустите их напрямую с помощью llama.cpp или любого другого проекта на основе llama.cpp Некоторые из этих квантов (Q3KXL, Q4KL и т. д.) являются стандартным методом квантования с вложениями и выходными весами, квантованными до Q8_0 вместо того, что они обычно используют по умолчанию. Если размер модели превышает 50 ГБ, она будет разделена на несколько файлов. In order to download them all to a local folder, run: You can either specify a new local-dir (rednote-hilabdots.llm1.inst-Q80) or download them all in place (./) Previously, you would download Q4044/48/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass. Now, however, there is something called «online repacking» for weights. details in this PR. If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly. As of llama.cpp build b4282 you will not be able to run the Q40XX files and will instead need to use Q40. Additionally, if you want to get slightly better quality for ,…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: bartowski
Теги: gguf, chat, en, zh, endpoints_compatible, imatrix, conversational
Лайков: 4  |  Загрузок: 362

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.