brucethemoose/CaPlatTessDolXaBoros-34B-200K-exl2-4bpw-fiction - Каталог нейросетей
Генерация текста

brucethemoose/CaPlatTessDolXaBoros-34B-200K-exl2-4bpw-fiction

Добавлено:
brucethemoose/CaPlatTessDolXaBoros-34B-200K-exl2-4bpw-fiction

Dolphin-2.2-yi-34b-200k, Nous-Capybara-34B, Tess-M-v1.4, Airoboros-31-yi-34b-200k, PlatYi-34B-200K-Q, and Una-xaberius-34b-v1beta** merged with a new, experimental implementation of «dare ties» via mergekit. Quantized with the git version of exllamav2 with 200 rows (400K tokens) on a long Orca-Vicuna format chat, a selected sci fi story and a fantasy story. This should hopefully yield better chat/storytelling performance than the short, default wikitext quantization. 4bpw is enough for ~47K context on a 24GB GPU. I would highly recommend running in exui for speed at long context. I go into more detail in this Reddit post Merged with the following config, and the tokenizer from chargoddard’s Yi-Llama: Various densities were tested with perplexity tests and high context prompts. Relatively high densities seem to perform better, contrary to the findings of the Super Mario paper. Dare Ties is also resulting in seemingly better, lower perplexity merges than a regular ties merge, task arithmetic or a slerp merge. Xaberuis is not a 200K model, hence it was merged at a very low density to try and preserve Yi 200K’s long context performance while still inheriting some of Xaberius’s…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: brucethemoose
Теги: llama, text-generation-inference, merge, en, endpoints_compatible
Лайков: 3  |  Загрузок: 11

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.