TheBloke/Chronos-13B-SuperHOT-8K-fp16 - Каталог нейросетей
Генерация текста

TheBloke/Chronos-13B-SuperHOT-8K-fp16

Добавлено:
TheBloke/Chronos-13B-SuperHOT-8K-fp16

Chat & support: my new Discord server Want to contribute? TheBloke’s Patreon page This is fp16 pytorch format model files for Elinas’ Chronos 13B merged with Kaio Ken’s SuperHOT 8K. Kaio Ken’s SuperHOT 13b LoRA is merged on to the base model, and then 8K context can be achieved during inference by using trustremotecode=True. Note that config.json has been set to a sequence length of 8192. This can be modified to 4096 if you want to try with a smaller sequence length. 4-bit GPTQ models for GPU inference 2, 3, 4, 5, 6 and 8-bit GGML models for CPU inference Unquantised SuperHOT fp16 model in pytorch format, for GPU inference and for further conversions Unquantised base fp16 model in pytorch format, for GPU inference and for further conversions Then run the following code. config.json has been default to a sequence length of 8192, but you can also configure this in your Python code. The provided modelling code, activated with trustremotecode=True will automatically set the scale parameter from the configured maxpositionembeddings. Eg for 8192, scale is set to 4. Provided in the repo is llamaropescaledmonkeypatch.py, written by @kaiokendev. It can be theoretically be added to any…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: TheBloke
Теги: llama, custom_code, text-generation-inference
Лайков: 3  |  Загрузок: 11

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.