michaelfeil/ct2fast-mpt-30b - Каталог нейросетей
Генерация текста

michaelfeil/ct2fast-mpt-30b

Добавлено:
michaelfeil/ct2fast-mpt-30b

Speedup inference while reducing memory by 2x-4x using int8 inference in C++ on CPU or GPU. Checkpoint compatible to ctranslate2>=3.16.0 and hf-hub-ctranslate2>=2.12.0 — computetype=int8float16 for device=»cuda» — computetype=int8 for device=»cpu»` This is just a quantized version. Licence conditions are intended to be idential to original huggingface repo. MPT-30B is a decoder-style transformer pretrained from scratch on 1T tokens of English text and code. This model was trained by MosaicML. MPT-30B is part of the family of Mosaic Pretrained Transformer (MPT) models, which use a modified transformer architecture optimized for efficient training and inference. MPT-30B comes with special features that differentiate it from other LLMs, including an 8k token context window (which can be further extended via finetuning; see MPT-7B-StoryWriter), support for context-length extrapolation via ALiBi, and efficient inference + training via FlashAttention. It also has strong coding abilities thanks to its pretraining mix. MPT models can also be served efficiently with both standard HuggingFace pipelines and NVIDIA’s FasterTransformer. The size of MPT-30B was also specifically chosen to make…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: michaelfeil
Теги: mpt, ctranslate2, int8, float16, Composer, MosaicML, llm-foundry, StreamingDatasets
Лайков: 3  |  Загрузок: 20

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.