michaelfeil/ct2fast-Llama-2-7b-hf - Каталог нейросетей
Генерация текста

michaelfeil/ct2fast-Llama-2-7b-hf

Добавлено:
michaelfeil/ct2fast-Llama-2-7b-hf

Speedup inference while reducing memory by 2x-4x using int8 inference in C++ on CPU or GPU. Checkpoint compatible to ctranslate2>=3.17.1 and hf-hub-ctranslate2>=2.12.0 — computetype=int8float16 for device=»cuda» — computetype=int8 for device=»cpu»` This is just a quantized version. Licence conditions are intended to be idential to original huggingface repo. Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 7B pretrained model, converted for the Hugging Face Transformers format. Links to other models can be found in the index at the bottom. Meta developed and publicly released the Llama 2 family of large language models (LLMs), a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. Our fine-tuned LLMs, called Llama-2-Chat, are optimized for dialogue use cases. Llama-2-Chat models outperform open-source chat models on most benchmarks we tested, and in our human evaluations for helpfulness and safety, are on par with some popular closed-source models like ChatGPT and PaLM. Variations Llama 2 comes in a…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: michaelfeil
Теги: llama, ctranslate2, int8, float16, facebook, meta, llama-2, en
Лайков: 3  |  Загрузок: 10

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.