This is quantized version of nvidia/Llama3-ChatQA-2-8B created using llama.cpp We introduce Llama3-ChatQA-2, a suite of 128K long-context models, which bridges the gap between open-source LLMs and leading proprietary models (e.g., GPT-4-Turbo) in long-context understanding and retrieval-augmented generation (RAG) capabilities. Llama3-ChatQA-2 is developed using an improved training recipe from ChatQA-1.5 paper, and it is built on top of Llama-3 base model. Specifically, we continued training of Llama-3 base models to extend the context window from 8K to 128K tokens, along with a three-stage instruction tuning process to enhance the model’s instruction-following, RAG performance, and long-context understanding capabilities. Llama3-ChatQA-2 has two variants: Llama3-ChatQA-2-8B and Llama3-ChatQA-2-70B. Both models were originally trained using Megatron-LM, we converted the checkpoints to Hugging Face format. For more information about ChatQA 2, check the website! Llama3-ChatQA-2-70B Evaluation Data Training Data Website Paper We evaluate ChatQA 2 on short-context RAG benchmark (ChatRAG) (within 4K tokens), long context tasks from SCROLLS and LongBench…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, nvidia, chatqa-2, chatqa, llama-3, en, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 681
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.