Gradient incorporates your data to deploy autonomous assistants that power critical operations across your business. If you’re looking to build custom AI models or agents, email us a message contact@gradient.ai. This model extends LLama-3 70B’s context length from 8k to > 524K, developed by Gradient, sponsored by compute from Crusoe Energy. It demonstrates that SOTA LLMs can learn to operate on long context with minimal training by appropriately adjusting RoPE theta. We trained on 210M tokens for this stage, and ~400M tokens total for all stages, which is < 0.003% of Llama-3's original pre-training data. — meta-llama/Meta-Llama-3-70B-Instruct as the base — NTK-aware interpolation [4] following scaling laws [2] to set optimal schedule for RoPE theta — Progressive training on increasing context lengths, similar to Large World Model [1] (See details below) We build on top of the EasyContext Blockwise RingAttention library [5] to scalably and efficiently train on very long contexts on Crusoe Energy high performance L40S cluster. We layered parallelism on top of Ring Attention with a custom network topology to better leverage large GPU clusters in the face of network bottlenecks from…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: LoneStriker
Теги: gguf, meta, llama-3, en, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 123
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.