llama-2-7b-chat-guanaco-hf-4bit
— BaseModel: Meta’s Llama 2 7B chat hf. — Dataset: timdettmers/openassistant-guanaco. We are unlocking the power of large...
— BaseModel: Meta’s Llama 2 7B chat hf. — Dataset: timdettmers/openassistant-guanaco. We are unlocking the power of large...
Speedup inference while reducing memory by 2x-4x using int8 inference in C++ on CPU or GPU. Checkpoint compatible...
MobileMoE — это семейство языковых моделей Mixture-of-Experts (MoE) на устройстве с субмиллиардными активными параметрами, предназначенных для продвижения границы...
— 💬 Enhanced Conversational Abilities: Fine-tuned on FineTome-100k for natural, engaging dialogue — 🚀 Efficient & Fast: —...
We are introducing MobileLLM-P1 or Pro, a 1B foundational language model in the MobileLLM series, designed to deliver...
The NeuraLake iSA-03-Mini-3B (Hybrid) is an advanced AI model developed by NeuraLake, specifically designed to integrate the best...
Наш последний метод квантования вводит прецизионно-адаптивное квантование для сверхнизкоразрядных моделей (1-2 бита) с проверенными улучшениями на Llama-3-8B. Этот...
This model was generated using llama.cpp at commit f5cd27b7. Our latest quantization method introduces precision-adaptive quantization for ultra-low-bit...
Наш последний метод квантования вводит прецизионно-адаптивное квантование для сверхнизкоразрядных моделей (1-2 бита) с проверенными улучшениями на Llama-3-8B. Этот...
We have a free Google Colab notebook for turning Llama 3.1 (8B) into a reasoning model: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1_(8B)-GRPO.ipynb All...