This model was generated using llama.cpp at commit 064cc596. Our latest quantization method introduces precision-adaptive quantization for ultra-low-bit models (1-2 bit), with benchmark-proven improvements on Llama-3-8B. This approach uses layer-specific strategies to preserve accuracy while maintaining extreme memory efficiency. All tests conducted on Llama-3-8B-Instruct using: — Standard perplexity evaluation pipeline — 2048-token context window — Same prompt set across all quantizations — Dynamic Precision Allocation: — First/Last 25% of layers → IQ4XS (selected layers) — Middle 50% → IQ2XXS/IQ3S (increase efficiency) — Critical Component Protection: — Embeddings/output layers use Q5K — Reduces error propagation by 38% vs standard 1-2bit Key: — PPL = Perplexity (lower is better) — Δ PPL = Percentag
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: Mungert
Теги: gguf, endpoints_compatible, conversational
Лайков: 4 | Загрузок: 243
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.