wepiqx/gemma-4-12B-OBLITERATED-SHQ7-GGUF - Каталог нейросетей
Генерация текста

wepiqx/gemma-4-12B-OBLITERATED-SHQ7-GGUF

Добавлено:
wepiqx/gemma-4-12B-OBLITERATED-SHQ7-GGUF

> OBLITERATED + selective IQ hybrid for 8 GB Pascal GPUs. Uncensored coding model. 7.7% smaller than Q4KM at the same quality. > > WARNING: This is an uncensored model. It may generate content that other models refuse. Use responsibly. The maintainer assumes no responsibility for misuse. > > Note: The model file is named SHQ7-IQ4XS.gguf for HF parser compatibility. This is NOT a pure IQ4XS quantization — it’s a hybrid using IQ4NL, Q4K, IQ4XS, and Q80 across different layers. Check the config for exact per-tensor types. > > Primarily optimized for coding tasks. May not perform well on general knowledge or non-code questions. A hand-optimized hybrid quantization of OBLITERATUS/Gemma-4-12B-OBLITERATED for 8 GB VRAM GPUs (tested on GTX 1070, works on any GPU with enough VRAM). > I’ve put a lot of work into hand-tuning these quants — let me know how they run on your hardware! Drop a comment or open a discussion with your setup and any feedback. TL;DR: Custom mixed quantization using IQ4NL on boundary FFN + Q4K on critical middle FFN + IQ4XS on all other FFN + Q80 on all attention. Fits in 8 GB VRAM with room for KV cache. > Q4KM baseline quantized with mradermacher’s imatrix (i1…

Модальности:
Генерация текста

Области применения:
Генерация кода Диалог / чат


Задача: Генерация текста
Автор: wepiqx
Теги: gguf, gemma-4, gemma, google, 12b, quantized, quantization, llama-cpp
Лайков: 4  |  Загрузок: 681

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.