Kimuraxhalu/gemma-4-12B-coder-fable5-composer2.5-MTP-NVFP4 - Каталог нейросетей
Генерация текста

Kimuraxhalu/gemma-4-12B-coder-fable5-composer2.5-MTP-NVFP4

Добавлено:
Kimuraxhalu/gemma-4-12B-coder-fable5-composer2.5-MTP-NVFP4

> TL;DR — A local Python-coding assistant that thinks before it codes. 8.25 GB, runs on one 16 GB Blackwell GPU, native in vLLM (no —quantization flag). Bundled speculative-decode draft included. 💚 Strong at: hard algorithms (DP, graphs, Fenwick/segment trees, bitmask DP), bug-fixing & refactoring (accurate root-cause + genuine O(n²)→O(n) rewrites that preserve semantics), and faithful open reasoning that matches the emitted code. Japanese prompts cause no measurable Python-quality drop. > ⚠️ Know the one sharp edge (verified): on quant / time-series code it can write a look-ahead bias (e.g. an unshifted position × a forward-shifted return), and its reasoning sometimes states the correct rule while the code does the opposite. Do not ship its pandas/numpy back-test or accounting code unreviewed — gate it. It’s a superb algorithm/debug specialist, not an unsupervised quant author. Aggregate throughput (no spec-decode; turn MTP off for batch): Choosing a layout on a fixed GPU budget: TP=4 gives the lowest latency, but TP=2 is more efficient per GPU (≈316 vs 195 tok/s/GPU at 16-way). For max farm throughput, run two data-parallel TP=2 replicas (≈1.3k tok/s on 4 GPUs) instead of one…

Модальности:
Генерация текста

Области применения:
Генерация кода Логика и рассуждение Диалог / чат

Языки программирования:
Python


Задача: Генерация текста
Автор: Kimuraxhalu
Теги: gemma4_unified, image-text-to-text, gemma4, coding, code, python, reasoning, thinking
Лайков: 4  |  Загрузок: 194

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.