This is a roleplay and creative-writing merge of google/gemma-4-31B-it. I built it because the usual merge recipe was failing in a familiar way: it got louder before it got better. The model would pick up flavor from the sources, but long generations started to ramble, repeat, or lose the thread. That looked less like a bad source model and more like a bad operation. The bet here is that the useful part of a finetune is not just its delta magnitude. It is the direction it turns the base. I have argued elsewhere that by the time a transformer reaches its readout, the useful linearity has migrated to the residual-to-lmhead` interface, and that RMSNorm makes that interface primarily directional: magnitude is suppressed as an independent degree of freedom in the last step before unembedding. If that is true, it has a consequence for merging that I had not seen drawn out. Combining two fine-tunes by adding their weight deltas and renormalizing the result works against that geometry. The operation it rewards is a rotation of the base toward them. This is the first model I built on that principle. The body is composed with an operator I am calling an isometric merge. Each source delta is…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: maldv
Теги: gemma4, image-text-to-text, gemma, merge, isometric-merge, roleplay, creative-writing, conversational
Лайков: 4 | Загрузок: 70
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.