> VIDRAFT attention + Qwen3-4B / Gemma4-E4B FFN crossbreed. > A Qwen3-4B × Gemma4-E4B hybrid — NOT from-scratch. Private research checkpoint. The Gen-1 adapter’s FFN is reconstructed by cross-breeding Qwen3-4B FFN with Gemma4-E4B FFN (ratio 0.15), then the attention is re-healed (VIDRAFT) to adapt to the fused FFN. This carries the Gen-1 attention forward while blending a second model’s knowledge — so the result is not reducible to any single parent. — attention: VIDRAFT healing (Qwen3-4B based) — FFN: Qwen3-4B 85% ⊕ Gemma4-E4B 15% (bilinear inter projection 10240→9728, layer map 42→36) — structure: 2560 / 9728 / 36L (Qwen3-4B coordinates) — re-healing: 0.5B tokens, attention-only, LR 1e-5 Gen1 measured on 6 subjects. All numbers are base zero-shot** — instruction-following quality is expected from a later SFT stage (cf. Gemma4-E4B base 26.7% → it 69.4%). → After blending 15% Gemma4 FFN, performance is maintained / slightly above the Gen-1 baseline and Gemma4-E4B base. Gemma knowledge is visibly incorporated (multilingual facts, «Germany is Berlin / Italy is …»), and the intermediate English degradation is recovered by re-healing. — Some Korean repetition remains in greedy…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: FINAL-Bench
Теги: qwen3, darwin, darwin-v9, darwin-chimera, ffn-crossbreed, cross-architecture, evolutionary-merge, gemma4
Лайков: 4 | Загрузок: 34
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.