This repository contains the quantized AMD Q4NX weights, AIE-ML firmware kernels, and runtime configuration for running openbmb/MiniCPM5-2B natively on the AMD XDNA 2 NPU (/dev/accel/accel0, 48 AIE-ML tiles) using FastFlowLM. > 📦 Porting Pipeline & Diagnostic Harnesses: The full conversion scripts, architectural adaptation code, and 42-layer ERT timeout reproduction suite are hosted on GitHub at julianmb/minicpm5-xdna2. > [!IMPORTANT] > Testing Disclaimer: While the compiled AIE kernels and FastFlowLM runtime target all XDNA 2 processors uniformly, this port was specifically benchmarked, tuned, and verified on the AMD Strix Halo (Ryzen AI Max+ 395 w/ Radeon 8060S, 128 GB UMA). Community feedback and verification reports on Strix Point and Krackan Point laptops are warmly welcomed. FastFlowLM requires compiled AIE kernels to reside in its xclbins/ directory. Copy or symlink them: Add the model definition to your FastFlowLM modellist.json (located in ~/.config/flm/modellist.json or /opt/flm/modellist.json`): Cause: FastFlowLM’s C++ loader requires explicit integer token IDs in tokenizerconfig.json. Fix: Fixed in the latest repository commit. Run git pull in your MiniCPM5-2B-NPU2…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: julianmb
Теги: qwen3, fastflowlm, xdna2, npu, amd, strix-halo, strix-point, krackan
Лайков: 4 | Загрузок: 3
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.