A local AI team, built around AMD Strix Halo. Ornith1.5 Ciru Halo Agent combines a custom quantization of Ornith’s 35B-A3B mixture-of-experts model with a purpose-built vLLM/ROCm runtime for fast coding, tool use, and concurrent agents. The design starts with the hardware: packed four-bit weights, four-bit activation paths, Strix Halo four-bit matrix instructions, specialized kernels, and adaptive DFlash2 speculative decoding. Prefix caching and a shared memory pool let agents return to long working histories. — 178 tok/s single-request decode on the ten-question coding speed screen, with 166 ms mean time to first token. — 295 tok/s aggregate at eight concurrent requests, completing the ten-question batch in 5.52 seconds. — 1,287 tok/s cold prefill at 64K and 668 tok/s near 256K—1.62× and 2.71× the recorded Q4KXL prefill rates, respectively. — 256K request context capacity, eight active sequences, and a 44 GiB shared KV/state pool. — 123 tok/s cached C1 decode at 63K history, with all ten return-task health checks passing. These figures describe specific workloads, rather than a universal generation rate. The benchmark tables below include quality results and the workloads where…
Модальности:
Генерация текста Компьютерное зрение
Области применения:
Вызов функций (Tool use)
Задача: Генерация текста
Автор: jcbtc
Теги: vllm, ornith, ciru, amd, strix-halo, gfx1151, rocm, agentic
Лайков: 4 | Загрузок: 0
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.