This model depends on the following PR merge status and mlx-ml release status: https://github.com/ml-explore/mlx-lm/pull/1261 You can either clone my fork: https://github.com/kyr0/mlx-lm/tree/feat/zaya-support or wait for mainline support for this model being merged by the MLX-ML team. ZAYA1 is an 800m active/8.3B total parameter MoE model, and the first trained entirely end-to-end on AMD’s hardware, software, and networking stack. Our ZAYA1 base model benchmark performance is extremely competitive with the SoTA Qwen3 series of models of comparable scale, and outperforms comparable western open-source models such as SmolLM3, and Phi4. ZAYA1-base excels especially at complex and challenging mathematical and STEM reasoning tasks, nearly matching the performance of SoTA Qwen3 thinking models under high pass@k settings even prior to explicit post-training for reasoning, and exceeds other strong reasoning models such as Phi4-reasoning, and Deepseek-R1-Distill. Details of our pretraining efforts, hardware specific optimizations, and ZAYA1 base model benchmarks are described in the accompanying technical report. ZAYA1’s architecture includes several innovations developed at Zyphra. These…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: kyr0
Теги: mlx, zaya, conversational, 8-bit
Лайков: 4 | Загрузок: 140
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.