This repository presents post FP8 quantized Falcon-H1R-7B-FP8 via NVIDIA Model Optimizer, enabling efficient inference while preserving the strong reasoning introduced in the paper Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling. Falcon-H1R-7B was trained via cold-start supervised fine-tuning with long reasoning traces and further enhanced by scaling RL with GRPO. The model demonstrates outstanding performance across various benchmark evaluations, including mathematics, programming, instruction following, and general logic. — Developed by: Technology Innovation Institute — Model type: Causal decoder-only — Architecture: Hybrid (Transformers + Mamba2) architecture — Language(s): English, Multilingual — License: Falcon-LLM License For more details on FP8 post-quantization for this model, please refer to the Falcon-H1R-FP8 technical blogpost. For more details about the training protocol of this model, please refer to the Falcon-H1R technical blogpost and Technical Report. Currently to use this model, you can either rely on Hugging Face transformers, vLLM or SGLang library. Make sure to install the latest version of transformers or vLLM…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: tiiuae
Теги: falcon_h1, falcon-h1r, conversational, en, endpoints_compatible, modelopt
Лайков: 4 | Загрузок: 228
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.