RedHatAI/Devstral-Small-2507-FP8-Dynamic - Каталог нейросетей
Генерация текста

RedHatAI/Devstral-Small-2507-FP8-Dynamic

Добавлено:
RedHatAI/Devstral-Small-2507-FP8-Dynamic

— Model Architecture: MistralForCausalLM — Input: Text — Output: Text — Model Optimizations: — Activation quantization: FP8 — Weight quantization: FP8 — Release Date: 08/28/2025 — Version: 1.0 — Model Developers: Red Hat (Neural Magic) This model was obtained by quantizing weights and activations of Devstral-Small-2507 to FP8 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%). Weight quantization also reduces disk size requirements by approximately 50%. This model was created with llm-compressor by running the code snippet below. This model can be deployed efficiently using the vLLM backend, as shown in the example below. The model was evaluated on popular coding tasks (HumanEval, HumanEval+, MBPP, MBPP+) via EvalPlus and vllm backend (v0.10.1.1). For evaluations, we run greedy sampling and report pass@1. The command to reproduce evals:

Модальности:
Генерация текста


Задача: Генерация текста
Автор: RedHatAI
Теги: mistral, neuralmagic, redhat, llmcompressor, quantized, FP8, compressed-tensors, en
Лайков: 4  |  Загрузок: 3,989

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.