RedHatAI/Qwen3-VL-235B-A22B-Instruct-FP8-dynamic - Каталог нейросетей
Генерация текста

RedHatAI/Qwen3-VL-235B-A22B-Instruct-FP8-dynamic

Добавлено:
RedHatAI/Qwen3-VL-235B-A22B-Instruct-FP8-dynamic

— Model Architecture: Qwen3VLMoeForConditionalGeneration — Input: Text/Image/Video — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Release Date: 09/28/202510 — Version: 1.0 — Model Developers:: Red Hat This model was obtained by quantizing the weights and activations of Qwen/Qwen3-VL-235B-A22B-Instruct to FP8 data type. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks of the language model are quantized. This model was quantized using the llm-compressor library as shown below. The model was evaluated on the OpenLLMv1 leaderboard task, using lm-evaluation-harness, on reasoning tasks using lighteval and on vision tasks using lmms-eval. vLLM was used for all evaluations. Category Metric Qwen3-VL-235B-A22B-Instruct Qwen3-VL-235B-A22B-Instruct-FP8-dynamic Recovery (%) OpenLLM V1 ARC-Challenge (Acc-Norm, 25-shot) 76.54 75.94 99.2 GSM8K (Strict-Match, 5-shot) 90.30 89.92 99.6 HellaSwag (Acc-Norm, 10-shot) 87.81 87.74 99.9 MMLU (Acc, 5-shot) 87.11 87.23…

Модальности:
Генерация текста Мультимодальность

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: RedHatAI
Теги: qwen3_vl_moe, fp8, quantized, llm-compressor, compressed-tensors, red hat, image, video
Лайков: 4  |  Загрузок: 29,004

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.