This repository hosts the Phi4-mini-instruct model quantized with torchao using int4 weight-only quantization and the awq algorithm. This work is brought to you by the PyTorch team. This model can be used directly or served using vLLM for 56% VRAM reduction (3.95 GB needed) and 1.17x speedup on H100 GPUs. The model is calibrated with 2 samples from mmlupro task to recover the accuracy for mmlupro specifically. It recovered accuracy from mmlupro` from INT4 checkpoint from 36.98 to 43.13, while bfloat16 baseline accuracy is 46.43. Install vllm nightly and torchao nightly to get some recent changes: and use a token with write access, from https://huggingface.co/settings/tokens We rely on lm-evaluation-harness to evaluate the quality of the quantized model. Here we only run on mmlu for sanity check. INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED BUT SOME MAY HAVE NO CONTENT
Модальности:
Генерация текста
Области применения:
Генерация кода Математика Диалог / чат Мультиязычность Следование инструкциям
Задача: Генерация текста
Автор: pytorch
Теги: phi3, torchao, phi, phi4, nlp, code, math, chat
Лайков: 4 | Загрузок: 233
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.