— Model Architecture: Phi-3 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: INT4 — Intended Use Cases: Intended for commercial and research use in English. Similarly to Phi-3-medium-128k-instruct, this models is intended for assistant-like chat. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 7/11/2024 — Version: 1.0 — License(s): MIT — Model Developers: Neural Magic Quantized version of Phi-3-medium-128-instruct. It achieves an average score of 72.38 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 74.46. This model was obtained by quantizing the weights of Phi-3-medium-128k-instruct to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 25%. Only the weights of the linear operators within transformers blocks are quantized. Symmetric group-wise quantization is applied, in which a linear scaling per group maps the INT4 and floating point representations of the quantized weights. The GPTQ algorithm is…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: RedHatAI
Теги: phi3, conversational, custom_code, en, text-generation-inference, endpoints_compatible, compressed-tensors
Лайков: 3 | Загрузок: 1,216
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.