Phi4-mini is quantized by the PyTorch team using torchao with 8-bit embeddings and 8-bit dynamic activations with 4-bit weight linears (INT8-INT4). The model is suitable for mobile deployment with ExecuTorch. We provide the quantized pte for direct use in ExecuTorch. (The provided pte file is exported with at maxseqlength/maxcontextlength of 1024; if you wish to change this, re-export the quantized model following the instructions in Exporting to ExecuTorch.) The pte file can be run with ExecuTorch on a mobile phone. See the instructions for doing this in iOS. On iPhone 15 Pro, the model runs at 17.3 tokens/sec and uses 3206 Mb of memory. ⚠️ Caveat: Our mobile demo apps have regressed support for the Phi-4 tokenizer, so this model will not currently run in our official apps. If you are using your own app and runner, you can still load and run the .pte file successfully. See https://github.com/pytorch/executorch/issues/14077 for details and tracking. We want to quantize the embedding and lm_head differently. Since those layers are tied, we first need to untie the model: and use a token with write access, from https://huggingface.co/settings/tokens We rely on…
Модальности:
Генерация текста
Области применения:
Генерация кода Математика Диалог / чат Мультиязычность Следование инструкциям
Задача: Генерация текста
Автор: pytorch
Теги: executorch, phi3, torchao, phi, phi4, nlp, code, math
Лайков: 3 | Загрузок: 828
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.