Phi-4-mini-instruct-INT8-INT4
Phi4-mini is quantized by the PyTorch team using torchao with 8-bit embeddings and 8-bit dynamic activations with 4-bit...
Phi4-mini is quantized by the PyTorch team using torchao with 8-bit embeddings and 8-bit dynamic activations with 4-bit...
This repository hosts the Phi4-mini-instruct model quantized with torchao using int4 weight-only quantization and the awq algorithm. This...
HuggingFaceTB/SmolLM3-3B квантуется с использованием Torchao с 8-битными вложениями и 8-битными динамическими активациями с 4-битными линейными весами (INT8-INT4). Затем...