Llama-2-7b-hf-8bit
Модель генерации текста Модальности:Генерация текста Задача: Генерация текста Автор: vivekraina Теги: llama, text-generation-inference, endpoints_compatible, 8-bitЛайков: 3 | Загрузок:...
Модель генерации текста Модальности:Генерация текста Задача: Генерация текста Автор: vivekraina Теги: llama, text-generation-inference, endpoints_compatible, 8-bitЛайков: 3 | Загрузок:...
MPT-30B — это трансформатор в стиле декодера, предварительно обученный с нуля на токенах 1T английского текста и кода....
This repository contains a LLaMA-7B further fine-tuned model on conversations and question answering prompts. ⚠️ I used LLaMA-7b-hf...
3B model converted to 8Bit by rockerBOO. May require bitsandbytes dependency. Tested on a 2080 8GB. StableLM-Tuned-Alpha is...
MLX (mlx-vlm tree) 8-bit build of GLM-5.3-Flash (320B-A18B, glm5next), converted by streaming dequant of the official FP8 release...
Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro,...
A surgically expert-pruned openai/gpt-oss-120b: 48 of 128 experts kept per layer (top-48 by REAP saliency), reducing the checkpoint...
This model depends on the following PR merge status and mlx-ml release status: https://github.com/ml-explore/mlx-lm/pull/1261 You can either clone...
MLX conversion of Qwen/Qwen3.6-27B, affine 8-bit (group_size 64), with the native Multi-Token-Prediction (MTP) head embedded in the main...
Модель генерации текста Модальности:Генерация текста Области применения:Диалог / чат Задача: Генерация текста Автор: mlx-community Теги: mlx, glm_moe_dsa, conversational,...