steelman-14b-ada-GGUF
Quantized GGUF of Steelman-14B-Ada v0.3 for use with Ollama, llama.cpp, or any GGUF-compatible runtime. Legacy GGUFs from prior...
Quantized GGUF of Steelman-14B-Ada v0.3 for use with Ollama, llama.cpp, or any GGUF-compatible runtime. Legacy GGUFs from prior...
NVFP4-quantized version of Qwen/Qwen2.5-72B-Instruct, produced by Enfuse. — Full NVFP4 (W4A4): Requires NVIDIA Blackwell GPU (B200, GB200, RTX...
This repository provides GGUF quantized versions of the Qwen/Qwen3.5-9B model, optimized for local execution using llama.cpp and compatible...
W4A16 (INT4 weights, BF16 activations) quantized version of OpenMOSE/Qwen3.5-REAP-262B-A17B, created using AutoRound. — Calibration dataset: NeelNanda/pile-10k (64 samples,...
Это модель Nanbeige4.1-3B, преобразованная в формат MLX с 4-разрядным квантованием (аффин, размер группы=64) для эффективного вывода на Apple...
Inventor: Konstantin Vladimirovich Grabko Email: grabko@cmsmanhattan.com Date: December 21, 2025 Needed: A sponsor for Llama 405b Distilation to...
— Development: Code assistance and generation — Writing: Content creation and editing — Analysis: Document summarization — Chat:...
4-bit quantized GGUF version of distil-qwen3-4b-text2sql for efficient local inference. Only 2.5GB — runs on most laptops and...
WeDLM is an 8B parameter instruction-tuned model by Tencent, supporting English and Chinese. It features QK Norm architecture...
The Model Wwayu/GLM-4.7-PRISM-mlx-2Bit was converted to MLX format from Ex0bit/GLM-4.7-PRISM using mlx-lm version 0.28.3. Модальности:Генерация текста Области применения:Диалог...