MiroThinker-v1.0-72B-AWQ-4bit
— Quantization Method: AWQ — Bits: 4 — Group Size: 32 — Calibration Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset — Quantization Tool:...
— Quantization Method: AWQ — Bits: 4 — Group Size: 32 — Calibration Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset — Quantization Tool:...
Breeze-3B is a specialized coding model fine-tuned on Qwen/Qwen2.5-Coder-3B-Instruct to automatically resolve Git merge conflicts with reasoning and...
“Local-first 3B model for VR / game companions that outputs strict {dialog, intent, microplan} JSON from a CONTEXT...
Welcome to NeuroSpectr13B NeuroSpectr13B is a large language model designed to reflect and introspect, aiming to enhance understanding...
JET-1.5B is designed to improve the efficient reasoning of LLMs by training the base DeepSeek-Distill-Qwen-1.5B model with a...
Nytheria-3B is a masterfully fine-tuned version of the Qwen/Qwen2.5-3B-Instruct model, engineered from the ground up to be a...
> Пальмирские мини-модели демонстрируют исключительные возможности в сложных областях рассуждений и решения математических задач. Его эффективность особенно примечательна...
Fleming-R1 — это модель рассуждений для медицинских сценариев, которая может выполнять пошаговый анализ сложных проблем и давать надежные...
Большая языковая модель Мин (Ming‑LLM) — это предметно-специализированная LLM для энергетического сектора. — Мы выпускаем как базовую модель,...
Qwen-2.5-7B-ConsistentChat is a 7B instruction-tuned chat model focused on multi-turn consistency. It is fine-tuned from the Qwen/Qwen2.5-7B base...