Hy3-JANG_2K-MTP
Quantized tencent/Hy3 for Apple Silicon MLX / JANG runtimes — a 295B-total / 21B-active text MoE, packed to...
Quantized tencent/Hy3 for Apple Silicon MLX / JANG runtimes — a 295B-total / 21B-active text MoE, packed to...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...
Meridian is a research-focused chat model built to work inside an app with live search, browsing, and source-reading...
> No matter your GPU. No matter your RAM. With ~4.5 GB of VRAM or unified memory free,...
Text generation model Модальности:Генерация текста Области применения:Диалог / чат Вызов функций (Tool use) Задача: Генерация текста Автор: roshangrewal...
We argue that efficient agentic reasoning benefits from decomposing deliberation into three interacting systems: reactive execution (System I)...
Hermes-style agentic fine-tune of Qwen3.6-27B, quantized to INT4 with a BF16 MTP overlay for speculative decoding. This model...
ToolOmni is a tool-use language model released for the ACL 2026 Main Conference paper ToolOmni: Enabling Open-World Tool...
kai-os/Carnice-9b quantized for hipfire, a Rust-native inference engine for AMD RDNA GPUs. Carnice is a Hermes tool-use finetune...
A fast, tiny function-calling model fine-tuned from LiquidAI/LFM2.5-1.2B-Instruct. Built on LFM2.5’s hybrid recurrent-attention architecture for significantly faster inference...