moonshotai_Kimi-K2.6-GGUF
— llama.cpp — ramalama — LM Studio — koboldcpp — Jan AI — Text Generation Web UI —...
— llama.cpp — ramalama — LM Studio — koboldcpp — Jan AI — Text Generation Web UI —...
> [!IMPORTANT] > Уведомление о наименовании (2026-04-10). Метод «HLWQ», используемый в этой модели, переименовывается в HLWQ (Hadamard-Lloyd Weight...
This is an FP8-quantized version of dnotitia/DNA-2.0-14B, optimized for efficient inference by DLM (Data Science Lab., Ltd.). FP8...
NVFP4-quantized version of Qwen/Qwen2.5-72B-Instruct, produced by Enfuse. — Full NVFP4 (W4A4): Requires NVIDIA Blackwell GPU (B200, GB200, RTX...
Plano-Orchestrator is a family of state-of-the-art routing and orchestration models that decide which agent(s) or LLM(s) should handle...
Orchestrator-8B is a state-of-the-art 8B parameter orchestration model designed to solve complex, multi-turn agentic tasks by coordinating a...
Orchestrator-8B is a state-of-the-art 8B parameter orchestration model designed to solve complex, multi-turn agentic tasks by coordinating a...
Сегодня мы выпускаем и с открытым исходным кодом MiniMax-M2, модель Mini, созданную для кодирования Max Имея всего 10...
Kimi K2 Thinking — это новейшая, наиболее способная версия модели мышления с открытым исходным кодом. Начиная с Kimi...
Kimi Linear: An Expressive, Efficient Attention Architecture (a) On MMLU-Pro (4k context length), Kimi Linear achieves 51.0 performance...