granite-4.2-3b-GGUF
Original model: https://huggingface.co/ibm-granite/granite-4.2-3b Model details: — Parameter count: 4B — Input support: text — Speculative decoding: no —...
Original model: https://huggingface.co/ibm-granite/granite-4.2-3b Model details: — Parameter count: 4B — Input support: text — Speculative decoding: no —...
Llama, but with integrated thinking and reasoning. Intelligent and self-proclaimed claude. Модальности:Генерация текста Области применения:Диалог / чат Задача:...
> | repeat-penalty 1.05 ✅ | correct (sweet spot) | > | repeat-penalty 1.0 | severe thinking loops...
LiteRT is Google’s on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch,...
> No matter your GPU. No matter your RAM. With ~4.5 GB of VRAM or unified memory free,...
Added a Jinja chat template so the model can format conversations correctly and work smoothly with mlx-lm chat-style...
> [!WARNING] > This model is a beta preview and a research artifact. It is not an open...
> TL;DR — A local Python-coding assistant that thinks before it codes. 8.25 GB, runs on one 16...
Этот репозиторий содержит квантования формата GGUF модели DavidAU/granite-4.1-8b-Claude-Opus-4.6-Thinking-MAX. Квантование проводили локально с использованием llama.cpp (сборка b9556). Модальности:Генерация текста...
Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI’s Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with...