Villanova-2B-Base-2603
Villanova — это семейство полностью открытых многоязычных многоязычных моделей (LLM), ориентированных на пять основных европейских языков. Все веса...
Villanova — это семейство полностью открытых многоязычных многоязычных моделей (LLM), ориентированных на пять основных европейских языков. Все веса...
JOSIE-1.1-4B-Instruct is a full-weight fine-tuned instruction-following model built on Qwen3-4B-Instruct, optimized for natural conversational interactions, problem-solving, and everyday...
JOSIE-1.1-4B-Thinking is a full-weight fine-tuned reasoning model built on Qwen3-4B-Thinking, optimized for extended context logical reasoning, mathematics, STEM...
Trinity Nano Preview is a preview of Arcee AI’s 6B MoE model with 1B active parameters. It is...
The post-training data has a cutoff date of November 28, 2025. The pre-training data has a cutoff date...
This model mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit was converted to MLX format from nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 using mlx-lm version 0.29.0. Модальности:Генерация текста Области применения:Диалог...
> [!NOTE] > Includes Unsloth chat template fixes! For llama.cpp, use —jinja > Unsloth Dynamic 2.0 achieves superior...
MPOA (Magnitude-Preserving Othogonalized Ablation, AKA norm-preserving biprojected abliteration) has been applied the many of the layers in this...
— Model Architecture: Qwen/Qwen3-235B-A22B-Instruct-2507 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP4 —...
Advanced Hybrid Reasoning Model with Tool-Calling Capabilities SAGE Reasoning Family Models are instruction-tuned, text-in/text-out generative systems released under...