Qwen3-1.7B-FC
A function calling model based on Qwen3-1.7B, fine-tuned using RLVR (Reinforcement Learning with Verifiable Rewards) to improve tool-use...
A function calling model based on Qwen3-1.7B, fine-tuned using RLVR (Reinforcement Learning with Verifiable Rewards) to improve tool-use...
We introduce RL-Struct, a lightweight Reinforcement Learning framework designed to solve the «Structure Gap»—the tension between probabilistic token...
Granite 4.0 Micro (≈3B) tuned for medical education & instruction following. Recipe: JEPA-LLM SFT on medmcqa-hard + personas...
Welcome to NeuroSpectr13B NeuroSpectr13B is a large language model designed to reflect and introspect, aiming to enhance understanding...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
Lightweight medical finetune on top of Arcee’s AFM-4.5B for education and research use. Trained using a straightforward 3-step...
This LoRA adapter enhances google/gemma-3-1b-it with structured reasoning capabilities using tags. Trained with GRPO (Group Relative Policy Optimization)...
Gguf version. 3.91B parameters, 2 experts active, 4 in total. This is the non-experimental version of Superthoughts Lite...
Улучшенная модель на Qwen2.5-1.5B-Thinking-v1.1. Он был обучен с использованием TRL. Эта модель была обучена с помощью GRPO, метода,...
SciJudge-4B-2605 — это модель Qwen3-4B-Instruct-2507, специально настроенная для оценки научных статей. Учитывая названия, аннотации и даты публикации двух...