deepseek-math-7b-lean-prover-dpo-olmo-3
This model is a fine-tuned version of formalmathatepfl/deepseek-math-7B-finetuned. It has been trained using TRL. This model was trained...
This model is a fine-tuned version of formalmathatepfl/deepseek-math-7B-finetuned. It has been trained using TRL. This model was trained...
The tune optimized for two things: — bringing warmth, emotional intelligence, general chat improvement to Qwen 3.5 series...
static quants of https://huggingface.co/RISys-Lab/RedSage-Qwen3-8B-DPO For a convenient overview and download list, visit our model page for this model....
Эта модель представляет собой точно настроенную версию allenai/Olmo-3.1-32B-Instruct, использующую humanline dpo для улучшения возможностей письма, ролевой игры и...
 ![Language]()  The model was sculpted through a rigorous multi-stage process: Objective: Instill...
Sungur-9B is a Turkish-specialized large language model derived from ytu-ce-cosmos/Turkish-Gemma-9b-v0.1, which itself is based on Gemma-2-9b. The model...
Lightweight medical finetune on top of Arcee’s AFM-4.5B for education and research use. Trained using a straightforward 3-step...
liberalis-cogitator-llama-3.1-8b is not just a machine for words — it is a forge for ideas. With 8 billion...
InfiAlign — это масштабируемая и эффективная в отношении данных структура пост-обучения, которая сочетает в себе контролируемую точную настройку...
This is a Japanese version of the Qwen/Qwen2.5-7B-Instruct model which was DPO trained using synthetic Japanese conversation data....