Qwythos-9B-Claude-Mythos-5-1M-GGUF
GGUF quantizations of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. Qwythos-9B is a...
GGUF quantizations of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. Qwythos-9B is a...
This model is a custom-code derivative of AxiomicLabs/GPT-X2-125M, adapted for experimental long-context causal language modeling and architecture research....
We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Overall...
This repository contains the δ-mem TSW adapter for Qwen/Qwen3-4B-Instruct-2507, as presented in the paper δ-mem: Efficient Online Memory...
GKA-primed-HQwen3-8B-Reasoner — это гибридная языковая модель, состоящая из 50% слоев внимания и 50% слоев стробированного KalmaNet (GKA), загрунтованных...
InCoder-32B (Industrial-Coder-32B) is the first 32B-parameter code foundation model purpose-built for industrial code intelligence. While general code LLMs...
A natively recursive language model based on Qwen3.5-35B-A3B, trained with Rejection Sampling SFT (RS-SFT) to solve long-context tasks...
Today we’re releasing DeepBrainz-R1, a family of reasoning-first Small Language Models (SLMs) designed for agentic AI systems in...
If you are using GENERator for sequence generation, please ensure that the length of each input sequence is...
Модель генерации текста Модальности:Генерация текста Области применения:Диалог / чат Задача: Генерация текста Автор: byroneverson Теги: llama, llm, long...