LongAlign-13B-64k-base
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for...
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for...
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for...
Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported...
MLX 4-bit conversion of Nex-N2.5-mini, a sparse MoE language model for local inference, coding, reasoning, and long-context work....
Построено из исходных весов InclusionAI с нашей собственной матрицей важности. Калибровочные корпуса за нашими сборками являются общедоступными. Ling...
> [!IMPORTANT] > Совместимость: для этих файлов GGUF требуется сборка llama.cpp с поддержкой архитектуры K2 Horizon. До тех...
A reproducible deployment and benchmark package for the official deepseek-ai/DeepSeek-V4-Flash-0731 checkpoint on two NVIDIA GB10-class nodes using vLLM...
Original model: https://huggingface.co/empero-ai/Qwythos-9B-v2 — llama.cpp — ramalama — LM Studio — koboldcpp — Jan AI — Text Generation...
This repository contains a bfloat16 MLX conversion of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for Apple Silicon inference with MLX, MLX-LM, MLX-VLM, and...
NVFP4 (NVIDIA FP4, weight-only NVFP4A16) quantization of empero-ai/Qwythos-9B-Claude-Mythos-5-1M — a Claude-Mythos/Fable-trace reasoning fine-tune of Qwen3.5-9B (qwen35: a dense...