claude-llama-8B-think-mlx-4Bit
Llama, but with integrated thinking and reasoning. Intelligent and self-proclaimed claude. Модальности:Генерация текста Области применения:Диалог / чат Задача:...
Llama, but with integrated thinking and reasoning. Intelligent and self-proclaimed claude. Модальности:Генерация текста Области применения:Диалог / чат Задача:...
The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of...
Work in progress of an RP finetune: — 60% cooked — Pretty uncensored, but NSFW needs prompting to...
RWKV7-2.9B-20260805 RWKV-7 «GOOSE» · постоянное рекуррентное языковое моделирование Нативная регистрация автокласса rwkv7 требует Transformers 5.15 или проверки источника...
Luciole-23B-Instruct-1.1-FP8 is a FP8-quantized version of Luciole-23B-Instruct-1.1 in the transformers / compressed-tensors (https://github.com/neuralmagic/compressed-tensors) format, quantized with LLM Compressor...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
Model Description Bias, Risks, and Limitations Recommendations Training Details Training Data Instruction template Training Procedure Evaluation Testing the...
> ⚠️ Предварительные результаты — 1 семя, дисперсия не контролируется. Каждое число ниже — это один тренировочный прогон,...
35-domain expert model built on Qwen3.5-35B-A3B (MoE, 256 experts, 3B active/token) with LoRA adapters and a cognitive layer...