Sumo-T9-7B-v0.1
Tensorplex Labs) is proud to announce that its latest top-performing model on Bittensor Subnet 9, Sumo-T9-7B, has outperformed...
Tensorplex Labs) is proud to announce that its latest top-performing model on Bittensor Subnet 9, Sumo-T9-7B, has outperformed...
TR-HASH MoE 200M is a 201.2M-parameter decoder-only base language model with grouped-query attention, a shared SwiGLU path, and...
Veyra2-Mango-15M-Base is a 15.7M-parameter Llama-like causal language model trained from scratch on approximately 30B tokens. It is a...
ObsidianSmall-Base is an 8.7M-parameter English base language model pretrained from scratch on 3,999,989,760 tokens. It is the first...
Escarda-86M-Base is a ~86M-parameter, from-scratch decoder-only language model — the base sibling of Quazim0t0/Escarda-86M (the chat-tuned model). It...
Argon-0.5B 一个自研的基线模型,复刻 DeepSeek-V4 模型训练典型优化器,并加入 Engram 模块。 本仓库计划上传模型权重、切分后的训练数据、训练代码、tokenizer 资产和完整配置,使 Argon-0.5B 成为一个可审计、可复现、可继续训练的研究型预训练样例。 这个项目的初衷是复刻 DeepSeek 技术栈中的关键训练流程,并尝试在 500M 参数规模上实现一个完整的预训练闭环。 — 复刻 、数据...
— Исследования — Бенчмаркинг — Эксперименты по предварительной подготовке — Эксперименты по тонкой настройке — Разработка модели малого...
This is a GGUF conversion of common-pile/comma-v0.1-2t for use with llama.cpp and Ollama. Original Model: Comma v0.1-2T Architecture:...
We present Tri-7B-Base, a foundation language model that serves as the pre-trained base for our Tri-7B model family....
Демо: https://huggingface.co/spaces/Banaxi-Tech/BananaMind-2-AI-Detect-Demo BananaMind 2 AI Detect — это детектор текста «человек против ИИ», созданный путем полной настройки Qwen/Qwen3.5-0.8B-Base...