A small, fast content-moderation classifier fine-tuned from Qwen2.5-3B-Instruct for Discord-style chat. It is the L1 fast-triage tier of the ModDog pipeline: it returns a structured JSON verdict (flag / category / confidence / reason) and is designed to be honestly uncertain on hard cases so they escalate to a larger model rather than being confidently mis-judged. This release (2026-06-23) is the model running in ModDog production. It replaces the previous upload; weights here are tensor-identical to the production checkpoint. Fast first-pass moderation triage on chat-style messages, for the judgment-call categories: toxicity, harassment, hatespeech, sexualcontent, selfharm, violence (plus benign). The verdict is meant to feed a graduated action ladder where low-confidence flags route to human review**, not automatic penalties. In the ModDog pipeline this model sits behind a deterministic rule layer («L0»), and several duties are deliberately delegated there — this model is neither trained nor evaluated for them: — Spam — invite links, scam phrases, mass mentions. Spam examples were excluded from this model’s training mix; the spam label exists in the verdict schema for pipeline…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: Hagrun
Теги: gguf, qwen2, content-moderation, moderation, safety, discord, qwen2.5, conversational
Лайков: 4 | Загрузок: 60
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.