HiveTraceGuard-Pro is a compact Russian-first guardrail built on Qwen3-0.6B for fast input and output classification. Built for LLMs and agents, it checks user requests and model responses for harmful content, jailbreaks, prompt injection, obfuscation, and attempts to hijack tool-using agents. The model is stateless and returns exactly one token: safe or unsafe. Model Requests Responses AEGIS 2.0 ToxicChat XSTest XSafetyEN OpenAIModeration AEGIS 2.0 BeaverTails HarmBench HiveTraceGuard-Pro (0.6B) 0.817 0.588 0.754 0.590 0.803 0.797 0.839 0.814 Shieldstral-1.0-3B 0.808 0.732 0.922 0.595 0.794 0.766 0.828 0.854 YuFeng-XGuard-Reason-0.6B 0.847 0.620 0.920 0.469 0.787 0.789 0.828 0.858 Qwen3Guard-Gen-0.6B 0.788 0.692 0.861 0.580 0.715 0.819 0.845 0.856 Llama-Guard-3-1B 0.733 0.385 0.837 0.368 0.766 0.635 0.652 0.794 OR-BenchToxic Multi
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: hivetrace
Теги: qwen3, guardrail, safety, moderation, content-moderation, prompt-injection, jailbreak, russian
Лайков: 4 | Загрузок: 355
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.