kl3m-002-170m is a (very) small language model (SLM) trained on clean, legally-permissible data. Originally developed by 273 Ventures and donated to the ALEA Institute, kl3m-002-170m was the first LLM to obtain the Fairly Trained L-Certification for its ethical training data and practices. The model is designed for legal, regulatory, and financial workflows, with a focus on low toxicity and high efficiency. Given its small size and lack of instruction-aligned training data, kl3m-002-170m is best suited for use either in SLM fine-tuning or as part of training larger models without using unethical data or models. — Architecture: GPT-NeoX (i.e., ~GPT-3 architecture) — Size: 170 million parameters — Hidden Size: 1024 — Layers: 16 — Attention Heads: 16 — Key-Value Heads: 8 — Intermediate Size: 1024 — Max Sequence Length: 4,096 tokens (true size, no sliding window) — Tokenizer: kl3m-001-32k BPE tokenizer (32,768 vocabulary size with unorthodox whitespace handling) — Language(s): Primarily English — Training Objective: Next token prediction — Developed by: Originally by 273 Ventures LLC, d
Модальности:
Генерация текста
Области применения:
Юриспруденция
Задача: Генерация текста
Автор: alea-institute
Теги: gpt_neox, kl3m, kl3m-002, legal, financial, enterprise, slm, gpt-neox
Лайков: 3 | Загрузок: 603
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.