This repository provides large language models developed by LLM-jp, a collaborative project launched in Japan. — torch>=2.3.0 — transformers>=4.40.1 — tokenizers>=0.19.1 — accelerate>=0.29.3 — flash-attn>=2.5.8 — Model type: Transformer-based Language Model — Total seen tokens: 256B — Pre-training: — Hardware: 128 A100 40GB GPUs (mdx cluster) — Software: Megatron-LM — Instruction tuning: — Hardware: 8 A100 40GB GPUs (mdx cluster) — Software: TRL and DeepSpeed The tokenizer of this model is based on huggingface/tokenizers Unigram byte-fallback model. The vocabulary entries were converted from llm-jp-tokenizer v2.2 (100k: code20Ken40Kja60K.ver2.2). Please refer to README.md of llm-ja-tokenizer for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary). — Model: Hugging Face Fast Tokenizer using Unigram byte-fallback model — Training algorithm: Marging Code/English/Japanese vocabularies constructed with SentencePiece Unigram byte-fallback and reestimating scores with the EM-algorithm. — Training data: A subset of the datasets for model pre-training — Vocabulary size: 96,867 (mixed vocabulary of Japanese, English, and…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: llm-jp
Теги: llama, conversational, en, ja, text-generation-inference
Лайков: 3 | Загрузок: 133
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.