llm-jp/llm-jp-3.1-8x13b-instruct4 - Каталог нейросетей
Генерация текста

llm-jp/llm-jp-3.1-8x13b-instruct4

Добавлено:
llm-jp/llm-jp-3.1-8x13b-instruct4

LLM-jp-3.1 is a series of large language models developed by the Research and Development Center for Large Language Models at the National Institute of Informatics. Building upon the LLM-jp-3 series, the LLM-jp-3.1 models incorporate mid-training (instruction pre-training), which significantly enhances their instruction-following capabilities compared to the original LLM-jp-3 models. This repository provides the llm-jp-3.1-8x13b-instruct4 model. For an overview of the LLM-jp-3.1 models across different parameter sizes, please refer to: — LLM-jp-3.1 Pre-trained Models — LLM-jp-3.1 Fine-tuned Models. For more details on the training procedures and evaluation results, please refer to this blog post (in Japanese). — torch>=2.3.0 — transformers>=4.40.1 — tokenizers>=0.19.1 — accelerate>=0.29.3 — flash-attn>=2.5.8 — Model type: Transformer-based Language Model — Architectures: The tokenizer of this model is based on huggingface/tokenizers Unigram byte-fallback model. The vocabulary entries were converted from llm-jp-tokenizer v3.0. Please refer to README.md of llm-jp-tokenizer for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our…

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: llm-jp
Теги: mixtral, conversational, en, ja, text-generation-inference
Лайков: 4  |  Загрузок: 125

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.