nicholasKluge/Aira-2-portuguese-124M - Каталог нейросетей
Генерация текста

nicholasKluge/Aira-2-portuguese-124M

Добавлено:
nicholasKluge/Aira-2-portuguese-124M

Aira-2 is the second version of the Aira instruction-tuned series. Aira-2-portuguese-124M is an instruction-tuned model based on GPT-2. The model was trained with a dataset composed of prompt, completions generated synthetically by prompting already-tuned models (ChatGPT, Llama, Open-Assistant, etc). — Size: 124,441,344 parameters — Dataset: Instruct-Aira Dataset — Language: Portuguese — Number of Epochs: 5 — Batch size: 24 — Optimizer: torch.optim.AdamW (warmupsteps = 1e2, learningrate = 5e-4, epsilon = 1e-8) — GPU: 1 NVIDIA A100-SXM4-40GB — Emissions: 0.35 KgCO2 (Singapore) — Total Energy Consumption: 0.73 kWh This repository has the source code used to train this model. Three special tokens are used to mark the user side of the interaction and the model’s response: O que é um modelo de linguagem?Um modelo de linguagem é uma distribuição de probabilidade sobre um vocabulário. — Hallucinations: This model can produce content that can be mistaken for truth but is, in fact, misleading or entirely false, i.e., hallucination. — Biases and Toxicity: This model inherits the social and historical stereotypes from the data used to train it. Given these biases, the model can produce toxic…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: nicholasKluge
Теги: gpt2, alignment, instruction tuned, text generation, conversation, assistant, pt, co2_eq_emissions
Лайков: 3  |  Загрузок: 137

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.