This Repository holds the model weights for the u-μP models trained at Aleph Alpha Research, in collaboration with Graphcore, for 72k steps (300B tokens). Please note, that the released checkpoints are not fully converged models and are intended for research use only. You can find all model weights at the following links: — umup-research-7b-bf16 — umup-research-7b-fp8 — sp-baseline-research-7b-bf16 — umup-research-3b-bf16 — umup-research-3b-fp8 — sp-baseline-research-3b-bf16 — umup-research-1b-bf16 — umup-research-1b-fp8 — sp-baseline-research-1b-bf16 The Maximal Update Parametrization (μP) aims to make the optimal hyperparameters (HPs) of a model-independent of its size, allowing them to be swept using a cheap proxy model rather than the We present a new scheme, u-μP, which improves upon μP by combining it with Unit Scaling, a method for designing models that makes them easy to train in low precision. The two techniques have a natural affinity: μP ensures that the scale of activations is independent of model size, and Unit Scaling ensures that activations, weights, and gradients begin training with a scale of one. This synthesis opens the door to a simpler…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: Aleph-Alpha
Теги: scaling
Лайков: 3 | Загрузок: 0
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.