marin-community/delphi-1e23-25Bparams-628Btokens - Каталог нейросетей
Генерация текста

marin-community/delphi-1e23-25Bparams-628Btokens

Добавлено:
marin-community/delphi-1e23-25Bparams-628Btokens

A 25B-parameter base model from the Delphi scaling suite. Trained at 1 × 10²³ FLOPs on 628B tokens with the Delphi recipe. Delphi is the Marin team’s first open scaling suite, inspired by Pythia. It has three parts: — a scaling recipe that maps compute budgets to model configurations, — a scaling suite of models trained from that recipe at IsoFLOP budgets from 3 × 10¹⁸ to 1 × 10²³ FLOPs, and — a scaling law which uses the smaller Delphi models to predict the larger ones. A pre-registered forecast from that scaling law predicted the final loss of the largest Delphi run (1 × 10²³ FLOPs, 25 B parameters, 600 B tokens) within 0.2%, using 300× less compute than the training run itself. The same process forecasts downstream benchmarks — MMLU, HumanEval, and GSM8K — via a two-step regression combining compute and observational scaling laws. See «Scaling Laws That Extrapolate 300× Past the Fit» for the recipe, fit, and downstream-eval projections. The full set of Delphi checkpoints — IsoFLOP grid points, held-out optima at 1e21/1e22/1e23 with multiple random seeds, and training intermediates — lives on marin-community on the Hub. AdamH, Adam with Hyperball, constrains every projection…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: marin-community
Теги: qwen3, marin, delphi, scaling-laws, pretrained, research-only, en, text-generation-inference
Лайков: 4  |  Загрузок: 237

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.