ChronoGPT is a series high-performance chronologically consistent large language models (LLMs) designed to eliminate lookahead bias and training leakage while maintaining good language understanding in time-sensitive applications. The model is pretrained on diverse, high-quality, open-source, and timestamped text to maintain chronological consistency. All models in the series achieve HellaSwag benchmark scores that surpass those of the GPT-2 124M model. This approach preserves the integrity of historical analysis and enables more reliable economic and financial modeling. — Developed by: Songrun He, Linying Lv, Asaf Manela, Jimmy Wu — Model type: Transformer-based autoregressive decoder (Modified modded-NanoGPT architecture) — Language(s) (NLP): English — License: MIT License ChronoGPT has the following features: — Type: Causal Language Models — Training Stage: Pretraining — Number of Parameters: ~1,552 Million — Encoder & Decoder Partitioning: 26 encoder and 26 decoder layers — Tokenizer: GPT2Tokenizer from HuggingFace — Context Length: 1,792
Модальности:
Генерация текста
Задача: Генерация текста
Автор: manelalab
Теги: ChronoGPT, chronologically consistent, modded-nanogpt, hellaswag, en
Лайков: 4 | Загрузок: 349
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.