PLLuM is a family of large language models (LLMs) specialized in Polish and other Slavic/Baltic languages, with additional English data incorporated for broader generalization. Developed through an extensive collaboration with various data providers, PLLuM models are built on high-quality text corpora and refined through instruction tuning, preference learning, and advanced alignment techniques. These models are intended to generate contextually coherent text, offer assistance in various tasks (e.g., question answering, summarization), and serve as a foundation for specialized applications such as domain-specific intelligent assistants. — Extensive Data Collection We gathered large-scale, high-quality text data in Polish (around 150B tokens after cleaning and deduplication) and additional text in Slavic, Baltic, and English languages. Part of these tokens (28B) can be used in fully open-source models, including for commercial use (in compliance with relevant legal regulations). — Organic Instruction Dataset We curated the largest Polish collection of manually created “organic instructions” (~40k prompt-response pairs, including ~3.5k multi-turn dialogs). This human-authored…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: CYFRAGOVPL
Теги: llama, conversational, pl, text-generation-inference, endpoints_compatible
Лайков: 4 | Загрузок: 106
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.