Yuren 13B is a large-scale language model that has been continuously trained based on Llama 2 13B. Focused on the field of information synthesis and built upon the data-centric work of Pleisto, this model achieves state-of-the-art levels in data synthesis scenarios such as information extraction in multiple languages, natural language generation of SQL, and structured data output, all with an equivalent parameter count. 羽人 13B 是在 Llama 2 13B 基础上进行持续训练的大语言模型,聚焦于信息合成领域并建立在 Pleisto 以数据为中心的工作上。该模型在以中英文为主的多种语言的信息抽取、自然语言生成 SQL、结构化数据输出等数据合成类场景下实现了同等参数量下的 SOTA 水平。 In the original Llama vocabulary, only a few hundred Chinese characters were included, and the remaining Chinese characters had to be generated by concatenating multiple Unicode bytes. This issue not only obviously affects the Chinese inference performance (generation speed), but also significantly creates a performance bottleneck in Chinese semantic understanding. We conducted a series of comparative experiments on different vocabulary expansion approaches and found the following: * Compared to the prevailing strategy of adding a large number of commonly used Chinese character words to the vocabulary, simply adding Chinese…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: pleisto
Теги: llama, llama2, zh, en, model-index, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 113
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.