ysdede/Phi-4-mm-inst-asr-turkish-3 - Каталог нейросетей
Генерация текста

ysdede/Phi-4-mm-inst-asr-turkish-3

Добавлено:
ysdede/Phi-4-mm-inst-asr-turkish-3

This model is a fine-tuned version of microsoft/Phi-4-multimodal-instruct on a 1300-hour Turkish audio dataset. The model was initially fine-tuned using the original ASR prompt: «Transcribe the audio clip into text.» This prompt is language agnostic—as described in the model paper: > The ASR prompt for Phi-4-Multimodal is “Transcribe the audio clip into text.”, which is language agnostic. We notice that the model can learn to recognize in the target language perfectly without providing language information, while Qwen2-audio and Gemini-2.0-Flash require the language information in the prompt to obtain the optimal ASR performance. However, we found that using a language-defining prompt, such as: «Transcribe the Turkish audio.» leads to better performance. See: ysdede/Phi-4-mm-inst-asr-turkish When benchmarked with the original ASR prompt «Transcribe the audio clip into text.», the evaluation results were as follows: Load generationconfig and processor` from the base model as a quick fix to use the default generation settings. — Transformers 4.46.1 — Pytorch 2.5.1+cu124 — Datasets 3.3.2 — Tokenizers 0.20.3

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: ysdede
Теги: tensorboard, phi4mm, generated_from_trainer, conversational, custom_code
Лайков: 3  |  Загрузок: 31

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.