clairaudience's Introduction

Zero-shot Domain-sensitive Speech Recognition with Prompt-conditioning Fine-tuning

Feng-Ting Liao, Yung-Chieh Chan, Yi-Chang Chen, Chan-Jan Hsu, Da-shan Shiu

In this work, we propose a method to create domain-sensitive speech recognition models that utilize textual domain information by conditioning its generation on a given text prompt. This is accomplished by fine-tuning a pre-trained, end-to-end model (Whisper) to learn from demonstrations with prompt examples. We show that this ability can be generalized to different domains and even various prompt contexts, with our model gaining a Word Error Rate (WER) reduction of up to 33% on unseen datasets from various domains, such as medical conversation, air traffic control communication, and financial meetings. Considering the limited availability of audio-transcript pair data, we further extend our method to text-only fine-tuning to achieve domain sensitivity as well as domain adaptation. We demonstrate that our text-only fine-tuned model can also attend to various prompt contexts, with the model reaching the most WER reduction of 29% on the medical conversation dataset.

Installation

Please first install openai's whisper repo and also the packages in requirements.txt

Training

To run the training example, ensure that Gigaspeech medium is downladed to data/hf_dd_data/gigaspeech/m and execute

python ./clairaudience/main.py ./configs/cfg_gigaspeech_ft_base.json

Evaluation

To run the training example, ensure that Gigaspeech medium is downladed to data/hf_dd_data/gigaspeech/m and execute

python ./clairaudience/main.py ./configs/cfg_gigaspeech_evaluation.json

Model Weight

See https://huggingface.co/MediaTek-Research/Clairaudience

clairaudience's People

Contributors

Stargazers

Watchers

clairaudience's Issues

configs for text-only training

I see the cfg_gigaspeech_ft_base.json in the [config] folder. I think that is the setting for audio-text training.
Can you provide the setting for text-only training? I would appreciate it very much.

training data prepare

Hi, I wnat to train model use GigaSpeech now, the Guidance doc said ensure that Gigaspeech medium is downladed to data/hf_dd_data/gigaspeech/m，i don’t know which format is available，origin long wav ? or short wav slice according to segments (timestamp json file)?
moreover，what means of dataset_name in func setup_dataset (data_process.py);
line 94: elif dataset_name == gigaspeech:
line 101: elif dataset_name == gigaspeech_extracted:
line 115: elif dataset_name == gigaspeech:

thx！

Recommend Projects

mtkresearch / clairaudience Goto Github PK

clairaudience's Introduction

Zero-shot Domain-sensitive Speech Recognition with Prompt-conditioning Fine-tuning

Installation

Training

Evaluation

Model Weight

clairaudience's People

Contributors

Stargazers

Watchers

Forkers

clairaudience's Issues

configs for text-only training

training data prepare

Recommend Projects

React

Vue.js

Typescript

TensorFlow

Django

Laravel

D3

Recommend Topics

javascript

web

server

Machine learning

Visualization

Game

Recommend Org

Facebook

Microsoft

Google

Alibaba

D3

Tencent