微調文字轉語音模型
你將使用 VCTK Corpus 來微調一個文字轉語音(text-to-speech)模型,以重現不同地區口音。這個資料集包含約 44 小時的語音資料,說話者為帶有各種英語口音的母語者。
dataset 已經載入並完成前處理,SpeechT5ForTextToSpeech 模組與 Seq2SeqTrainingArguments、Seq2SeqTrainer 模組也都已載入。資料整理器(data_collator)已預先定義。
請不要在 trainer 設定上呼叫 .train() 方法,因為在此環境執行會導致逾時。
本練習屬於課程
使用 Hugging Face 的多模態模型
練習說明
- 使用
SpeechT5ForTextToSpeech載入microsoft/speecht5_tts的預訓練模型。 - 建立
Seq2SeqTrainingArguments實例,設定:gradient_accumulation_steps為8、learning_rate為0.00001、warmup_steps為500、max_steps為4000。 - 使用新的訓練參數,並搭配提供的
model、資料與processor,來設定 trainer。
動手互動練習
試著完成這個範例程式碼,體驗一下這個練習。
# Load the text-to-speech pretrained model
model = ____.____(____)
# Configure the required training arguments
training_args = ____(output_dir="speecht5_finetuned_vctk_test",
gradient_accumulation_steps=____, learning_rate=____, warmup_steps=____, max_steps=4000, label_names=["labels"],
push_to_hub=False)
# Configure the trainer
trainer = ____(args=training_args, model=model, data_collator=data_collator,
train_dataset=dataset["train"], eval_dataset=dataset["test"], tokenizer=processor)