Bắt đầu ngayBắt đầu miễn phí

Tạo speech embedding

Đến lúc mã hóa một mảng âm thanh thành speaker embedding! Speaker embedding chứa thông tin về cách cá nhân hóa audio được tạo theo một người nói cụ thể, và là thành phần thiết yếu để tạo ra audio được fine-tune.

Mô hình pretrained spkrec-xvect-voxceleb (speaker_model) và bộ dữ liệu VCTK (dataset) đã được nạp sẵn cho bạn.

Bài tập này là một phần của khóa học

Mô hình đa phương thức với Hugging Face

Xem khóa học

Hướng dẫn bài tập

  • Hoàn thiện định nghĩa hàm create_speaker_embedding() bằng cách tính embedding thô từ waveform bằng speaker_model.
  • Trích xuất mảng âm thanh từ điểm dữ liệu tại chỉ số 10 của dataset.
  • Tính một speaker embedding từ mảng âm thanh bằng hàm create_speaker_embedding().

Bài tập tương tác thực hành trực tiếp

Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.

def create_speaker_embedding(waveform):
    with torch.no_grad():
        # Calculate the raw embedding from the speaker_model
        speaker_embeddings = ____.____(torch.tensor(____))
        
        speaker_embeddings = torch.nn.functional.normalize(speaker_embeddings, dim=2)
        speaker_embeddings = speaker_embeddings.squeeze().cpu().numpy()
    return speaker_embeddings

# Extract the audio array from the dataset
audio_array = dataset[10]["____"]["____"]

# Calculate the speaker_embedding from the datapoint
speaker_embedding = ____(____)
print(speaker_embedding.shape)
Chỉnh sửa và Chạy Mã