Bắt đầu ngayBắt đầu miễn phí

Sinh bản dịch

Bây giờ bạn sẽ sinh bản dịch tiếng Pháp bằng mô hình suy luận được huấn luyện với Teacher Forcing.

Mô hình này (nmt_tf) đã được huấn luyện 50 epoch trên 100.000 câu và đạt khoảng 98% accuracy trên tập kiểm định hơn 35.000 mẫu. Bài tập này có thể khởi tạo lâu hơn vì cần tải mô hình đã huấn luyện. Bạn đã có sẵn hàm sents2seqs(). Ngoài ra còn có hai hàm mới:

word2onehot(tokenizer, word, vocab_size) với các tham số:

  • tokenizer - Đối tượng Tokenizer của Keras
  • word - Chuỗi biểu diễn một từ trong từ vựng (ví dụ: 'apple')
  • vocab_size - Kích thước từ vựng

probs2word(probs, tok) với các tham số:

  • probs - Đầu ra từ mô hình có dạng [1,<French Vocab Size>]
  • tok - Đối tượng Tokenizer của Keras

Bạn có thể xem nhanh mã nguồn của các hàm này bằng cách gõ print(inspect.getsource(word2onehot))print(inspect.getsource(probs2word)) trong console.

Bài tập này là một phần của khóa học

Machine Translation với Keras

Xem khóa học

Hướng dẫn bài tập

  • Dự đoán trạng thái khởi tạo của decoder (de_s_t) bằng encoder.
  • Dự đoán đầu ra và trạng thái mới từ decoder, sử dụng dự đoán (đầu ra) trước đó và trạng thái trước đó làm đầu vào. Nhớ đệ quy cập nhật trạng thái mới.
  • Lấy chuỗi từ (word string) từ đầu ra xác suất bằng hàm probs2word().
  • Chuyển chuỗi từ thành chuỗi one-hot bằng hàm word2onehot().

Bài tập tương tác thực hành trực tiếp

Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.

en_sent = ['the united states is sometimes chilly during december , but it is sometimes freezing in june .']
print('English: {}'.format(en_sent))
en_seq = sents2seqs('source', en_sent, onehot=True, reverse=True)
# Predict the initial decoder state with the encoder
de_s_t = ____.predict(____)
de_seq = word2onehot(fr_tok, 'sos', fr_vocab)
fr_sent = ''
for i in range(fr_len):    
  # Predict from the decoder and recursively assign the new state to de_s_t
  de_prob, ____ = ____.predict([____,____])
  # Get the word from the probability output using probs2word
  de_w = probs2word(____, fr_tok)
  # Convert the word to a onehot sequence using word2onehot
  de_seq = word2onehot(fr_tok, ____, fr_vocab)
  if de_w == 'eos': break
  fr_sent += de_w + ' '
print("French (Ours): {}".format(fr_sent))
print("French (Google Translate): les etats-unis sont parfois froids en décembre, mais parfois gelés en juin")
Chỉnh sửa và Chạy Mã