Sinh bản dịch
Bây giờ bạn sẽ sinh bản dịch tiếng Pháp bằng mô hình suy luận được huấn luyện với Teacher Forcing.
Mô hình này (nmt_tf) đã được huấn luyện 50 epoch trên 100.000 câu và đạt khoảng 98% accuracy trên tập kiểm định hơn 35.000 mẫu. Bài tập này có thể khởi tạo lâu hơn vì cần tải mô hình đã huấn luyện. Bạn đã có sẵn hàm sents2seqs(). Ngoài ra còn có hai hàm mới:
word2onehot(tokenizer, word, vocab_size) với các tham số:
- tokenizer - Đối tượng
Tokenizercủa Keras - word - Chuỗi biểu diễn một từ trong từ vựng (ví dụ:
'apple') - vocab_size - Kích thước từ vựng
probs2word(probs, tok) với các tham số:
- probs - Đầu ra từ mô hình có dạng
[1,<French Vocab Size>] - tok - Đối tượng
Tokenizercủa Keras
Bạn có thể xem nhanh mã nguồn của các hàm này bằng cách gõ print(inspect.getsource(word2onehot)) và print(inspect.getsource(probs2word)) trong console.
Bài tập này là một phần của khóa học
Machine Translation với Keras
Hướng dẫn bài tập
- Dự đoán trạng thái khởi tạo của decoder (
de_s_t) bằng encoder. - Dự đoán đầu ra và trạng thái mới từ decoder, sử dụng dự đoán (đầu ra) trước đó và trạng thái trước đó làm đầu vào. Nhớ đệ quy cập nhật trạng thái mới.
- Lấy chuỗi từ (word string) từ đầu ra xác suất bằng hàm
probs2word(). - Chuyển chuỗi từ thành chuỗi one-hot bằng hàm
word2onehot().
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
en_sent = ['the united states is sometimes chilly during december , but it is sometimes freezing in june .']
print('English: {}'.format(en_sent))
en_seq = sents2seqs('source', en_sent, onehot=True, reverse=True)
# Predict the initial decoder state with the encoder
de_s_t = ____.predict(____)
de_seq = word2onehot(fr_tok, 'sos', fr_vocab)
fr_sent = ''
for i in range(fr_len):
# Predict from the decoder and recursively assign the new state to de_s_t
de_prob, ____ = ____.predict([____,____])
# Get the word from the probability output using probs2word
de_w = probs2word(____, fr_tok)
# Convert the word to a onehot sequence using word2onehot
de_seq = word2onehot(fr_tok, ____, fr_vocab)
if de_w == 'eos': break
fr_sent += de_w + ' '
print("French (Ours): {}".format(fr_sent))
print("French (Google Translate): les etats-unis sont parfois froids en décembre, mais parfois gelés en juin")