Bắt đầu ngayBắt đầu miễn phí

Quiz 2 - Question 1

Imagine you train a byte pair encoding (BPE) tokenizer on English and Amharic texts. This means that they share a single vocabulary consisting of English and Amharic subword tokens. You apply this tokenizer to the following Amharic sentence:

ስለተዋወቅን ደስ ብሎኛል

The tokenizer splits this sentence into 14 tokens. When you tokenize its English translation, “Nice to meet you”, it splits it into 7 tokens.

Which explanation is most plausible given how BPE learns merges and determines its subword token vocabulary.

Bài tập này là một phần của khóa học

Google DeepMind: Represent Your Language Data

Xem khóa học

Bài tập tương tác thực hành

Biến lý thuyết thành hành động với một trong các bài tập tương tác của chúng tôi

Bắt đầu bài tập