Quiz 2 - Question 1
Imagine you train a byte pair encoding (BPE) tokenizer on English and Amharic texts. This means that they share a single vocabulary consisting of English and Amharic subword tokens. You apply this tokenizer to the following Amharic sentence:
ስለተዋወቅን ደስ ብሎኛል
The tokenizer splits this sentence into 14 tokens. When you tokenize its English translation, “Nice to meet you”, it splits it into 7 tokens.
Which explanation is most plausible given how BPE learns merges and determines its subword token vocabulary.
Bài tập này là một phần của khóa học
Google DeepMind: Represent Your Language Data
Bài tập tương tác thực hành
Biến lý thuyết thành hành động với một trong các bài tập tương tác của chúng tôi
Bắt đầu bài tập