시작하기무료로 시작하기

Quiz 2 - Question 1

Imagine you train a byte pair encoding (BPE) tokenizer on English and Amharic texts. This means that they share a single vocabulary consisting of English and Amharic subword tokens. You apply this tokenizer to the following Amharic sentence:

ስለተዋወቅን ደስ ብሎኛል

The tokenizer splits this sentence into 14 tokens. When you tokenize its English translation, “Nice to meet you”, it splits it into 7 tokens.

Which explanation is most plausible given how BPE learns merges and determines its subword token vocabulary.

이 연습은 강의의 일부입니다

Google DeepMind: Represent Your Language Data

강의 보기

실습형 인터랙티브 연습문제

이론을 실습으로 바꾸는 인터랙티브 연습 중 하나를 만나보세요

연습 시작