ทำความคุ้นเคยกับข้อมูลข้อความ
ในแบบฝึกหัดนี้ จะได้ลองทำงานกับข้อมูลข้อความโดยวิเคราะห์คำพูดของ Sheldon Cooper จากซีรีส์ The Big Bang Theory ซึ่งจะช่วยให้เข้าใจว่าการจัดการกับข้อมูลข้อความในชีวิตจริงนั้นเป็นอย่างไร
จะใช้ dictionary comprehension เพื่อสร้าง dictionary ที่แมปคำไปยังดัชนีและในทิศทางกลับกัน ที่เลือกใช้ dictionary แทน pandas.DataFrame เนื่องจากเข้าใจได้ง่ายกว่าและไม่เพิ่มความซับซ้อนที่ไม่จำเป็น
ข้อมูลอยู่ใน sheldon_quotes โดยสองประโยคแรกถูกแสดงผลไว้ให้แล้ว
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Recurrent Neural Networks (RNNs) สำหรับ Language Modeling ด้วย Keras
คำแนะนำการฝึกหัด
- ใช้
joinเพื่อรวมประโยคทั้งหมดเข้าเป็นตัวแปรเดียว จากนั้นดึงคำทั้งหมดออกมาและเก็บไว้ในลิสต์all_words - ลบคำที่ซ้ำกันโดยใช้
list(set())กับลิสต์คำ แล้วเก็บผลลัพธ์ไว้ในunique_words - สร้าง dictionary โดยใช้ดัชนีเป็น key และคำเป็น value ด้วย dictionary comprehension
- สร้าง dictionary โดยใช้คำเป็น key และดัชนีเป็น value ด้วย dictionary comprehension
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
# Transform the list of sentences into a list of words
all_words = ' '.____(sheldon_quotes).split(' ')
# Get number of unique words
unique_words = list(set(all_words))
# Dictionary of indexes as keys and words as values
index_to_word = {____ for i, wd in enumerate(sorted(unique_words))}
print(index_to_word)
# Dictionary of words as keys and indexes as values
word_to_index = {wd:i for ____ in enumerate(sorted(unique_words))}
print(word_to_index)