การประเมินโมเดลสร้างข้อความแบบ pre-trained
ทีม PyBooks ได้นำโมเดล GPT-2 แบบ pre-trained ที่คุณทดลองใช้มาสร้างข้อความจาก prompt ที่กำหนด ขณะนี้พวกเขาต้องการประเมินคุณภาพของข้อความที่สร้างขึ้น โดยมอบหมายให้คุณประเมินข้อความที่สร้างเทียบกับข้อความอ้างอิง
โหลด BLEUScore และ ROUGEScore ไว้ให้แล้ว
แบบฝึกหัดนี้เป็นส่วนหนึ่งของหลักสูตร
Deep Learning สำหรับข้อความด้วย PyTorch
คำแนะนำการฝึกหัด
- เริ่มต้นด้วยการ initialize เมตริกทั้งสอง (BLEU และ ROUGE) จาก
torchmetrics.text - ใช้เมตริกที่ initialize แล้วเพื่อคำนวณคะแนนระหว่างข้อความที่สร้างขึ้นกับข้อความอ้างอิง
- แสดงคะแนน BLEU และ ROUGE ที่คำนวณได้
แบบฝึกหัดเชิงโต้ตอบแบบลงมือทำ
ลองทำแบบฝึกหัดนี้โดยเติมโค้ดตัวอย่างนี้ให้สมบูรณ์
reference_text = "Once upon a time, there was a little girl who lived in a village near the forest."
generated_text = "Once upon a time, the world was a place of great beauty and great danger. The world of the gods was the place where the great gods were born, and where they were to live."
# Initialize BLEU and ROUGE scorers
bleu = ____()
rouge = ____()
# Calculate the BLEU and ROUGE scores
bleu_score = bleu([____], [[reference_text]])
rouge_score = rouge([generated_text], [[____]])
# Print the BLEU and ROUGE scores
print("BLEU Score:", bleu_score.____())
print("ROUGE Score:", rouge_score)