Đánh giá chuỗi
Đến lúc bạn thật sự đánh giá đầu ra cuối cùng bằng cách so sánh với câu trả lời do chuyên gia nội dung viết. Bạn sẽ dùng lớp LangChainStringEvaluator của LangSmith để thực hiện bài đánh giá so sánh chuỗi này.
Một prompt_template để đánh giá chuỗi đã được viết sẵn cho bạn như sau:
You are an expert professor specialized in grading students' answers to questions.
You are grading the following question:{query}
Here is the real answer:{answer}
You are grading the following predicted answer:{result}
Respond with CORRECT or INCORRECT:
Grade:
Đầu ra từ RAG chain được lưu trong predicted_answer và phản hồi của chuyên gia được lưu trong ref_answer.
Tất cả các lớp cần thiết đã được import sẵn cho bạn.
Bài tập này là một phần của khóa học
Retrieval Augmented Generation (RAG) với LangChain
Hướng dẫn bài tập
- Tạo bộ đánh giá chuỗi QA của LangSmith bằng
eval_llmvàprompt_templateđã cung cấp. - Đánh giá đầu ra RAG,
predicted_answer, bằng cách so sánh với phản hồi của chuyên gia choquery, được lưu trongref_answer.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Create the QA string evaluator
qa_evaluator = ____(
"____",
config={
"llm": ____,
"prompt": ____
}
)
query = "How does RAG improve question answering with LLMs?"
# Evaluate the RAG output by evaluating strings
score = qa_evaluator.evaluator.____(
prediction=____,
reference=____,
input=____
)
print(f"Score: {score}")