Thêm dữ liệu vào collection
Đến lúc thêm các phim và chương trình TV trên Netflix vào collection của bạn! Bạn đã có sẵn danh sách ID của document và phần văn bản, lần lượt lưu trong ids và documents, được trích xuất từ netflix_titles.csv bằng đoạn mã sau:
ids = []
documents = []
with open('netflix_titles.csv') as csvfile:
reader = csv.DictReader(csvfile)
for i, row in enumerate(reader):
ids.append(row['show_id'])
text = f"Title: {row['title']} ({row['type']})\nDescription: {row['description']}\nCategories: {row['listed_in']}"
documents.append(text)
Ví dụ về thông tin sẽ được embedding, đây là document đầu tiên trong documents:
Title: Dick Johnson Is Dead (Movie)
Description: As her father nears the end of his life, filmmaker Kirsten Johnson stages his death in inventive and comical ways to help them both face the inevitable.
Categories: Documentaries
Tất cả hàm và gói cần thiết đã được import, và một persistent client đã được tạo rồi gán vào client.
Bài tập này là một phần của khóa học
Nhập môn Embeddings với OpenAI API
Hướng dẫn bài tập
- Tạo lại collection
netflix_titlescủa bạn. - Thêm các document và ID tương ứng vào collection.
- In ra số lượng document trong
collectionvà 10 mục đầu tiên.
Bài tập tương tác thực hành trực tiếp
Hãy thử làm bài tập này bằng cách hoàn thành đoạn mã mẫu này.
# Recreate the netflix_titles collection
collection = client.____(
name="netflix_titles",
embedding_function=OpenAIEmbeddingFunction(model_name="text-embedding-3-small", api_key="")
)
# Add the documents and IDs to the collection
____
# Print the collection size and first ten items
print(f"No. of documents: {____}")
print(f"First ten documents: {____}")