Bắt đầu ngayBắt đầu miễn phí

Module 2 Quiz: Design and transformations — Question 4

You are building a new pipeline using Serverless for Apache Spark. The source data is highly structured, with well-defined columns like userid and purchaseamount. The primary task involves complex filtering and calculations on these columns. According to modern Spark best practices, which core API should you use to represent and manipulate this data?

Bài tập này là một phần của khóa học

Build Batch Data Pipelines on Google Cloud

Xem khóa học

Bài tập tương tác thực hành

Biến lý thuyết thành hành động với một trong các bài tập tương tác của chúng tôi

Bắt đầu bài tập