SortByKey と Collect
キーに基づいてペア RDD を並び替える(この章の後半で登場する word count など)ことはよくあります。この演習では、前の演習で作成したペア RDD Rdd_Reduced をキーで降順にソートし、最終的な出力を表示します。
なお、作業スペースには SparkContext sc と Rdd_Reduced がすでに用意されています。
この演習はコースの一部です
PySparkで学ぶBig Data入門
演習の手順
Rdd_Reducedをキーで降順にソートします。- 中身を collect して、反復処理で出力を表示します。
実践的なインタラクティブ演習
このサンプルコードを完成させて、この演習に挑戦してみましょう。
# Sort the reduced RDD with the key by descending order
Rdd_Reduced_Sort = Rdd_Reduced.____(ascending=False)
# Iterate over the result and retrieve all the elements of the RDD
for num in Rdd_Reduced_Sort.____():
print("Key {} has {} Counts".format(____, num[1]))