Wikipedia clustern, Teil II

Jetzt ist es Zeit, deine Pipeline aus der vorherigen Übung einzusetzen! Du erhältst ein Array articles mit tf-idf-Wortfrequenzen einiger populärer Wikipedia-Artikel und eine Liste titles mit deren Titeln. Verwende deine Pipeline, um die Wikipedia-Artikel zu clustern.

Eine Lösung der vorherigen Übung wurde bereits für dich geladen, sodass eine Pipeline pipeline, die TruncatedSVD mit KMeans verknüpft, verfügbar ist.

Diese Übung ist Teil des Kurses

<Kurs>Unsupervised Learning in Python</Kurs>

Übungsanweisungen

Importiere pandas als pd.
Passe die Pipeline an das Wortfrequenz-Array articles an.
Sage die Cluster-Labels voraus.
Richte die Cluster-Labels an der Liste titles der Artikeltitel aus, indem du ein DataFrame df mit labels und titles als Spalten erstellst. Das wurde bereits für dich erledigt.
Verwende die Methode .sort_values() von df, um das DataFrame nach der Spalte 'label' zu sortieren, und gib das Ergebnis aus.
Drücke auf Antwort senden und nimm dir einen Moment, um dein großartiges Clustering der Wikipedia-Seiten anzuschauen!

Interaktive praktische Übung

Versuche dich an dieser Übung, indem du diesen Beispielcode vervollständigst.

# Import pandas
____

# Fit the pipeline to articles
____

# Calculate the cluster labels: labels
labels = ____

# Create a DataFrame aligning labels and titles: df
df = pd.DataFrame({'label': labels, 'article': titles})

# Display df sorted by cluster label
print(____)

Code bearbeiten und ausführen

Diese Übung ist Teil des Kurses

<Kurs>Unsupervised Learning in Python</Kurs>

Mittlere SchwierigkeitSchwierigkeitsgrad

4.8+

Kurs kostenlos starten

Learn how to discover the underlying groups (or "clusters") in a dataset. By the end of this chapter, you'll be clustering companies using their stock market prices, and distinguishing different species by clustering their measurements.

Exercise 1: Unsupervised Learning Exercise 2: How many clusters?Exercise 3: Clustering 2D points Exercise 4: Inspect your clustering Exercise 5: Evaluating a clustering Exercise 6: How many clusters of grain?Exercise 7: Evaluating the grain clustering Exercise 8: Transforming features for better clusterings Exercise 9: Scaling fish data for clustering Exercise 10: Clustering the fish data Exercise 11: Clustering stocks using KMeans Exercise 12: Which stocks move together?

In this chapter, you'll learn about two unsupervised learning techniques for data visualization, hierarchical clustering and t-SNE. Hierarchical clustering merges the data samples into ever-coarser clusters, yielding a tree visualization of the resulting cluster hierarchy. t-SNE maps the data samples into 2d space so that the proximity of the samples to one another can be visualized.

Exercise 1: Visualizing hierarchies Exercise 2: How many merges?Exercise 3: Hierarchical clustering of the grain data Exercise 4: Hierarchies of stocks Exercise 5: Cluster labels in hierarchical clustering Exercise 6: Which clusters are closest?Exercise 7: Different linkage, different hierarchical clustering!Exercise 8: Intermediate clusterings Exercise 9: Extracting the cluster labels Exercise 10: t-SNE for 2-dimensional maps Exercise 11: t-SNE visualization of grain dataset Exercise 12: A t-SNE map of the stock market

Dimension reduction summarizes a dataset using its common occuring patterns. In this chapter, you'll learn about the most fundamental of dimension reduction techniques, "Principal Component Analysis" ("PCA"). PCA is often used before supervised learning to improve model performance and generalization. It can also be useful for unsupervised learning. For example, you'll employ a variant of PCA will allow you to cluster Wikipedia articles by their content!

Exercise 1: Visualisierung der PCA-Transformation Exercise 2: Korrelierte Daten in der Natur Exercise 3: Dekorrelation der Getreidemessungen mit PCA Exercise 4: Principal components (Hauptkomponenten)Exercise 5: Intrinsische Dimension Exercise 6: Die erste Hauptkomponente Exercise 7: Varianz der PCA-Merkmale Exercise 8: Intrinsische Dimension der Fischdaten Exercise 9: Dimensionsreduktion mit PCA Exercise 10: Dimensionsreduktion der Fischmessungen Exercise 11: Ein tf-idf-Worthäufigkeit-Array Exercise 12: Wikipedia clustern, Teil I Exercise 13: Wikipedia clustern, Teil II

Aktuelle Übung

In this chapter, you'll learn about a dimension reduction technique called "Non-negative matrix factorization" ("NMF") that expresses samples as combinations of interpretable parts. For example, it expresses documents as combinations of topics, and images in terms of commonly occurring visual patterns. You'll also learn to use NMF to build recommender systems that can find you similar articles to read, or musical artists that match your listening history!

Exercise 1: Non-negative matrix factorization (NMF)Exercise 2: Non-negative data Exercise 3: NMF applied to Wikipedia articles Exercise 4: NMF features of the Wikipedia articles Exercise 5: NMF reconstructs samples Exercise 6: NMF learns interpretable parts Exercise 7: NMF learns topics of documents Exercise 8: Explore the LED digits dataset Exercise 9: NMF learns the parts of images Exercise 10: PCA doesn't learn parts Exercise 11: Building recommender systems using NMF Exercise 12: Which articles are similar to 'Cristiano Ronaldo'?Exercise 13: Recommend musical artists part I Exercise 14: Recommend musical artists part II Exercise 15: Final thoughts