Explore a distribuição dos dados

Quando queremos anonimizar um conjunto de dados amostrando dados de forma bem realista, precisamos adquirir algum conhecimento de domínio e estatístico sobre os dados. Como vimos, encontrar a distribuição de probabilidade da coluna de interesse é essencial.

Neste exercício, você vai explorar a coluna business_travel de uma versão simplificada do conjunto de dados de RH da IBM.

O DataFrame foi importado como hr e o numpy como np. Como mencionado no capítulo anterior, o pandas foi importado como pd para este e os demais capítulos do curso.

Este exercicio faz parte do curso

Privacidade de Dados e Anonimização em Python

exercicio interativo prático

Tente este exercicio completando este código de exemplo.

# Print the absolute frequencies of each unique value
print(____)

Editar e Executar Código

Este exercicio faz parte do curso

Privacidade de Dados e Anonimização em Python

AvançadoNível de habilidade

4.9+

Comece o curso gratuitamente

Get ready to apply anonymization techniques such as data suppression, masking, synthetic data generation, and generalization. In this chapter, you’ll learn how to distinguish between sensitive and non-sensitive personally identifiable information (PII), quasi-identifiers, and the basics of the GDPR. You'll also encounter real-life examples of what can go wrong if you don't follow these best practices.

Exercise 1: What's private, and why do we care?Exercise 2: Privacy is power Exercise 3: Is it sensitive or non-sensitive?Exercise 4: Suppression of sensitive attributes Exercise 5: Data masking and data generation with Faker Exercise 6: Masking sensitive PII Exercise 7: Removing names with faker Exercise 8: Anonymizing with data generalization Exercise 9: Reducing identification risk with generalization Exercise 10: Data aggregation and data generalization Exercise 11: Top and bottom coding White House salaries

Discover how to anonymize data by sampling from datasets following the probability distribution of the columns. You’ll then learn how to apply the k-anonymity privacy model to prevent linkage or re-identification attacks and use hierarchies to perform data generalization in categorical variables.

Exercise 1: Anonimizando dados categóricos Exercise 2: Explore a distribuição dos dados

Exercicio Atual

Exercise 3: Amostrando da mesma distribuição de probabilidade Exercise 4: Anonimizando dados contínuos Exercise 5: Distribuições diferentes Exercise 6: Amostragem da melhor distribuição contínua Exercise 7: Introdução ao k-anonymity Exercise 8: Atributos de privacidade Exercise 9: Generalizando em intervalos Exercise 10: Generalizando dados usando hierarquias Exercise 11: Usando hierarquias para dados categóricos Exercise 12: Aplicando k-anonimidade a um conjunto de dados

Learn about differential privacy, the model used by major technology companies such as Apple, Google, and Uber. In this chapter, you’ll explore data by generating private histograms and computing private averages in data. You’ll also create differentially private machine learning models that allow businesses to increase the utility of their data.

Exercise 1: Introduction to differential privacy Exercise 2: Epsilon (ϵ): the magic number Exercise 3: Histograms with differential privacy Exercise 4: Privacy budgets Exercise 5: Using privacy budgets Exercise 6: When no budget is left Exercise 7: Exploring data with a privacy budget accountant Exercise 8: Differentially private machine learning models Exercise 9: Build a differentially private classifier Exercise 10: Predicting salaries Exercise 11: Differentially private clustering models Exercise 12: Pre-processing data Exercise 13: Segmenting customers

In this final chapter, you’ll learn how to apply dimensionality reduction methods such as principal component analysis (PCA) to anonymize large multi-column datasets. You’ll then use Faker to generate realistic and consistent datasets, and scikit-learn to create synthetic datasets that follow a normal distribution. Lastly, you’ll tie everything you learned in this course together as you combine multiple techniques to safely release datasets to the public.

Exercise 1: PCA for anonymization Exercise 2: Anonymization of high-dimensional data Exercise 3: Data masking with PCA Exercise 4: Generating realistic datasets with Faker Exercise 5: Consistent synthetic dataset Exercise 6: Datasets with the same probabilistic distribution Exercise 7: Creating synthetic datasets using scikit-learn Exercise 8: Generating datasets for classification Exercise 9: Generating datasets for clustering Exercise 10: Safely release datasets to the public Exercise 11: Exploring and pseudonymizing a dataset Exercise 12: Preparing employee data for safe release Exercise 13: Great work!