Citirea fragmentelor ca data.frame
În exemplul anterior, am citit fiecare fragment în funcția de procesare ca matrice, folosind mstrsplit(). Această abordare funcționează bine când citim date rectangulare în care toate elementele dintr-o coloană au același tip. Când nu este cazul, poate fi mai util să citim datele ca data.frame. Poți face acest lucru fie citind fragmentul ca matrice și apoi convertindu-l în data.frame, fie folosind direct funcția dstrsplit().
Acest exercițiu face parte din cursul
Procesarea scalabilă a datelor în R
Instrucțiuni pentru exercițiu
- În funcția
make_msa_table(), citește fiecare fragment ca data frame. - Apelează
chunk.apply()pentru a citi datele pe fragmente. - Calculează totalul valorilor din fiecare coloană adunând toate rândurile.
Exercițiu interactiv practic
Încearcă acest exercițiu completând acest cod de exemplu.
# Define the function to apply to each chunk
make_msa_table <- function(chunk) {
# Read each chunk as a data frame
x <- ___(chunk, col_types = rep("integer", length(col_names)), sep = ",")
# Set the column names of the data frame that's been read
colnames(x) <- col_names
# Create new column, msa_pretty, with a string description of where the borrower lives
x$msa_pretty <- msa_map[x$msa + 1]
# Create a table from the msa_pretty column
table(x$msa_pretty)
}
# Create a file connection to mortgage-sample.csv
fc <- file("mortgage-sample.csv", "rb")
# Read the first line to get rid of the header
readLines(fc, n = 1)
# Read the data in chunks
counts <- ___(fc, ___, CH.MAX.SIZE = 1e5)
# Close the file connection
close(fc)
# Aggregate the counts as before
___