ÎncepețiÎncepe gratuit

Citirea fragmentelor ca data.frame

În exemplul anterior, am citit fiecare fragment în funcția de procesare ca matrice, folosind mstrsplit(). Această abordare funcționează bine când citim date rectangulare în care toate elementele dintr-o coloană au același tip. Când nu este cazul, poate fi mai util să citim datele ca data.frame. Poți face acest lucru fie citind fragmentul ca matrice și apoi convertindu-l în data.frame, fie folosind direct funcția dstrsplit().

Acest exercițiu face parte din cursul

Procesarea scalabilă a datelor în R

Vezi cursul

Instrucțiuni pentru exercițiu

  • În funcția make_msa_table(), citește fiecare fragment ca data frame.
  • Apelează chunk.apply() pentru a citi datele pe fragmente.
  • Calculează totalul valorilor din fiecare coloană adunând toate rândurile.

Exercițiu interactiv practic

Încearcă acest exercițiu completând acest cod de exemplu.

# Define the function to apply to each chunk
make_msa_table <- function(chunk) {
    # Read each chunk as a data frame
    x <- ___(chunk, col_types = rep("integer", length(col_names)), sep = ",")
    # Set the column names of the data frame that's been read
    colnames(x) <- col_names
    # Create new column, msa_pretty, with a string description of where the borrower lives
    x$msa_pretty <- msa_map[x$msa + 1]
    # Create a table from the msa_pretty column
    table(x$msa_pretty)
}

# Create a file connection to mortgage-sample.csv
fc <- file("mortgage-sample.csv", "rb")

# Read the first line to get rid of the header
readLines(fc, n = 1)

# Read the data in chunks
counts <- ___(fc, ___, CH.MAX.SIZE = 1e5)

# Close the file connection
close(fc)

# Aggregate the counts as before
___
Editează și rulează codul