Tekstgegevens voorbereiden voor modelinvoer

Eerder leerde je hoe je woordenboeken maakt van index-naar-woord en andersom. In deze oefening splits je de tekst in tekens en ga je verder met het voorbereiden van de data voor supervised learning.

Teksten in tekens opdelen lijkt misschien vreemd, maar het wordt vaak gedaan voor tekstgeneratie. Bovendien is het voorbereidingsproces hetzelfde; alleen de manier waarop je de tekst splitst verandert.

Je maakt de trainingsdata met een lijst van teksten van vaste lengte en hun labels, namelijk de bijbehorende volgende tekens.

Je blijft werken met de gegevensset met citaten van Sheldon (The Big Bang Theory), beschikbaar in de variabele sheldon_quotes.

De functie print_examples() print de paren, zodat je kunt zien hoe de data is getransformeerd. Gebruik help() voor meer details.

Deze oefening maakt deel uit van de cursus

Recurrent Neural Networks (RNN's) voor taalmodellen met Keras

Cursus bekijken

Oefeninstructies

Definieer step gelijk aan 2 en chars_window gelijk aan 10.
Voeg de volgende zin toe aan de variabele sentences.
Voeg de juiste positie van de tekst sheldon toe aan de variabele next_chars.
Gebruik de functie print_examples() om 10 zinnen en volgende tekens te printen.

Praktische interactieve oefening

Probeer deze oefening eens door deze voorbeeldcode in te vullen.

# Create lists to keep the sentences and the next character
sentences = []   # ~ Training data
next_chars = []  # ~ Training labels

# Define hyperparameters
step = ____          # ~ Step to take when reading the texts in characters
chars_window = ____ # ~ Number of characters to use to predict the next one  

# Loop over the text: length `chars_window` per time with step equal to `step`
for i in range(0, len(sheldon_quotes) - chars_window, step):
    sentences.____(sheldon_quotes[i:i + chars_window])
    next_chars.append(sheldon_quotes[____])

# Print 10 pairs
print_examples(____, ____, 10)

Code bewerken en uitvoeren

Deze oefening maakt deel uit van de cursus

Recurrent Neural Networks (RNN's) voor taalmodellen met Keras

SkillTag.level.advancedSkillTag.label

4.8+

Begin de cursus gratis

In this chapter, you will learn the foundations of Recurrent Neural Networks (RNN). Starting with some prerequisites, continuing to understanding how information flows through the network and finally seeing how to implement such models with Keras in the sentiment classification task.

Exercise 1: Introductie van de cursus Exercise 2: Het aantal parameters van RNN en ANN vergelijken Exercise 3: Sentimentanalyse Exercise 4: Sequence-to-sequence-modellen Exercise 5: Introductie tot taalmodellen Exercise 6: Wennen aan tekstdata Exercise 7: Tekstgegevens voorbereiden voor modelinvoer

Huidige oefening

Exercise 8: Nieuwe tekst transformeren Exercise 9: Introductie tot RNN in Keras Exercise 10: Keras-modellen Exercise 11: Keras-preprocessing Exercise 12: Je eerste RNN-model

You will learn about the vanishing and exploding gradient problems, often occurring in RNNs, and how to deal with them with the GRU and LSTM cells. Furthermore, you'll create embedding layers for language models and revisit the sentiment classification task.

Exercise 1: Vanishing and exploding gradients Exercise 2: Exploding gradient problem Exercise 3: Vanishing gradient problem Exercise 4: GRU and LSTM cells Exercise 5: GRU cells are better than simpleRNN Exercise 6: Stacking RNN layers Exercise 7: The Embedding layer Exercise 8: Number of parameters comparison Exercise 9: Transfer learning Exercise 10: Embeddings improves performance Exercise 11: Sentiment classification revisited Exercise 12: Better sentiment classification Exercise 13: Using the CNN layer

Next, in this chapter you will learn how to prepare data for the multi-class classification task, as well as the differences between multi-class classification and binary classification (sentiment analysis). Finally, you will learn how to create models and measure their performance with Keras.

Exercise 1: Data pre-processing Exercise 2: Prepare label vectors Exercise 3: Pre-process data Exercise 4: Transfer learning for language models Exercise 5: Transfer learning starting point Exercise 6: Word2Vec Exercise 7: Multi-class classification models Exercise 8: Exploring 20 News Groups dataset Exercise 9: Classifying news articles Exercise 10: Assessing the model's performance Exercise 11: Precision-Recall trade-off Exercise 12: Precision or Recall, that is the question Exercise 13: Performance on multi-class classification

This chapter introduces you to two applications of RNN models: Text Generation and Neural Machine Translation. You will learn how to prepare the text data to the format needed by the models. The Text Generation model is used for replicating a character's way of speech and will have some fun mimicking Sheldon from The Big Bang Theory. Neural Machine Translation is used for example by Google Translate in a much more complex model. In this chapter, you will create a model that translates Portuguese small phrases into English.

Exercise 1: Sequence to Sequence Models Exercise 2: Text generation examples Exercise 3: NMT example Exercise 4: The Text Generating Function Exercise 5: Predict next character Exercise 6: Generate sentence with context Exercise 7: Change the probability scale Exercise 8: Text Generation Models Exercise 9: Create vectors of sentences and next characters Exercise 10: Preparing the data for training Exercise 11: Creating the text generation model Exercise 12: Neural Machine Translation Exercise 13: Preparing the input text Exercise 14: Preparing the output text Exercise 15: Translate Portuguese to English Exercise 16: Congratulations!