1 Pick the text
This is the only thing your model will ever read. Each of these is a different kind of structure rather than a different book, because what a model really learns is a grammar — and a chess score has a much stricter one than a rhyme.
2 Break it into pieces
Nothing in the network reads letters, so every distinct character is assigned a number. Below is the window of text the model actually gets handed, and underneath it the guessing game it is asked to play: for each character, name the next one.
3 Build the network
Choose a starting size, then change the parts. Drag a piece onto the block, or click a piece and then click where it belongs. Where you put the normalization is the point of the exercise: above the layer is one design, below the ⊕ is another.
4 Look inside
The same network three ways — as the shape of the data moving through it, as source code marked up against the plain build, and as the maths. Change a part in stage 3 and all three move together.
5 Teach it
Hand it a window of text, let it guess every next character at once, score how wrong it was, and adjust. That loop is the whole of pre-training. Every run you do is kept, so you can put two builds side by side.
—
6 See what it looks at
Pick any character in the string below and the bars show how much your trained model leaned on each earlier character when predicting it. This is the mechanism the rest of the architecture exists to serve.
7 Make it write
Give it a start and it predicts one character, sticks it on the end, and goes round again. Everything you have ever typed at a chatbot came out of this loop.
—
8 Make it better
Training is a loop: look at the curve, guess what is holding it back, change one thing, run it again. This reads the runs you have actually done and says what to change. The levers underneath are roughly in the order worth trying.