Step 07 · The recurrent step
State In, State Out
- One hidden state
- One input
- Two matrices
Sequences
Thirty-two steps on data where order is the signal: what a recurrent step actually computes, why long sequences lose the beginning, and how attention solved it — worked on sequences short enough to follow by hand.
Secure checkout · Instant download · 60-day guarantee
Thirty-two steps across five chapters — order as signal, the recurrent step, the failure of long chains, attention, and reading a transformer block.
Sound familiar?
Queries, keys and values. Three words that sound like a database and explain nothing about what gets multiplied by what. So attention stays a picture you have seen rather than a computation you could perform.
On a sequence of four tokens, with vectors of length three, the whole mechanism fits on one page and takes about fifteen minutes to work through. After that the diagrams finally mean something.
Sample pages
Every step uses a sequence short enough to write out in full — the state, the update, the scores, the weights and the result.
Step 07 · The recurrent step
Step 16 · Why it forgets
Step 24 · The whole mechanism
32 steps in total
Language models reproduce the patterns in their training text, including the unpleasant ones. The final chapter is direct about what that means.
What's inside
It derives the problem before the solution — you feel why recurrence struggles before attention is offered as the answer.
Four tokens, vectors of length three. Small enough that every number in the attention matrix is on the page.
The vanishing gradient is computed, not asserted. You watch the number collapse over twenty steps before anything replaces it.
Query, key and value are named for what they multiply, with the shapes written out, rather than by analogy to databases.
What you get
Thirty-two steps across five chapters — order as signal, the recurrent step, the failure of long chains, attention, and reading a transformer block.
The offer
When Order Is The Signal takes sequences apart in thirty-two steps — the recurrent step, the gradient that collapses, and attention computed in full on a sequence short enough to check by hand.
Instant digital download. One-time payment of $10.00. Educational material about how neural networks work — not a course, not a certification, and not a promise about employment or salary.
Questions
It explains the mechanism they are built from, on a scale you can verify by hand. It does not cover training at scale, which is an engineering problem rather than a conceptual one.
Because attention is the answer to a problem, and the problem is worth feeling. The recurrent chapters are short and they make the rest obvious.
Yes — plain NumPy for every mechanism, and a PyTorch version at the end of each chapter so you can see which parts the library provides.
No. Shapes and gradients are re-derived where they are needed.
No. That is the transfer-learning book. This one is about how the mechanism computes.
No — it's a digital guide (PDF), emailed to you right after purchase. Nothing is shipped.
You're protected by a 60-day money-back guarantee. Email us within 60 days for a full refund — no questions asked.
Thirty-two steps from one recurrent update to attention computed in full on four tokens.
Secure checkout · Instant download · 60-day guarantee
This is educational material about how neural networks work. It is not a course, a certification, a bootcamp or a career programme, and it makes no promise about employment, salary or professional outcomes of any kind. The code and the worked figures are teaching examples, written for clarity rather than for production: read them, adapt them, and test anything you reuse. Library APIs change often, so the method is what carries over, not the exact call signature. Nothing here is professional, legal or financial advice. Models reproduce the patterns and the biases of the data they are trained on; deploying one that affects people carries responsibilities this book does not cover.