aidactic.

Shorts / Transformers

2 min 54

Transformers

How a transformer, the neural network design behind many large language models, turns text into tokens and vectors, lets each token weigh the others through attention, stacks layers to end in a probability for the next token, and why the cost of attention bounds how much text a model can take in at once.

Free to watch, no sign-in, captions.

More shorts