I had been using LLMs for over a year before I sat down to actually understand what one is. I kept dodging it. The math looked intimidating. The papers were written for researchers. I kept settling for analogies and hoping nobody would ask me to go a layer deeper.
Then I came across Andrej Karpathy's deep-dive video. Three and a half hours of him walking through what is actually going on when you type into a chat box. He used to lead AI at Tesla and was a co-founder at OpenAI. He decided to explain the whole thing on YouTube, slowly, with the assumption that the listener has no special background. I watched it twice.
Some of what I took away, in the simplest version I can give it.
An LLM is, at base, a next-token predictor. You give it a sequence of text. It guesses what fragment of a word comes next. Then it does the same thing again, with its previous guess included. It is a very fast, very compressed pattern-matcher trained on a huge amount of text.
It does not know things the way a person knows things. It has internalised statistical regularities of language. Some of those regularities happen to encode facts. Many of them encode patterns of how facts are usually expressed, which is why it can sound certain about things that are not true. The hallucination problem is not a bug they will eventually fix. It is a feature of the architecture.
The model is not searching its memory. It is generating one token at a time, in the shape of an answer.
There is a second stage on top, where the base model is fine-tuned with examples and human feedback to make it act like a helpful assistant. This stage is what makes ChatGPT feel like ChatGPT, rather than like a stochastic auto-complete. Most of what we mean when we say 'an LLM is good' is actually 'the post-training is good.'
What this changed for me as an operator: when a model fails at a task, I now ask whether the failure is upstream of the model. Was the context wrong. Was the prompt missing structure. Did I expect judgment from a system that does not have judgment. Sometimes the answer is yes. The model was doing exactly what its architecture allows. I was the one expecting more.
I am not pretending I now understand the inside of a transformer. I do not. What I have is a working mental model of what an LLM is and is not, and that has been enough to procure differently and brief teams differently. The math will keep waiting.
If you have not watched Karpathy and you are making AI decisions for a business, I think it is worth the three hours. What is the simplest version of an LLM you could explain to your board the morning after?