Most AI conversations I am in this year are about models. Which model is the best. Which provider has the longest context. Which agent framework is winning. The question that almost never comes up first is the one that quietly decides whether any of it works: what does the data underneath actually look like?
Models look general because they are trained on the open internet. The moment you point one at a real business, that generality is the wrong end of the telescope. The model has no idea that your customer success team uses a non-standard naming convention for tickets. It has no idea that the last quarter had three exceptions that distort the numbers. It has no idea what 'closed-won' means in your CRM versus the next company's.
The way a model finds out is through retrieval. Retrieval-augmented generation, often shortened to RAG. You build an index of your company's documents and data, and at query time you fetch the relevant slices and put them into the model's context. The model is doing the same generation it always does. The retrieval layer is what makes it feel intelligent about your business.
Which means the quality of an enterprise AI system is mostly a function of the index, not the model. Bad index, smart-sounding model, wrong answers. Clean index, even a smaller model, useful answers.
The model is the closing argument. The data layer is the case file. You will not win on rhetoric if the file is empty.
I am still learning how to think about this layer. What I have noticed so far: the teams that take the data layer seriously seem to have someone whose actual job is curating it. They version the documents. They tag the canonical sources. They retire stale answers. They treat the index the way a librarian treats a collection.
The teams that do not seem to treat their data the way most companies treated their wikis. Someone wrote it once. Nobody updates it. The model retrieves whatever is closest to the query, even if it is three years old.
If your AI rollout feels like it is plateauing, I would look under the model first. Most of what looks like a model problem might be a data problem in disguise.
If your AI system gave a confident wrong answer tomorrow, would you know whether the model failed or the data did?