Dr. Aditya Vikram Kashyap, AI governance and innovation leader exploring how autonomous systems reshape institutions, markets and power.
We spent the AI boom learning to rank models, but the systems built around them now decide whether the answers we get are useful, safe or simply wrong.
Someone in finance asks the company’s new assistant whether a supplier contract can be canceled without penalty. The answer, which arrives in seconds, is fluent, confident and cites a clause. It is wrong. The clause was renegotiated last spring; the assistant was reading the old version, because that is what the system handed it.
The natural verdict is that the AI got it wrong. A more useful one is that the model did exactly what it was built to do with what it was given. The failure lived in plumbing that nobody sees.
For years, the public conversation about AI has been about models: which is smartest, which company is ahead. But what you experience when you use an AI assistant is not the model alone. It is a system built around the model, including instructions, document retrieval, memory, tools, permissions, testing and rules for when a person should step in.
Two organizations can license the same model and get opposite results. One system knows the current policy and asks before it acts. The other is a fluent liability. Much of what separates them is what each built around the model.
Nobody judges an aircraft by benchmarking its engine. The engine is indispensable, but it says little about whether a flight will arrive; that depends on everything else on board. In this case, the model is the engine. The rest of the aircraft decides whether you arrive.
A model never meets your company or your contract. It meets a representation assembled for it: the documents a search step picked out in the order the code put them. Everything else is absent. A Stanford-led study found that models can use relevant information less reliably when it appears in the middle of a long input than near the beginning or end. Many apparent failures of intelligence are failures of information architecture.
Memory sounds like a straightforward improvement until a stored mistake continues shaping answers after the original error has been forgotten. A preprint tested five agent-memory systems by giving them a revoked policy and its replacement. None consistently enforced the revocation. When both versions were retrieved, the systems favored the old policy every time, and the agents acted on it in about four out of 10 trials, regardless of the model’s capability tier. When an AI remembers something incorrectly, what mechanism makes it forget?
Stakes change once a system stops answering and starts acting. In February 2026, the director of alignment at Meta’s superintelligence lab connected an agent to her inbox, asked it to suggest actions it could take and told it to do nothing without her approval. By her account, the agent’s working memory filled, an automatic step condensed the conversation and the instruction to wait didn’t survive. It started deleting emails while she typed “stop.” The underlying model hadn’t changed. A sentence seemingly went missing from its memory.
A typical benchmark asks a model a question once and grades the answer. A deployed system may handle the same kind of request hundreds or thousands of times. One customer-service benchmark found that the best model tested completed about 61% of retail tasks on a single attempt; its pass^8 score, which measures consistency across eight repeated trials, fell below 25%.
A single score cannot see that gap. NIST’s agent standards initiative starts from the observation that an agent’s usefulness is constrained by how it interacts with outside systems and data, and its work runs from interoperability standards to research on agent security and identity.
The obvious objection is that if the system around the model matters so much, why spend so much on better models? The premise is half right. A large jump in what models can do changes what can be built at all. Researchers showed that small gains in per-step accuracy can compound into large gains in the length of task a model can complete. Frontier capability still expands the frontier. What reaches the user, however, depends on how a particular system puts that capability to work.
When Air Canada’s chatbot told a grieving passenger that a bereavement discount could be claimed after flying, contrary to Air Canada’s policy, the airline argued, in effect, that the chatbot was a separate legal entity responsible for its own actions. A British Columbia tribunal called that a remarkable submission and said it made no difference whether the misinformation came from a static page or a chatbot. A small-claims ruling about one refund illustrates the broader principle: a company can be held responsible for what its system tells customers, whether that information comes from a webpage or a chatbot.
Set the examples side by side. The contract answer was a retrieval failure: the system fetched the wrong document. The revoked policy was a memory failure. So, on her account, was the deleted inbox, although that incident exposed a deeper problem: only a sentence in a conversation stood between the agent and the delete command.
The refund case points to a different kind of failure. At the institutional level, nobody had adequately checked what the chatbot was telling customers. In each case, the user experiences the same thing: the AI gave a wrong answer or took the wrong action. But the underlying causes are different. If an organization treats every breakdown as a model problem, it will keep reaching for the same solution: a smarter model. That will not fix a stale document, a failed memory mechanism or an instruction that disappears before an agent acts.
Return to the box. You type; an answer appears. The interface invites you to imagine that the intelligence lives in the thing that produced the words. Behind the box sits a set of decisions someone made: which documents the model gets to see, what it may remember, what it may do and when a human has to be asked. The model is still extraordinary, but its intelligence is only part of the question. The rest is the system built around it, and whether anyone is checking that system when it fails.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

