How AI Agents Learned to Run the Show
Published: 2026-04-05T16:37:33 · Updated: 2026-04-22T08:35:34Z
The first AI agents were not agents in any meaningful sense of the word. They were lookup tables dressed up in clever language. In the 1950s, when Alan Turing proposed his famous imitation test and researchers at Dartmouth coined the term artificial intelligence, the field was animated by a deceptively simple idea: what if a machine could reason the way a person does. What followed was decades of rule-based systems, expert systems, and decision trees that could only navigate problems their creators had explicitly anticipated. ELIZA, developed at MIT in 1966, could mimic a therapist convincingly enough to fool some users, but it understood nothing. It matched patterns and returned responses. There was no reasoning happening, no memory of what was said three sentences earlier, and absolutely no capacity to act on anything.
That limitation held for longer than most people realize. Through the 1980s and 1990s, AI research produced impressive-sounding systems that remained fundamentally brittle. IBM's Deep Blue beat Garry Kasparov at chess in 1997, which generated enormous headlines, but Deep Blue could not play checkers. It was the world's most expensive specialist. The agents being theorized in academic circles during this period, systems that could perceive an environment, form goals, and take actions to achieve them, remained largely conceptual. The computer was not there. The data pipelines were not there. And the models underpinning the whole enterprise were not sophisticated enough to handle ambiguity, which is most of what real-world tasks consist of.

The shift started to become visible around 2017, when the transformer architecture emerged from Google Brain and rewrote what was possible in language understanding. Models trained on massive text corpora began exhibiting something that looked less like pattern matching and more like comprehension. By 2020, OpenAI's GPT-3 could write coherent paragraphs, answer questions, and switch between tasks based on natural language instructions alone. That was genuinely new. But GPT-3 was still a closed system. You could talk to it. You could not hand it a task and walk away. It had no memory between conversations, no ability to browse the internet, no capacity to execute code or interact with other software. It was brilliant at generating text and useless at doing anything with the world outside its context window. The gap between generating language and taking action turned out to be enormous.
The real inflection point came in late 2024, when Anthropic released the Model Context Protocol, a standard that allowed developers to connect large language models to external tools in a consistent, predictable way. Before MCP, connecting an AI model to a database or an API was bespoke work every single time. After MCP, it became composable. You could snap tools together and hand the whole thing to a model that could decide, on its own, which tool to use and when. Google then released Agent2Agent in April 2025, which addressed a different problem entirely: not how agents use tools, but how agents communicate with each other. Suddenly, you did not just have one agent calling APIs. You had networks of agents delegating subtasks to each other, checking each other's work, and routing problems to whichever agent was best equipped to handle them. The architecture of AI had fundamentally changed.
What happened to the economics of all this is just as important as the technical story. Per-million-token pricing fell from around thirty dollars in early 2023 to between ten cents and two dollars fifty in early 2026. That kind of cost collapse does not just make existing use cases cheaper. It creates entirely new ones. Agentic workflows that were financially impractical at thirty dollars per million tokens became obvious at ten cents. Founders who had been building static AI features started rebuilding their entire products around agents that could run autonomously for hours. On autonomous task benchmarks, frontier models that could sustain independent work for about four minutes in early 2024 extended that to fourteen and a half hours by early 2026. That is not incremental. That is a different category of tool.

What autonomous agents can do right now is worth being specific about, because the hype and the reality are still not perfectly aligned. Agentic browsers emerged in 2025, tools that reframed the browser as an active participant rather than a passive interface, capable of booking a vacation rather than just searching for one. AI agents can now perform financial trading, plan and execute multi-step processes, and integrate with external systems to solve problems in real time. Small teams are using them to run entire customer support pipelines, content operations, lead generation workflows, and software development cycles with minimal human intervention. The Klarna case is instructive here. The company deployed an AI agent that handled 2.3 million conversations in its first month, equivalent to 700 customer service agents, and cut headcount by 40 percent. Then customer complaints spiked, satisfaction scores dropped, and the CEO publicly admitted they went too far before beginning to rehire human agents for complex cases. The lesson is not that agents do not work. It is that agents work very well for the predictable parts of a business and struggle with the judgment calls that require real contextual understanding and empathy.
IBM's Chris Hay put it plainly: in 2024, agents were small and specialized, the email writer, the research helper. By 2026, with reasoning capabilities improving, they can plan, call tools, and complete complex tasks across multiple environments simultaneously. The question for founders right now is not whether to build with agents. It is how to structure your business so that agents handle the repeatable work while humans concentrate on the decisions that actually require them. Gartner projects that by 2028, at least 15 percent of work decisions will be made autonomously by AI agents, up from virtually zero in 2024. That trajectory is aggressive, but given how fast the last two years moved, it is hard to argue it is unrealistic. What began as a lookup table in a university lab has spent seventy years becoming something that can run your operations while you sleep. The honest question now is not whether your business will use agents. It is whether you will build with them before someone who understands them better does it for you.