Two Years In: From a Chatbox to a Fleet
What changed between the morning I typed questions into a box and the morning an agent briefed me before I woke.
This morning, before I opened my laptop, an AI agent running on a small rented server sent me a briefing. What happened overnight across my projects. What needs my decision today. What it already handled without me.
Two years ago, AI was a textbox I pasted questions into.
The distance between those two mornings is the most interesting thing I have been part of professionally, and I want to write it down while it is still fresh. Partly for the people who keep asking me where to start. Partly because writing it down made me realize how far this has come.
Year one: I was the integration layer
In the beginning the model was brilliant and the workflow was primitive. I copied context in, copied answers out, and stitched everything together by hand. Every session started from zero. The model knew nothing about my work unless I retyped it.
The first real lesson had nothing to do with prompts. The bottleneck was context. The model was smart. It was also blind.
The first shift: giving AI eyes
I started building so the model could see my world instead of me describing it. A browser extension that turns any webpage into clean Markdown with one keystroke, because feeding AI clean context beats feeding it noise. An open-source MCP server that gives any AI assistant direct access to 200 million peer-reviewed papers, built for my doctoral research and now part of my daily workflow.
Then MCP matured and everything clicked. Instead of pasting context into a chat, I plug models directly into the sources: my git repositories, my notes, my email, my calendar, live research databases, analytics. More than ten connections now. The model went from an outsider I briefed to a colleague with access.
The second shift: from typing to directing
Agents changed my unit of work. On a normal evening I have 15 to 20 of them running, each with its own job. One is refactoring a project. One is researching. One is drafting. One is validating an automation before it ships. I review, redirect, approve.
My job stopped being typing and became directing. I read far more than I write now. The skill that matters most is the one nobody talks about: deciding what is worth doing, and stating it clearly enough that someone else, silicon or human, can run with it.
The third shift: always-on
The latest step was taking myself out of the loop for the routine 80 percent. I gave an agent a server and a phone number. It runs around the clock on a small VPS, has read-only access to 17 git repositories, executes scheduled jobs through the night, and reports to me on my phone. Every morning I get a digest: done, blocked, needs you. I answer from my phone like I am texting a colleague who never sleeps.
Underneath all of it sits a second brain: interlinked notes holding every project, decision, and idea from the last two years. I recorded a flythrough of the knowledge graph recently and it looks like neurons firing. Fitting, because that is how it functions. The agents read it, write to it, and connect things I forgot I knew. Models come and go. The context is mine, and it compounds.
This is not a hobby story
The same playbook runs at a much larger scale in the work I do during the day. Shipping agents that operations teams use every day. Building prompt packs for entire departments. Training non-technical people. Redesigning processes around what the technology can now do.
The lesson is identical at every scale: the technology is ready before the operations are. Buying licenses changes very little. Redesigning how work flows through people and agents changes everything. That is operations work and leadership work, and very few people are doing it seriously yet.
So is it 10x or 100x?
People ask me this constantly. Honest answer after two years: both, and neither, depending on the task.
Research that took weeks takes an afternoon. First drafts that took days take minutes. Software I would have needed a team for, I now ship alone on weekends. On those tasks, 10x is conservative and 100x is sometimes literal.
Judgment did not get faster. It is now the bottleneck, which means it is now the job.
Deciding what to build, earning trust, knowing what good looks like. Those still run at human speed. The multiplier applies to execution. Direction stayed scarce.
Where this is heading
Chat was the demo. Agents are the product. Fleets are the operating model. Three things I am confident about.
One. Always-on agents become normal. Within a couple of years, serious operators will run them the way they run email today. The morning agent digest will be as unremarkable as the morning inbox.
Two. Context becomes the moat. Models keep improving underneath you. A new model generation arrives and the entire system gets smarter overnight, without me changing a line. What cannot be rented is accumulated context: the notes, the data, the documented processes. Organizations that treat context as infrastructure will pull away from organizations that treat AI as a chat window.
Three. Every organization will need someone who defines what AI does for it. Which workflows, which guardrails, which sequence, which teams get trained first. That is a leadership role sitting between strategy and operations. It barely existed two years ago. It will be on every org chart in five.
Two years ago I typed questions into a box. Today, before breakfast, I direct a small organization where most of the staff is silicon. I have never been more enthusiastic about a technology, and the people who will get the most out of it are the ones who start treating it as an operating model now, not a tool.
If you are wondering where to begin: pick one workflow you own end to end, give the model real context, ship the result, repeat weekly. The distance from chatbox to fleet is shorter than you think.