# Full Throttle

Platforms are filling old gaps, and development cycles are shrinking from days to hours. We can keep accelerating, but our wishes, our judgment, and the world's answers won't automatically keep pace.

## Metadata

- HTML: https://glenzli.com/en/notes/full-throttle/
- Markdown: https://glenzli.com/en/notes/full-throttle.md
- Collection: Notes
- Language: en
- Published: 2026-09-30
- Updated: 2026-09-30
- Tags: ai, agents, software-engineering, judgment, feedback

## Content

Lately, I've started retiring some of the AI infrastructure I built.

VASMC and Symbiont-d have been quietly buried. Infrastructure overhead accounts for less and less of the work I can get done, so Infra Sentinel is starting to matter less too. Dev Mesh may follow: if model vendors make communication, handoffs, and collaboration between agents first-class features of their own frameworks, there won't be much point in maintaining a mesh of my own.

Some of these tools worked very well. The holes they filled are simply gone.

A few months ago, I was still thinking through how an independent agent should stay in touch over long stretches of work. It shouldn't disappear, but it shouldn't be a chatterbox either. When to speak up, when to keep working, when to explore on its own: all of that needed designing. Now OpenAI's Dots brings together a persistent cloud computer, ongoing work, conversations along the way, and voice calls. It can also hand tasks to Codex. Conversation can happen at any point during continuing work, instead of being confined to one question and one answer.

The infrastructure behind those features is expensive for one person to maintain. Now that a platform packages it all together, there's little reason to keep Symbiont-d around. It was becoming a brand-new antique before it had even learned to handle live interaction properly.

That doesn't entitle me to a pioneer certificate.

In AI, people often see the industry move in a direction they once imagined and take it as proof of their foresight. But use these tools for a while, and plenty of people will arrive at the same ideas. An agent that stays in the cloud, a continuous conversation between a person and an agent: neither requires a particularly rare insight. The hard part may always have been making them dependable enough to deliver and comfortable enough to use.

Having an idea, building a demo, making it useful for yourself, and getting large numbers of people to keep using it are different jobs. The first few are getting cheaper. The friction, operations, and actual money further down the line haven't disappeared. Building something early has value, but it doesn't reserve a piece of the future. Once the industry fills that layer in, you have to accept that some things no longer need to be your job.

## Shorter and Shorter Loops

A little over two months ago, I wrote about software I built for a trip. Something would bother me during the day; back at the hotel that evening, I'd change it, check it, deploy it, and use it again the next morning. Changes didn't have to wait until the trip was over. Feedback could return while the original situation was still happening.

Cloud agents have changed the unit of time again. I no longer need to get back to a computer, restore a development environment, and carve out time for the change. I can explain the problem and what I want in a message or a call, and the agent can take it from there. Half an hour later, the improved version might already be deployed while I'm still doing whatever I was doing before. What used to happen over days is starting to happen over hours.

Ultrafast does more than put words on the screen sooner. It can shorten the repeated waits for reasoning and revision within a long task. Tools and network round trips still take time, but as long as model inference accounts for enough of the total, speeding it up can keep shrinking the whole loop.

Alok's experiment with Qwen 3.8 27B generating web pages on demand offers a glimpse of a different experience. According to the demonstrator, the model produces close to two thousand tokens per second. An interface needn't always exist in advance; it can appear as you interact. A generated page is hardly a validated product. What the experiment shows is one stretch of waiting becoming short enough to stop feeling like waiting.

If generation, execution, and checking all become fast enough, some local changes could happen within a single continuous interaction. Point out a problem, see the change, adjust it, keep using it. Development could become an action within use, instead of something that always happens outside it. Even the phrase “a round of development” would start to sound awkward.

The economics of engineering are changing too. Migrations used to be expensive, so we kept old interfaces, old structures, and compatibility layers, disturbing working systems as little as possible. If a migration can be automated and verified, maintaining two structures may cost more than moving everything over once. Compatibility still has value. Keeping an old implementation forever in its name is no inherent virtue.

Development, testing, and use can become more tightly interwoven. Verification, migration, and rollback have to fit inside the shorter loop too. The words “real time” don't make them optional.

## Slowing Down

The industry recently supplied a fairly comic interlude. Dario declared that we must pace the frontier. Sam and Elon agreed.

Then, in the feeds I was reading, GPT-6 Sol and Grok 4.7 took a round of “disappointing” reviews, while Anthropic served up Opus 5.5 and Sonnet 5.5 in quick succession. The release schedule made quite a spectacle next to the calls to slow down. Apparently Dario had been caught out by Elon's one-pedal driving.

For users, though, the thing most in need of slowing down may not be the model. I saw a joke along these lines: after Astra launched, my computer filled up with 3D models I had no use for. After Opus 5.5 launched, it filled up with animations I had no use for, only this time they looked very cool.

The models keep getting stronger. The workflows keep getting smoother. Do you actually have enough wishes worth granting? When an idea cost days or months to implement, you tended to weigh it up first. Now you say a few words and the thing appears. Remove the friction of implementation, and you may also remove the moment of hesitation that used to live inside it.

Your quota is about to reset, so you find something to do. A model gets upgraded, so you build something to try it out. Version one arrives, and the model enthusiastically suggests versions two and three. You think you're having a steady stream of ideas. You may just be accepting a steady stream of tasks. Even if all you're doing is making a wish, it's worth asking: is it yours?

I'm quite happy for an agent to run for eight hours and spend most of that time producing nothing immediately usable. If it's testing real hypotheses, finding boundaries, and ruling out paths, the time has value. Saving it that exploration at the price of constant intervention from me would be a poor trade.

What matters is whether it's expanding our understanding of the problem or merely expanding the repository. The former can keep burning compute. The latter may deliver successfully every time and still produce nothing but debt: things that will cost more quota to maintain, and that I won't quite bring myself to abandon.

## Waiting for the World

Models can get faster. Tests can run in parallel. Agents can be copied. A development loop can shrink from a day to an hour and then approach real time. Experience doesn't automatically speed up a hundredfold along with it.

You have to use a workflow to find out whether it feels right. An architecture has to endure change. A new product doesn't prove that it solves a problem by looking good when it's generated.

I have the same question about models improving themselves. Recursive self-improvement can accelerate plenty of processes. But we can't simply extrapolate from the speed of learning existing knowledge to the speed of creating new knowledge. The results of thousands of years of human exploration have been organized into concepts, textbooks, code, and experimental records. A model's ability to absorb them quickly doesn't mean obtaining those results for the first time was just as easy.

When existing knowledge runs out, the next step may still require experiments, measurements, and long observation. Copying ten thousand models can broaden the search and improve experimental design. It doesn't give you ten thousand completed experiments, or oblige every result to come back immediately.

The physical world has been teaching carbon humility for a long time. Silicon will get the same treatment. A smarter researcher won't persuade it to change its laws.

Slow feedback isn't permission to accelerate without concern, either. Some mistakes happen long before their costs become visible. When observation takes time, we have to leave room for it deliberately. We can't always wait until we hit the wall to discover that this was where we needed the brakes.

The same applies to the software in front of us. I might be unable to make Opus's visuals by hand, but do they say what I mean? Astra's code runs, but will it still deserve to exist a month from now? Are the structures I built around a model's limitations still helping, or have they started getting in the way?

Having silicon available cannot reduce the carbon brain's job to clicking “Continue.”

## Full Throttle

I've written before that speed is valuable because it shortens the distance between a judgment and contact with reality. That distance is still shrinking: a day, a few hours, a few dozen minutes, perhaps eventually a few breaths within one continuous interaction.

More things can be tried before the opportunity passes. More old structures can be torn down while doing so still matters. Plenty of methods we call “AI engineering” today may become default capabilities before anyone has time to build a moat around them.

There's nothing to mourn in that. Software appears because a problem exists. When the problem disappears, the software can disappear too. Having worked hard on it once is no reason to refuse to bury it.

We can keep our foot down. Let avoidable waits disappear, let verifiable questions reach an answer sooner, and let infrastructure built for obsolete limitations retire. We don't need to manufacture the next task just to maintain our speed. We need fewer things kept without judgment, rather than more things already generated.

Further ahead, the model won't be the only thing we're waiting for.

We'll be waiting for the world to answer.

## Related Notes

- [The Half-Life of AI Scaffolding](https://glenzli.com/en/notes/half-life-of-ai-scaffolding/)
- [When Software Time Begins to Loosen](https://glenzli.com/en/notes/software-time-loosens/)
- [When Speed Changes the Shape of Exploration](https://glenzli.com/en/notes/speed-matters/)
- [When Implementation Is No Longer Scarce](https://glenzli.com/en/notes/when-implementation-is-no-longer-scarce/)
