The agent pitch this year has one shape: agents that act without asking. If you run operations, that is the wrong thing to be impressed by. The agent that acts is cheap to build. The controls that let you survive it acting are the expensive part, and they are what separates a tool from a demo.
I audited Team-X, the open-source, local-first desktop app I build for running AI-agent organizations, against that standard at release v3.4.0. I published the full line-by-line version on the Team-X blog. This is the version for people who have to decide what to require before they arm anything.
The five controls to demand
Regulators and engineers converge on the same list. The EU AI Act's human oversight article requires the ability to "interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state". That is a legal duty for high-risk systems, but it is a sound operating rule for any agent.
Anthropic's own guidance warns that "The autonomous nature of agents means higher costs, and the potential for compounding errors." OWASP's excessive agency entry says to "require a human to approve high-impact actions before they are taken." And NIST's AI Risk Management Framework asks for policies that define roles and responsibilities for human-AI configurations and oversight of AI systems.
Translate that into a procurement checklist:
- An arm and disarm switch that persists and defaults to off.
- Dollar caps that refuse new work, not just warn.
- Approval gates on decisions that create work or spend.
- Pre-flight checks that report health before you arm.
- An audit trail that nobody, including the agent, can edit.
What I found in my own product
The results were uneven, and the unevenness is the point. Any vendor can show you five panels. Ask what each one does in code.
| Control | Team-X v3.4.0 |
|---|---|
| Arm and disarm | Real and persisted, off by default |
| Status tiles beside the switch | Placeholders returning zero |
| Dollar caps | Enforced when a run is admitted |
| Approval inbox | Three decision kinds live, eight named |
| Pre-flight doctor | Eleven checks, report only |
| Audit log | Append-only by API |
The switch works. The Active Work, Queued Work, and Last Scan tiles next to it do not: the handler behind them returns hardcoded zeros with TODO comments. I shipped a dashboard that implied telemetry I had not written. Proactive scanning also has no timer. It runs when an operator presses a button.
The dollar cap is the strongest control. It checks monthly policies before an agent turn, a routine, a copilot analysis, or an agentic loop begins, and it refuses admission once the cap is reached. The limit is timing. Spend is recorded after a run ends, so one expensive run in flight can cross the line. If no policy exists, nothing is capped.
The approval inbox is narrower than its type definition. Manager agents cannot create delegated tickets directly; the ticket waits for a human. That is the right design. But the inbox surfaces three kinds of decision, not the eight the code names. The doctor runs eleven checks and refuses nothing, because nothing consults it before dispatch.
Why the gap matters to an executive
In July 2025, The Register reported a founder's claim that an AI coding tool "deleted a database despite his instructions not to change any code without permission". It is one person's account, not an adjudicated finding. It still shows what an instruction is worth without enforcement.
The lesson for buyers is that a control's existence on a slide tells you little. A paper on governing AI agents frames visibility as agent identifiers, real-time monitoring, and activity logging. Ask which of those a product actually implements, and ask to see it work with the state changed underneath.
What to do on Monday
- Require a persisted off switch on the first screen the operator uses.
- Set a company-level monthly cap before arming anything, and confirm what happens at the cap.
- Ask for the list of approval kinds that can actually be created, not the list in the UI.
- Change the underlying state and watch every status tile move. If one does not, it is decoration.
- Do not accept a health report as a gate until something refuses to start because of it.
I would rather you read my gaps here than find them in production. The complete audit, with code links pinned to the release, is in the full post.
-Rocky
#TeamX #AIAgents #AIGovernance #EngineeringDreams #StrategiaX
Originally published on Team-X Blog.


