Most AI agents are governed by a paragraph of persona text. "You are a senior engineer." If the model ignores it, nothing happens. For an operations lead, that is not a control. It is a hope.
I build Team-X, an open-source, local-first desktop app for running AI-agent organizations. Its core bet is that an agent should be defined as a role spec: a file with schema-validated fields the software can enforce, plus a prompt. I audited the shipped code at release v3.2.1 to see whether the bet holds. The answer is partly yes, and the gaps are the useful part.
Authority belongs outside the model
The OWASP Top 10 for LLM Applications 2025 defines Excessive Agency as "the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction." Its prevention advice is the sentence every executive sponsoring an agent program should memorize: "Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not."
Anthropic describes agents as "systems where LLMs dynamically direct their own processes and tool usage." When the model chooses the tools, the boundary around those choices cannot live in the model's instructions. It has to live in the runtime.
What the code enforces, and what it does not
I sorted every field in the role spec. Here is the result at v3.2.1.
| Field | Status at v3.2.1 |
|---|---|
| Tool allow and deny lists | Enforced twice for MCP tools: filtered before the model sees them, refused again at call time |
| Hierarchy level | Enforced for hiring, project decomposition, and manager assignment, with a spelling gap I describe below |
| Capabilities | Used to score who fits a task; grants nothing |
| Decision authority, KPIs, escalation targets | Validated, then descriptive. I found no code that reads them |
| Preferred model tier and providers | Not read from the role by the provider factory |
The practical consequence: a field named decision_authority that says "final" is a label, not a permission. If your agent platform has fields like that, ask which ones the software actually checks. I had to ask that of my own product.
Least privilege, and where my defaults fall short
The principle is old. Saltzer and Schroeder wrote in 1975 that "Every program and every user of the system should operate using the least set of privileges necessary to complete the job." NIST's glossary restates it as each entity being granted "the minimum system resources and authorizations that the entity needs to perform its function."
Team-X has the mechanism. It does not yet have the defaults. In 56 of the 57 shipped role files, both tool lists are empty, and an empty allowlist means unrestricted. Only the CEO role sets lists. That is the reverse of least privilege: the safe configuration exists, and the shipped roles mostly do not use it.
Two further findings from the audit. First, the built-in orchestrator tools, such as messaging a colleague, are exempt from the lists. Second, the VP level is spelled with an underscore in the role files and a hyphen in several runtime gates, and the normalizer does not bridge the two. Traced through the source, that refuses a VP at the hire gate and at the approval check for project decomposition. It is a finding from reading the code, and I have not yet reproduced it in a running build.
Why the design still matters to a buyer or builder
A study of multi-agent failures, Why Do Multi-Agent LLM Systems Fail?, builds a taxonomy of 14 failure modes in three categories, one of which is system design issues. I am not quoting a failure rate from it. The point is narrower: how you specify agents is a design decision with its own failure category, not a prompt-tuning detail.
A role spec that separates what the model is told from what the runtime allows gives you something auditable. You can diff it, sign it, and review it. A persona paragraph gives you none of that.
What I would do in your position
- Ask every agent vendor where authorization runs when the model requests a tool. If the answer is the system prompt, treat it as a suggestion.
- Inventory the fields in any role or agent definition and mark each as enforced, scored, or descriptive. Delete or enforce the descriptive ones.
- Start from deny. Give each role an explicit allowlist for anything that touches money, customers, files, or the network.
- Test a denial on purpose, and check that the platform logs a violation event.
- Pin your audits to a version. I audited a tagged release, not a moving branch, so every claim here can be checked.
I wrote up the full field-by-field audit, with code links pinned to the release, in A role is not a prompt. The code is open, and I would rather you find the next gap than have me hide it.
-Rocky
#TeamX #AIAgents #LeastPrivilege #EngineeringDreams #StrategiaX
Originally published on Team-X Blog.
