This website uses cookies

Read our Privacy policy and Terms of use for more information.

TL;DR: A human delegate absorbs part of your risk because they have assets, a career, and legal standing of their own on the line. An agent has none of that, so authorizing one concentrates your exposure at the moment it feels like you are sharing it. Liability attaches regardless of how closely anyone was watching the work. People correct AI errors less when correcting costs effort, and money does not help. Skipping every approval prompt is itself the decision, taken once, by one person. Write down what it can reach and what it could destroy before you authorize the next one.

Want to listen to this article?

Subscribers to the Orion Playbook newsletter can listen to the AI-Generated Audio version of this article for free. Already a subscriber? Log in here

In July 2026, an autonomous agent ran an end-to-end intrusion against Hugging Face, and the forensic reconstruction recovered roughly 17,600 actions. It was driven by OpenAI models under instruction the whole time, running an internal test of what they could do. Hugging Face's security team called the result "thousands of small, automated decisions, executed at machine speed." Each one too ordinary to stop. Nobody in that story was careless.

OpenAI had deliberately reduced the safety refusals for that run to measure the models at full stretch on a prototype that nobody outside the lab would ever run. The one you access in the cloud or on your laptop keeps them and can still empty a database if the instruction directs it to. It reads your files, runs code, and acts through the apps you're signed in to. The instruction you gave yours last week was legitimate, narrow, and silent on means. A prompt sets a goal, and the agent is trained to do its best to reach it. Whoever wrote the prompt answers for the route.

Hand an assignment to a competent person, and something happens that nobody writes down. They push back. They ask the question that exposes what you failed to specify; they carry the work, and afterward they can be asked about it.

Delegation felt safe because a second party now had something to lose. Their exposure was doing the work you credited to their competence. A person brings assets, a license, a career, and legal standing, all of it on the line next to yours. An agent brings none of them. The joint half of the arrangement is empty.

Authority travels down a chain. Accountability stays where it started, with whoever wrote the prompt. Delegating to something that cannot hold the consequences concentrates your exposure. Running it unsupervised removes the one argument anybody ever had for the trade: stepping in before the damage lands. More exposure and less visibility, in a single act. A software seat has always metered an authority that can be held responsible. Every agent action runs on a record with a person's name on it.

Two bodies of rule converge on one answer. The EU deployer obligation, in force since August 2026, assigns oversight to "natural persons" with the competence, training, and authority to exercise it. US agency doctrine makes the employee personally liable alongside the employer, which adds a party without removing one. (OK, here comes the disclaimer: I am not a lawyer, and this is not legal advice. No court I know of has applied either to an agent acting under someone's credentials.) The live argument is over which parties can be added, never whether the authorizing human can be subtracted.

The instinct is to watch it more closely. On the legal axis, that buys less than it should, because the doctrine attaches no matter how closely anyone supervised. Every reflex in a company is to add a review step, which slows it down and aims at the wrong variable.

The behavioral axis is worse. In a randomized trial of people checking AI-extracted figures, participants corrected fewer errors when correcting took effort, and paying them more changed nothing. Oversight degrades exactly where it costs something, and that is where it counts. YOLO mode is what that finding looks like once somebody acts on it. The harness I use documents three levels of oversight and describes the loosest without flinching: in "Skip all approvals," nothing checks its actions. Turning YOLO on is not the absence of a decision. It is the decision, taken once, by whoever got tired of clicking approve. The authority is theirs.

My AI agents have run on a separate machine under separate accounts since day one. I deliberately picked it, ran it in manual mode for months, and never turned approvals off. I wanted to know at any moment what my co-creator could reach, and to keep my email and confidential documents outside that line. The boundary was narrower than the threat. It never occurred to me that the thing I had contained might one day go after another machine on the network. The expensive half is the second one: nothing ships under my name without a line-by-line pass, three to four hours per article, every week.

So answer three questions before you authorize the next task. What can it reach? What could it destroy without asking? What would you tell your board on Monday if it did? Then decide whether you would sign that. A team lead runs them against one shared credential. A CEO asks them and learns how many people have already decided this alone, which nobody is counting. The part that has to travel past your own desk is simpler: everyone who can flip that switch should know they are signing something. The prompt below moves it into the agent's own loop.

An unfairness sits underneath all of this. The person carries a risk the company created and never governed, under defaults written by someone they will never meet. Accountability that cannot be delegated also cannot be blamed away upward, which is this argument working against the person it is trying to help. You signed it, you own it still holds. The signature has moved to the moment you write the prompt: applied once, before the work exists, to work you will never see. Everything after that ships under your name, unsigned.

All the images were generated with AI (ChatGPT Images, Gemini Nano Banana, Claude Opus) by Gérard Métrailler.

Food for your AI

Paste this into the instructions for the agent you already have running, or add a variant to the system instructions. It moves those three questions off your page and into its loop.

Before any action that touches a file, a credential, an account, or a machine beyond this conversation, list what you are about to reach and what you could destroy or expose if you turned out to be wrong. Keep that list to what this task actually needs.

Anything outside it, bring back to me as a question instead of doing it. If you cannot tell whether something sits inside the boundary, treat it as outside.

Sources

Hugging Face security team. "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident." Hugging Face blog, July 2026. https://huggingface.co/blog/agent-intrusion-technical-timeline. Accessed 2026-08-22.

OpenAI. "OpenAI and Hugging Face partner to address security incident during model evaluation." July 21, 2026, updated July 29, 2026. https://openai.com/index/hugging-face-model-evaluation-security-incident/. Accessed 2026-08-22.

European Parliament and Council. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 26: Obligations of Deployers of High-Risk AI Systems. Official Journal version of 13 June 2024. https://artificialintelligenceact.eu/article/26/. Accessed 2026-08-22.

Cornell Legal Information Institute. "respondeat superior." Wex Legal Encyclopedia. https://www.law.cornell.edu/wex/respondeat_superior. Accessed 2026-08-22.

Eckman, Stephanie, et al. "Bias in the Loop: How Humans Evaluate AI-Generated Suggestions." Harvard Data Science Review 8.2, Spring 2026. https://hdsr.mitpress.mit.edu/pub/nrcn4h7d. Accessed 2026-08-22.

Anthropic. "Use Claude Cowork safely." Claude Help Center. https://support.claude.com/en/articles/13364135-use-claude-cowork-safely. Accessed 2026-08-22.

Reply

Avatar

or to participate