TL;DR: Everyone has the same models, and nobody has the same space around them, which is where the real variable sits. The framework for measuring that gap ran in a management journal back in 2005, written for an entirely different purpose. Two questions decide it: can you codify the work, can you measure the result. What ruled work out then, untransferable team knowledge, is exactly what you can now supply. A colleague is onboarded once; an agent without memory is onboarded every time. Specify once, where the agent reads by default; unclear thinking now runs a meter.
Want to listen to this article?
Subscribers to the Orion Playbook newsletter can listen to the AI-Generated Audio version of this article for free. Already a subscriber? Log in here
Ask a chat assistant for an article, and the median conversation runs thirteen rounds of back-and-forth. Put the request to an agent with custom skills that runs inside your own files, and the median session contains a single human prompt. Same artifact, same request. The obvious explanation is model choice, yet the gap survives when the model is held constant.
So the variable is the space the work happened in. What matters is how much of the description the machine can reach on its own, instead of being handed it again every session. Everyone has the same models. Nobody has the same space.
The tools do not learn from feedback. Too much context has to be supplied by hand every session. They break on the edge cases. A colleague who cannot remember what the client likes, repeats corrections you already made, and needs the whole background again every Monday is a hire nobody onboarded. Companies have an entire discipline for that failure, and it is not machine learning.
A corporate lawyer at a mid-sized firm, quoted in MIT's study of enterprise deployments, said it plainly: "It repeats the same mistakes and requires extensive context input for each session. For high-stakes work, I need a system that accumulates knowledge and improves over time." An onboarding gap, described by someone with no word for it. Extensive context input for each session. Every session.
I have argued before that offshorability was an early read on automatability, because both waves measure one property. What I left out was the test itself. Ravi Aron and Jitendra Singh published it in Harvard Business Review in 2005, as the framework for deciding which work could leave your building. Some of you ran it yourselves, with a transition binder and a team in another time zone. It is the framework you need now, and it has been sitting unread.

The test was two questions. Can you codify the work, write down how it is done across every situation that occurs, including the ugly ones? And can you measure the result, meaning tell whether the output is right without redoing it? Their instruction was to fix the measurement in-house before sending the work anywhere. Measurement is the one that bites. I write this newsletter with an AI co-creator, and the hard part has never been the drafting. It is saying what done looks like, then improving that definition, then improving it again. We are still at it, and I expect we always will be.
What ruled a process out then was knowledge that lived in your team and could not travel: the client history, the feel for how a market behaves. The axis has flipped. The thing that used to close the door is exactly what you can now supply. The framework was never wrong. It was waiting. So write down what happened, what done looks like, and how it is done. Write it once, into the place the agent reads by default, and check it every so often for drift, because a specification you never reread stops matching the work. The steps that fail the test are the ones worth your week.
Then run it on the right list. The instinct is to start with the task you most resent. Resentment marks where you act as a transport layer instead of a judge, and it is still a poor ranking instrument. In Anthropic's usage survey, the more experience someone has, the less of their own work they believe a machine could take. Wisdom or blind spot, it is unhelpful. Recurrence sits on a calendar and in a sent folder. Resentment picks the candidates; recurrence ranks them. None of this replaces buying something. It is what makes you competent at buying, because you cannot ask a vendor to fit a process you have never written down.
One objection here should stand: capability keeps absorbing specification work. The people who build these systems already concede that better models need less prescriptive engineering, which makes much of today's context-engineering scaffolding around a temporary limitation. Scaffolding is not a moat. About the artifacts, the objection is probably right: your prompts will rot. The ability to take your own work apart and say what finished looks like survived the move from VBA macros to offshore teams, and it will survive this one.
A colleague is onboarded once. An agent without a memory is onboarded on every execution. The bill will not tell you which is which, because token spend tracks the value of the work as much as the waste in it. What would tell you is tokens per completed run of one recurring task, tracked over a few months. Almost nobody keeps that number, and it is cheap to start. Unclear thinking used to cost you nothing at the moment you did it. Now it runs a meter.

Images source: ChatGPT Images / Claude Opus / Gérard Métrailler
Sources
Aron, Ravi, and Jitendra V. Singh. "Getting Offshoring Right." Harvard Business Review, December 2005. https://hbr.org/2005/12/getting-offshoring-right. Accessed 2026-08-01.
Anthropic Applied AI team. "Effective context engineering for AI agents." Anthropic Engineering, 29 September 2025. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents. Accessed 2026-08-01.
Challapally, Aditya, Chris Pease, Ramesh Raskar, and Pradyumna Chari. The GenAI Divide: State of AI in Business 2025. MIT Project NANDA, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf. Accessed 2026-08-01.
Massenkoff, Maxim, Eva Lyubich, Szymon Sacher, Zoe Hitzig, Shaoyi Zhang, Ryan Heller, and Peter McCrory. "Anthropic Economic Index report: Cadences." Anthropic, 26 June 2026. https://www.anthropic.com/research/economic-index-june-2026-report. Accessed 2026-08-01.


