The Five Levels of Working With AI: What Kitchens, Cockpits, and Job Sites Already Know | WhatAICanDo Skip to content

The Five Levels of Working With AI: What Kitchens, Cockpits, and Job Sites Already Know

Devin
Published date:
23 min read

Core argument: “How should I organize my AI assistants?” sounds like a brand-new question. It is not. Every demanding trade that has survived a century — the professional kitchen, the airline cockpit, the operating room, the construction site — has had to solve the harder version of it: how to organize capable work so that it survives contact with reality, with change, and with human limits. This essay proposes a five-level framework — skill, process, methodology, architecture, philosophy — for organizing what those trades already know, and for deciding what carries over when the worker is an AI. Each level has real failures behind it, the kind you can count in bodies and dollars. The essay walks the ladder the long way — what ordinary work discovered first, the invariant underneath, and the concrete translation to AI collaboration — because the AI question is not new. It is the oldest question in productive work, asked about a new kind of worker.

This site has covered parts of this stack before — the prompt engineering formula, prompts upgraded into protocols, what agents are made of (not a feature, but a loop). This essay is upstream of all of them: a way of seeing the whole ladder, borrowed from work that existed long before software.

1. The short answer

Ask how work gets organized where failure is expensive — Saturday service, a night flight, an operating list, a high-rise — and you find the needs met in layers that rhyme across trades. This essay sorts them into five:

A quick test for telling the floors apart: each level’s product differs. Skill produces the work; process produces the checklist and the review; methodology produces the playbook and the memory; architecture produces the constraint; philosophy produces the signature.

One honesty note before the climb. This is a map, not a law. Other fields cut the same territory differently: systems science ranks the places you can intervene in a system, as in Donella Meadows’ ladder from tuning parameters up to transcending paradigms [37]; psychology’s Dreyfus model cuts across skills, from rule-following novice to intuitive expert [38]. The five levels here are cut where they are for one reason: each names a layer that AI visibly bends — a floor whose keepers AI either replaces or forces to change shape. Readers who prefer the other maps will lose nothing; the territory is the same.

The AI question — how to organize assistants that can already do the skill-looking part — lands on this exact ladder. What follows climbs it one level at a time. At each level: what ordinary work learned, the invariant under it, and what it changes about how you work with Claude, ChatGPT, or a coding agent.

2. Level 1 — Skill: the recipe and the principle

Psychology drew a fault line inside skill long before AI did. In the Dreyfus model of skill acquisition, novices follow context-free rules — if X, do Y — while experts read whole situations with intuition built from long, varied experience [38]. A related gap shows in kitchens in plain sight: a line cook can execute one station on a known menu, but moved to a strange station on a slammed Saturday night, the cook running on memorized solutions flounders, while the one who understands why heat, fat, and time behave as they do improvises. Memorized solutions and understood principles are different assets, and ordinary work treats them as such.

AI collapses the gap between the two routes from the outside. A large model arrives holding an enormous library of memorized solutions — recipes, code samples, precedents — with little of the principle-level grounding that lets a professional judge a situation. The result is the pattern practitioners know well: output that looks journeyman-grade right up to the moment it is confidently wrong, and the last 30% — edge cases, context, correctness — stays on the human [1]. The prompt, at this level, is a work order handed to a worker with an enormous library and no shop floor.

The trap of Level 1 is not using it; it is staying here. Collecting prompt templates is hiring a new stranger for every task. The rest of this essay is what the other floors are for.

3. Level 2 — Process: the checklist was invented because memory fails

On October 30, 1935, Boeing’s Model 299 — the prototype of the B-17 — took off at Wright Field, stalled, and crashed, killing Major Ployer Hill, chief of the Army’s Flying Branch and the pilot leading the evaluation. The investigation found no design flaw: the pilot had forgotten to release the gust lock. The plane needed at least thirty steps before takeoff, and, as one newspaper put it in Atul Gawande’s telling, was “too much airplane for one man to fly.” A group of test pilots responded by inventing the pilot’s checklist. With it, the same design flew 1.8 million miles without an accident [2][3].

That story is the whole argument for Level 2 in miniature: the failure was not ignorance, and it was not lack of skill. It was ineptitude in Gawande’s sense — knowledge that exists but is not reliably applied [4]. And aviation drew the loop even tighter than most professions realize. Before every flight, a pilot runs IMSAFE — Illness, Medication, Stress, Alcohol, Fatigue, Eating — a formal self-assessment of the pilot’s own state, not the machine’s [5]. And after hard missions, the U.S. Army formalized the After-Action Review (TC 25-20, 1993): a structured debrief asking what was planned, what actually happened, why, and what to do differently — run by the unit itself, feeding the next mission’s plan [6]. Preflight checks the world and yourself; acceptance defines “safe to fly” before the wheels move; the retrospective turns one flight’s lessons into the next flight’s checklist.

Medicine reproduced the numbers. In eight hospitals across eight cities, the WHO Surgical Safety Checklist cut deaths from 1.5% to 0.8% and complications from 11.0% to 7.0% across nearly 7,700 patients [7]. In 103 Michigan ICUs, a five-item central-line bundle drove the median infection rate from 2.7 per 1,000 catheter-days to zero, sustained across eighteen months [8]. Even the model teams see this. OpenAI’s own testing found that letting GPT-5 carry its reasoning from previous turns back into context — same model, same task, but stateful context instead of stateless requests — lifted a realistic retail-agent benchmark from 73.9% to 78.2% [36]. What moved the number was retained context, not a cleverer sentence.

The translation to AI collaboration is direct, and it is already codified by the teams that do this at scale:

And ordinary work supplies the warning label too. When Ontario mandated checklists across 101 hospitals, mortality and complications did not significantly move [11]. The study measured outcomes, not behavior, so it cannot say exactly why — but it shows that mandating the artifact is not the same as changing the practice, and NASA’s human-factors work names one candidate mechanism: under workload, crews revert to memory and skip items [12]. A process works when it changes what people do, not when it exists. Your saved “AI workflow” document counts for exactly as much.

4. Level 3 — Methodology: kitchens run on method, not recipes

Ask a chef what makes a professional kitchen different from a home kitchen and the answer is usually not a dish. It is mise en place — “everything in its place,” the discipline of preparing and organizing every ingredient and tool before service, in the professional curriculum since at least the CIA’s The Professional Chef [13]. Mise en place is not a recipe. It is what makes recipes swappable: a new dish is a variation on a station that already works. One level up sits the method that improves methods. Toyota’s improvement kata is a repeating four-part routine — understand the direction, grasp the current condition, set the next target, experiment toward it — coached daily by managers [14]. The result Toyota has reported is staggering in scale: its chairman Eiji Toyoda claimed 1.5 million worker suggestions a year with 95% implemented, as reported by Masaaki Imai [15]. The point is not any one suggestion. It is that the shop has a routine for improving routines, and it runs every day.

But the deepest thing Level 3 knows is about how skill actually moves between people. Michael Polanyi opened The Tacit Dimension with the sentence this whole subject turns on: “we can know more than we can tell” [16]. The journeyman never learned the trade from the manual; the learning happened through graduated participation and correction — what Lave and Wenger documented across tailors, midwives, and quartermasters [17]. Entire professions run on standards that were never written down, because the apprentice absorbed them by being there.

Here is the inversion that makes AI different from every previous tool: the AI cannot stand in your kitchen and absorb the osmosis. It will not learn what your team means by “good,” which vendors burned you, why the third step matters — unless someone writes it down. Every organization that adopts AI agents discovers that it has been forced into the first honest knowledge-management audit of its existence: the tacit standards that lived in one person’s hands must become explicit artifacts. The industry’s names for those artifacts are memory, playbooks, skills, and standing instruction files — the open AGENTS.md format alone claims over 60,000 open-source projects [18], and the memory-systems field (structured, self-edited memory blocks [19]; with reported double-digit quality gains and over 90% token savings versus stuffing context [20]) exists precisely to industrialize what a good shop foreman used to carry in his head.

The pathology here is the binder nobody follows — methodology that produces documents instead of behavior. The test is the one mise en place passes every night: the abstraction must emit a working process and survive an unfamiliar order.

5. Level 4 — Architecture: codes are written in blood

On March 25, 1911, 146 workers died in the Triangle Shirtwaist factory fire — many because the doors were locked; the reform wave that followed, OSHA writes, “marks a century of reforms that make up the core of OSHA’s mission” [21]. In 1942, 492 people died at Boston’s Cocoanut Grove nightclub, and the exit provisions that came out of it — doors that swing outward, panic hardware, capacity limits — were absorbed into the codes that govern almost every public building since [22]. And on July 17, 1981, in Kansas City, two suspended walkways at the Hyatt Regency fell and killed 114 people: a fabricator’s change had split each continuous hanger rod in two, doubling the load on the connection, and the change sat inside shop drawings returned stamped with the engineering review seal — nobody recalculated. Both engineers lost their licenses [23]. Time and again, the codes that bite are past failures encoded as constraints, so the next crew doesn’t relearn them at the same price.

That is what architecture means in ordinary work — not elegance, but designed constraints that make ordinary skill safe. High-reliability organizations operationalize it into habits: preoccupation with failure, reluctance to simplify, sensitivity to operations, commitment to resilience, deference to expertise [24].

AI collaboration has grown its own version of this layer faster than most people notice, under the name context engineering [25] — because what a model sees determines what it does: model recall degrades as context grows (Anthropic calls it context rot), so the craft is curating “the smallest possible set of high-signal tokens” [26]. The artifacts are architectural, exactly like the kitchen’s station layout: what the agent can see and touch (permission scopes), which rules are enforced deterministically rather than by asking nicely (hooks), how parallel work stays isolated (workspaces), where state lives between sessions. And the field’s most instructive moment is architectural too. On June 12, 2025, Cognition argued that architectures which break shared context should be ruled out by default, because “actions carry implicit decisions” [27]. One day later, Anthropic published its orchestrator-worker research system, reporting that multi-agent runs use about 15× more tokens than chat, with token usage explaining 80% of performance variance [28] — while cautioning that complex, interdependent tasks are exactly where the pattern struggles. The disagreement is narrower than “parallelism: yes or no,” and the useful extraction is that the dependency structure of the tasks chooses the architecture: loosely coupled work profits from separate contexts; tightly coupled work punishes them. Either way, the argument is not about prompts. It is about the shape of the system — which is exactly what Level 4 means.

The pathology is familiar from every construction boom: speculative complexity — building codes for a skyscraper when the job is a garden shed. The antidote comes from the same labs building the biggest systems: “find the simplest solution possible, and only increasing complexity when needed” [29]. A constraint earns its place by closing a class of failure you have actually had.

6. Level 5 — Philosophy: the signature you cannot delegate

A professional engineer’s license exists so that somewhere in the system there is one name attached to the work. The model law defines it as responsible charge — “the direct control and personal supervision of engineering work” — and is blunt about the part most at risk in the AI era: “reviewing drawings or documents after preparation without involvement in the design and development process does not satisfy the definition” [30]. Medicine carries the same line in its training requirements: “although the attending physician is ultimately responsible for the care of the patient,” supervision cannot be delegated [31]. In mature fields, accountability is the one thing that does not subcontract.

Ordinary work even predicted the way AI tempts us to fake it. Human-factors research since the early 1990s has shown that when people monitor highly reliable automation, their attention decays and they miss the failures — automation-induced “complacency” is measurable in the lab [32], with the canonical framing being misuse: overreliance on automation that still needs a human somewhere [33]. A rubber stamp is not a review, and thirty years of aviation psychology says the urge to stamp will win unless the system is designed against it.

And when execution becomes abundant, judgment is what runs out. Herbert Simon said it in 1971, about information: “a wealth of information creates a poverty of attention” [34]. AI is that sentence applied to work: a wealth of output creates a poverty of judgment. The questions on the human side of the line — what is worth building, what counts as good enough to ship, who is accountable when automated work causes damage — are exactly the ones OpenAI’s agent guidance tells builders to design for, by planning for human intervention with explicit escalation triggers on failure thresholds and high-risk actions [35]. In ordinary-work language: every system has a seat marked decides, and the design question is who sits in it. The pathology at this level is twofold — paralysis (endless discourse, no artifacts) and rubber-stamping (the signature without the charge). Both are cured the same way ordinary professions cure them: by staying connected down the ladder, philosophizing with your hands on the work.

7. One task through the ladder: this essay

To show the levels doing work rather than just naming things, here is one real task run through all five — the production of the piece you are reading.

Level 1. The first prompt — “draft an essay on organizing AI work” — returned a fluent draft in minutes. It read like every other AI essay: clean sentences, borrowed frameworks, no receipts. Nothing in the output itself could even tell you why it wasn’t publishable.

Level 2. What changed that was a loop, not a better prompt. Preflight: gather primary sources before writing, and set the bar in writing — every load-bearing claim needs a source a reader can open. Acceptance: a second reader goes through the draft against that bar, not against vibes. This is not hypothetical: an earlier draft of this essay attributed a benchmark gain to the wrong mechanism, and the error was caught in exactly that review step — by opening the source instead of trusting the summary. Retrospective: the correction fed the next draft, the way an after-action review feeds the next mission.

Level 3. What survives the task is not the essay; it is the residue. The keyword families already researched sit in a registry, so next quarter’s scan doesn’t re-litigate them. The reviewer’s corrections become a standing checklist; the house style guide absorbs every “never again.” The next essay starts from those artifacts, not from a blank prompt — that is the difference methodology makes.

Level 4. Some constraints no longer depend on anyone’s memory. Drafts live as reviews until the editor approves, and the publishing pipeline will not carry an unapproved piece — a gate installed after the one slipped through. Frontmatter, internal links, and citation format are enforced by the build rather than by vigilance. Like the exit door that swings outward, the gate works on the day everyone is tired.

Level 5. The editor decides what publishes — and by design that reader is not the author grading their own homework. No fluency the model adds changes what the piece is for, or who is accountable for it being true.

Nothing in that story was solved by a better model. Each level added structure around the same model — which is the claim of this essay, now with receipts.

8. The failure modes, in one table

LevelOrdinary work’s name for itThe cure lives at
1 — SkillThe recipe memorized, the principle never learned2 — wrap the task in a loop
2 — ProcessThe mandated checklist that changed nothing [11]3 — ask what the process is for
3 — MethodologyThe binder nobody opens2 — make it emit one working process
4 — ArchitectureCodes for a skyscraper, job is a shed1–3 — simplest thing that passes the check
5 — PhilosophyThe rubber stamp; or the debate that never ships4 — design the seat marked “decides”

Read the table and one pattern appears: no level audits itself. Skill cannot see its own brittleness; a process cannot judge whether its checklist still serves; methodology drifts without a working process to test itself on; architecture overbuilds without contact with real tasks; philosophy without artifacts becomes talk. And the corrections run in both directions — up, when a skill needs a loop around it; down, when an abstraction needs to be re-tested against a real case. The ladder is less a promotion track than a feedback circuit you keep wired together: the pilot who still runs the walk-around, the chef who still works the line.

9. The transfer move: four questions to steal

This is where 举一反三 — reasoning from one case to the rest — becomes a practice. You can audit any AI-assisted work, including your own, with four questions borrowed from four workshops:

Run the four questions on your last week of AI work. Most people find they live on Level 1 with decorations. The climb itself is ordinary: one checklist that changes behavior, one written standard that leaves your head, one constraint that closes a real failure class, one signature you actually earn.

What AI adds to the ladder is a change in economics, not structure. Skill-looking output is now rentable by anyone; the artifacts of Levels 2 and 3 can be drafted by the machine itself; what cannot be rented is everything this essay traced to ordinary work — the acceptance authority, the taste, the signature, the seat that decides. The technologies before this one moved muscle off the critical path; this one moves craft off it. The ladder is older than all of us. For the first time, the top of it is the whole job.

10. Questions people actually ask

Is this just project management with new branding? Project management is one ancestor, and the systems-science behind it (Meadows’ leverage points, from tuning parameters up to transcending paradigms) is fifty years older [37]. What is genuinely new is the assignment: the execution layer itself is now machine-rentable, so every layer above it is exposed for what it always was. The ladder is old; its AI-era payoff is not.

I just use ChatGPT in a browser. Does any of this apply? All of it, at smaller scale. A process is a preflight note and an acceptance checklist in a doc. Methodology is a file of your standards you paste at session start. Architecture is deciding what you never paste in and what you always verify yourself. Philosophy is remembering that the model’s confident answer is not a signature. None of the levels require code.

Where do I start if I’m at Level 1? Pick one recurring task. Before you do it again, write down: what inputs it needs, what a good outcome must contain, what went wrong last time. That sheet is your Model 299 moment — the checklist that made “too much airplane” flyable. Run it twice. You are now on Level 2, and it took you one page.

Won’t better models make these levels obsolete? They keep dissolving Level 1 from underneath — that is real, and it is why prompt collections age so fast. But the levels above are not about the model’s limits; they are about the structure of capable work: verification, transfer, constraint, accountability. Ontario’s hospitals did not fail to benefit from checklists because their surgeons lacked skill; mandating the card had not changed the practice. Better models raise the ceiling of Level 1 — and the stakes of every level above it.

Sources

  1. Addy Osmani, “The 70% problem: Hard truths about AI-assisted coding”, 2025.
  2. Dave Kindy, “On. Set. Checked.”, Air & Space Quarterly (Smithsonian National Air and Space Museum), Winter 2023.
  3. Atul Gawande, “The Checklist”, The New Yorker, December 10, 2007.
  4. Atul Gawande, The Checklist Manifesto: How to Get Things Right, Metropolitan Books, 2009.
  5. FAA, Pilot’s Handbook of Aeronautical Knowledge (FAA-H-8083-25C, 2023), Chapter 2: Aeronautical Decision-Making, IMSAFE checklist; see also AOPA preflight resources.
  6. U.S. Army, Training Circular TC 25-20, A Leader’s Guide to After-Action Reviews, September 30, 1993; business adaptation: Darling, Parry & Moore, “Learning in the Thick of It”, Harvard Business Review, July–August 2005.
  7. Haynes AB et al., “A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population”, N Engl J Med 360:491-499, January 29, 2009.
  8. Pronovost P et al., “An Intervention to Decrease Catheter-Related Bloodstream Infections in the ICU”, N Engl J Med 355:2725-2732, December 28, 2006.
  9. Anthropic, “Claude Code: Best practices for agentic coding”, April 2025, continuously maintained.
  10. Den Delimarsky, GitHub Blog, “Spec-driven development with AI: Get started with a new open source toolkit”, September 2, 2025.
  11. Urbach DR et al., “Introduction of Surgical Safety Checklists in Ontario, Canada”, N Engl J Med 370:1029-1038, March 20, 2014.
  12. A. Degani & E. Wiener, “Human Factors of Flight-Deck Checklists: The Normal Checklist”, NASA Contractor Report 177549, May 1991.
  13. The Culinary Institute of America, The Professional Chef, 9th ed., Wiley, 2011, Chapter 4: “Mise en Place”; definition per Merriam-Webster.
  14. Mike Rother, Toyota Kata: Managing People for Improvement, Adaptiveness and Superior Results, McGraw-Hill, 2009; the improvement and coaching kata.
  15. Masaaki Imai, Kaizen: The Key to Japan’s Competitive Success, McGraw-Hill, 1986 — reporting Eiji Toyoda’s figure of 1.5 million suggestions per year, 95% implemented.
  16. Michael Polanyi, The Tacit Dimension, Doubleday, 1966, p. 4: “we can know more than we can tell.”
  17. Jean Lave & Étienne Wenger, Situated Learning: Legitimate Peripheral Participation, Cambridge University Press, 1991.
  18. agents.md — the open agent-instructions standard, stewarded under the Linux Foundation.
  19. Letta, “Memory Blocks: The Key to Agentic Context Management”, May 14, 2025.
  20. Prateek Chhikara et al., “Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory”, arXiv:2504.19413, April 2025.
  21. David von Drehle, “The Triangle Shirtwaist Factory Fire”, OSHA; see also New York State DOL.
  22. NFPA Journal, “Fitting Tribute”, November 21, 2023; NIST, Bukowski, “The Basis for Egress Provisions in U.S. Building Codes”, 2004.
  23. Online Ethics Center, “Hyatt Regency Walkway Collapse”, University of Virginia; Missouri Board finding, November 1984.
  24. Karl E. Weick & Kathleen M. Sutcliffe, Managing the Unexpected, Jossey-Bass, 2001 (3rd ed. 2015) — the five HRO principles.
  25. Simon Willison, “Context engineering”, June 27, 2025 — chronicling Tobi Lütke’s and Andrej Karpathy’s June 2025 formulations.
  26. Anthropic, “Effective context engineering for AI agents”, September 29, 2025.
  27. Walden Yan, Cognition, “Don’t Build Multi-Agents”, June 12, 2025.
  28. Anthropic, “How we built our multi-agent research system”, June 13, 2025.
  29. Erik Schluntz & Barry Zhang, Anthropic, “Building effective agents”, December 19, 2024.
  30. NCEES Model Law (2021) and NSPE Position Statement No. 10-1778, “Responsible Charge”, revised May 2024.
  31. ACGME Common Program Requirements, Section VI: supervision — “the attending physician is ultimately responsible for the care of the patient.”
  32. Parasuraman R, Molloy R, Singh IL, “Performance Consequences of Automation-Induced ‘Complacency’”, The International Journal of Aviation Psychology 3(1):1-23, 1993.
  33. Parasuraman R & Riley V, “Humans and Automation: Use, Misuse, Disuse, Abuse”, Human Factors 39(2):230-253, 1997.
  34. Herbert A. Simon, “Designing Organizations for an Information-Rich World”, in Computers, Communications, and the Public Interest, Johns Hopkins Press, 1971, pp. 37-72.
  35. OpenAI, “A practical guide to building agents”, April 2025.
  36. OpenAI, “GPT-5 prompting guide”, August 2025 — persistence alone moved Tau-bench Retail from 73.9% to 78.2%; the structure around the model does the work.
  37. Donella Meadows, “Leverage Points: Places to Intervene in a System”, The Sustainability Institute, 1999.
  38. Stuart Dreyfus & Hubert Dreyfus, A Five-Stage Model of the Mental Activities Involved in Directed Skill Acquisition, US Air Force Office of Scientific Research, 1980; elaborated in Mind Over Machine, 1986.

For the engineering layer under this essay: what agents are made of (not a feature), why the loop matters more than the tools (the loop), and a full reference architecture for building one (from scripts to systems).

Next
What Is Artificial General Intelligence, Actually?