Evidence in action

Should firms reorganize around AI?

A diagnostic walk through the Galbraith Star Model, stress-tested against three recent findings: P&G's field experiment, human-LLM-OR inventory complementarity, and Coase versus Claude.

The question

The question isn't "should we use AI." It's "should the org chart change."

Most AI debates argue about tools. This one argues about structure. If an AI-augmented individual can match a two-person team, if the best decisions come from humans and machines together, and if AI collapses the coordination costs that justified the firm in the first place, then the live question isn't adoption anymore. It's reorganization.

How to read this page: three findings supply the evidence, the Coasean lens supplies the theory, and the Galbraith Star supplies the diagnostic. The verdict at the end is deliberately balanced: reorganize the parts the evidence touches, and resist the urge to reorganize everything.

The three readings

What the evidence actually says

Two recent empirical studies and one theoretical essay. Skim the abstracts and results; the numbers below are pulled straight from the papers.

Field experiment

The Cybernetic Teammate

Dell'Acqua, Lakhani et al. — 776 P&G product innovators, pre-registered

On the product-innovation tasks studied, individuals with AI matched the solution quality of teams without AI. Proposals also became more balanced across technical and commercial perspectives. This does not establish a headcount equivalent for a whole role or company.

NBER w33641 →
Benchmark + experiment

AI Agents for Inventory Control

Baek et al. — InventoryBench, 1,000+ inventory instances

Hybrid methods improved benchmark performance. A separate classroom experiment found that humans making decisions with OR and LLM recommendations outperformed the compared modes on average. These task-specific results do not demonstrate company-wide staffing reductions.

arXiv 2602.12631 →
Theory

Coase vs. Claude

Howard Yu — the theoretical anchor

Coase asked in 1937 why firms exist at all: because some activities are cheaper to internalize than to buy on the market. Yu argues that when AI drives coordination, search, and monitoring costs toward zero, that boundary moves, and the firm starts to look like a platform of micro-enterprises.

Read the essay →
The data

Two findings, in numbers

The headline results, charted from the published point estimates. Both have the same shape: AI changes who can do the work, and with whom. Speed is almost the least interesting part.

AI lifts the individual to team level

Cybernetic Teammate — solution quality vs. a solo worker without AI

0 0.1 0.2 0.3 0.4 +0.37 SD +0.39 SD Solo worker + AI Two-person team + AI Baseline (0) = a solo worker without AI
Teams with AI were 3× more likely to produce a top-tier solution.

Point estimates: +0.37 SD (individuals) and +0.39 SD (teams), both p<0.01. The study reports comparable solution quality for individuals with AI and teams without AI in this task; that comparison is not a whole-role productivity multiple.

Complementary tools improved this inventory benchmark

InventoryBench — mean normalized reward, Gemini 3 Flash

0 0.2 0.4 0.6 0.445 0.538 +21% OR algorithm alone OR → LLM pipeline
In the classroom experiment, human + AI teams beat humans or AI alone on profit.

Table 1(a): OR 0.445 (95% CI ±0.018); OR-to-LLM 0.538 (±0.016), across 1,320 benchmark instances. The human result comes from a separate classroom experiment. Source table and method.

Both studies point the same way: AI is most powerful as a teammate. Take that seriously and it changes how teams get sized and staffed, not merely how fast they move.
The theoretical lens

Coase asked why firms exist. AI rewrites the answer.

Firms exist, Coase argued, because internalizing some activities is cheaper than buying them on the market. Yu's claim is that AI collapses the very costs (coordination, search, monitoring) that made internalizing cheaper. As that boundary moves, some work flows back out to the market. And here's the twist: a different set of activities becomes more compelling to keep inside, not less.

Push back to the market

Where AI drives transaction costs below the cost of internal management.

  • Routine coordination and scheduling
  • Supplier sourcing and contract drafting
  • Quality monitoring and status reporting
  • Modular work units an agent can orchestrate
AI collapses coordination, search & monitoring costs

Pull deeper inside

Where value rises precisely because everything else got cheap and noisy.

  • Brand identity and trust as a filter
  • Capital to absorb failed experiments at scale
  • Platform accountability and the standard
  • High-stakes human judgment

Yu's examples push the point: Shein atomized design into micro-decisions where Zara held an integrated collection; Figma made the element, not the file, the unit of work; Haier reorganized 80,000+ employees into 4,000+ self-managing micro-enterprises under one platform brand. The firm doesn't vanish. It becomes a platform orchestrating a network.

The diagnostic

Walking the Galbraith Star

Galbraith's model says an organization is healthy only when five points stay aligned: Strategy, Structure, Processes, Rewards, and People. Change one and the others must rebalance, or the system misfires. AI hits some points hard, and companies routinely forget to re-tune the others. I've watched it happen more than once.

Strategy

AI pressure

What becomes possible when a solo operator has team-level output? New service levels, faster cycles, smaller bets.

Structure

AI pressure

Smaller teams, fewer hand-offs, blurred functional silos. The "two-person team" finding lands hardest here.

Processes

AI pressure

Human + LLM + OR hand-offs, review gates, and the determinism bridge replace single-owner workflows.

Rewards

Often un-rebalanced

Still pay for individual heroics and headcount, and AI leverage stalls. The point companies most often forget.

People

Often un-rebalanced

Roles shift from doing to directing and reviewing — and one new seat appears: the meaning steward, accountable for keeping institutional definitions current as the business changes.

For the room

Four questions to pre-think

Bring one or two. These anchor the small-group discussion; each one turns a finding into a decision about your own firm.

From Cybernetic Teammate

Smaller teams, fewer specialists, or just better tooling?

AI-augmented individuals matched two-person teams, and R&D and Commercial converged on similar-quality solutions. If that holds in your firm, does it argue for smaller teams, fewer functional specialists, or simply better tooling on top of today's structure? And where does the finding break down?

From Baek

Who is the third teammate you don't yet have?

The best inventory team was human + LLM + OR, not any one alone. Pick a function in your company: who's the third member of that team you haven't hired yet? What changes about hiring, training, and reporting if you take the finding seriously?

From Coase vs. Claude

What do you push out, and what do you pull in?

If AI collapses coordination, search, and monitoring costs toward zero, which activities would you push back out to the market? And which become more compelling to keep inside, not less?

Walking the Star

Which point gets hit, and which gets forgotten?

Across Strategy, Structure, Processes, Rewards, and People: which point does AI hit hardest in your industry? Which one gets left un-rebalanced, so the rest of the system misfires?

The verdict

So, should firms reorganize around AI?

Test workflow changes first. The studies support task-specific hypotheses about collaboration; they do not prescribe a new organization chart. Structural change needs sustained evidence from your own work, quality, coordination costs, and employee outcomes.

Test the workflow

Map one bounded job with its domain translator. Trial different human and AI handoffs in parallel, with a named operator and consistent quality gates. Compare net effort and coordination before changing team structure.

Re-tune first

Rewards and People, the points companies forget. Pay for leverage and judgment, not headcount and heroics; retrain roles from doing to directing before the structure changes.

Hold the line

Keep accountability for brand, capital, and consequential decisions clear. Test whether routine coordination improves with agents or external providers; include oversight and failure costs before changing ownership.

The Day One Leader read: the same discipline applies here as everywhere in the guide. Start with the work, not the org chart. Pick the point on the Star where the evidence is strongest, change it deliberately, re-tune the points around it, and then measure whether the system got healthier. Faster alone doesn't count.

From research to a test in your business
EvidenceLimitYour test
Product-innovation field experimentOne task setting, not a staffing study across all functions.Compare approved deliverables, quality, and total effort on the same job before changing roles.
Inventory benchmark and classroom experimentBenchmark rewards and classroom decisions are not measured company-wide operating savings.Trial a supervised inventory workflow on historical data and track overrides, cost, and service outcomes.
Transaction-cost theoryA lens for hypotheses; it does not establish that coordination costs disappear.Measure internal and external coordination, review, and incident costs before moving the boundary.
Sources

Read the originals