Workforce Equivalency

Measuring a workforce of one

Everyone with an AI subscription now claims a multiple. "I'm 10× more productive." Ten times at what, measured how, against whom? This page is my attempt to answer that question honestly for one person: me. I call the result Workforce Equivalency (WFE) — how many of yesterday's me today's me equals — and I publish the formula, the numbers, and the parts that don't flatter.

1.71×My fixed-bundle WFE estimate
2.57×Same-hours output index, a different question
3.2×Fixed-composition model ceiling
1.11×My judgment estimate, awaiting measurement
The honest setup

Two arbitrary choices, admitted up front

Every productivity multiple hides two decisions, and most people quoting one can't tell you either. First, the baseline: multiplied compared to what? Mine is pre-AI me doing the same work, not a market-average employee. That makes WFE personal and checkable instead of a hiring claim. Second, the unit of output: each dimension of the work needs its own definition of "done," and each definition needs a quality gate, or the multiple rewards volume over value. A 7× documentation rate means nothing if nobody uses the documents.

The test for any multiple: if someone quotes you a productivity number and can't name their baseline and their unit of output, they're quoting a feeling. The feeling might even be right. But you can't manage a feeling, and you can't put one in front of a board.

The framework

Five dimensions, five multipliers

I provisionally split my pre-AI working week into five kinds of output, then hold that baseline role mix fixed. Each multiplier estimates today's output rate against the pre-AI rate for the same dimension. Both the shares and multipliers are working estimates, not yet a complete instrumented baseline — the protocol at the bottom of this page is how estimates become measurements. Judgment is still the constraining estimate, but 1.11× is my prior, not a result supplied by the literature.

Dimension What counts as output Multiplier Estimated pre-AI share Quality gate
Judgment & decisions Decisions made; cycle time from raised to resolved 1.11× 35% The decision still holds at a 30-day review
Systems: build & deploy Working systems shipped to production 3.0× 20% Running and used, not demoed
Documentation Docs, manuals, plans, and specs shipped 7.0× 15% Someone other than me actually uses the page
Communications Substantive emails, posts, briefs, meetings closed 1.4× 20% Bounded by the human on the other end. First reconstruction measured 2.0× on the personal channel (upper bound) — 1.4× kept as the conservative model input
Research & analysis Questions answered with sourced, verified findings 2.5× 10% Claims traced to sources, not summaries of vibes

Multipliers compare today's rate with the pre-AI rate for the same output. The 35/20/15/20/10 shares are provisional baseline estimates and must be replaced with reconstructed or logged pre-AI shares before this becomes a publishable WFE result. Once established, those shares stay frozen; current allocation is tracked separately.

Reality check

What the published research says

My estimates run above what controlled studies have measured, and you should know that before trusting any of this. The studies below are dated snapshots in specific settings, not validation of my multipliers.

Dimension My estimate What controlled studies found Source
Systems 3.0× Cuts both ways. Developers finished one scoped coding task 55.8% faster with an AI assistant. In METR's early-2025 historical trial, experienced open-source developers on familiar repositories were 19% slower; METR's 2026 follow-up says newer data weakly point toward speedup but cannot support a reliable current estimate. Peng et al. 2023; METR 2025; METR 2026 update
Documentation 7.0× Professional writing tasks took 40% less time with ChatGPT, with quality up 18% — roughly a 1.7× rate, nowhere near 7×. Noy & Zhang, Science 2023
Communications 1.4× Customer-support agents resolved 15% more issues per hour on average; the biggest gains went to novices, and top performers saw small quality declines. This supports strong task and skill heterogeneity, but it does not validate my 1.4× communications estimate. Brynjolfsson, Li & Raymond, QJE 2025
Research & analysis 2.5× 758 BCG consultants inside AI's capability frontier: 12.2% more tasks, 25% faster, 40% higher quality. Outside the frontier, AI users did worse than the control group. Dell'Acqua et al. 2023
Judgment 1.11× Little direct evidence establishes a general multiplier for decision quality. The strongest adjacent finding here is that individuals with AI matched teams without AI on product-innovation performance — a capacity result, not evidence for precisely 1.11× judgment. That number remains my personal prior until measured. Dell'Acqua et al., NBER w33641

Why my estimates run hotter than the studies

The gap has a structural explanation, not just optimism. The RCTs measure assistance: one human doing one assigned task with AI help. My estimates describe delegation: agents doing whole units of work in parallel while I review — and I choose which work goes to them, which keeps me inside the capability frontier the BCG study mapped. Delegation plus selection can honestly produce multiples that assistance studies never show. It can also fool me. Both are true, and only measurement separates them.

The strongest argument for measuring at all: in METR's early-2025 randomized trial, experienced developers believed AI made them 20% faster after it had measurably made them 19% slower in that setting. METR no longer treats that result as current, but the perception warning survives: every self-reported multiplier on this page, including mine, stays an estimate until a log exists.

The formulas

One set of inputs, two different questions

Feed the same five multipliers and frozen baseline role shares into two formulas and you get two numbers. They are not an optimistic and conservative version of the same claim. The harmonic calculation estimates how quickly today's system could reproduce a fixed pre-AI bundle of work. The arithmetic calculation indexes the output produced if the old time allocation stays in place, which changes the mix toward the fastest dimensions.

Reading the formulas: wi is the dimension's share of the frozen pre-AI baseline week — judgment was 35% of mine, so w = .35. mi is today's output rate divided by the pre-AI rate for the same dimension. The Σ means "add this up across all five dimensions." The shares must come from the same baseline period; current allocation is a separate diagnostic.

Same-hours output index

Arithmetic index

2.57×
Index = Σ (wₖ × mₖ)
= .35(1.11) + .20(3.0) + .15(7.0) + .20(1.4) + .10(2.5)

The question it answers: if I keep allocating the same hours to each dimension, what is the weighted index of the resulting output rates?

The hidden assumptions: unlike outputs can be added meaningfully, their value weights are equal, and demand for the fastest output remains available. That makes this a directional index, not an FTE claim.

Fixed-bundle WFE estimate

Harmonic throughput

1.71×
WFE = 1 ÷ Σ (wₖ ÷ mₖ)
= 1 ÷ (.35/1.11 + .20/3.0 + .15/7.0 + .20/1.4 + .10/2.5)

The question it answers: how quickly could today's system reproduce one fixed bundle of the pre-AI role's outputs?

Why it is the WFE candidate: it converts each baseline share into the time now required to reproduce it, then takes the reciprocal. It is only defensible while the baseline bundle and quality gates stay fixed.

What I publish: 1.71× as the fixed-bundle WFE estimate, beside a 2.57× same-hours output index. I do not present them as a confidence range because they answer different questions. The gap still diagnoses how much the output mix changes when some dimensions accelerate much faster than others.

How to run it on any role — five steps

  1. Split the baseline role into output dimensions.Not activities — outputs. What actually left your desk before the AI-assisted period: decisions, systems, documents, messages, findings. The frozen baseline shares must sum to 100%.
  2. Reconstruct the baseline with AI, then start the clock.Point an agent at the archives that already keep history — git, sent mail, published posts, your docs store — and have it compute your weekly output rate before AI. Where no archive exists (judgment, usually), the clock starts today: log forward and let the first eight weeks become the reference.
  3. Set a multiplier per dimension — measured beats estimated.Multiplier = current weekly rate ÷ baseline rate, same unit, same quality gate. Where you can't measure yet, estimate — but label it, state the bias direction, and let the log embarrass it later. Mine got beaten within a day (communications: estimated 1.4×, measured 2.0×).
  4. Compute both diagnostics, but do not call them a range.The harmonic (1 ÷ Σ wᵢ/mᵢ) estimates fixed-bundle throughput. The arithmetic index (Σ wᵢ×mᵢ) shows the changed output mix under the old time allocation. Publish the assumptions beside both.
  5. Read the constraints.The gap shows how unevenly leverage is distributed. The fixed-composition ceiling (judgment multiplier ÷ baseline judgment share) shows the limit of this model if other dimensions approach infinite speed. The lowest measured multiplier suggests where to investigate next; it is not automatically the highest-value intervention.

Same math, three role shapes

The formulas respond differently to the frozen baseline role mix — that's the RLAI made concrete. Same method, three illustrative role shapes:

Role shapeBaseline shares & estimated multipliersFixed-bundle WFEOutput indexModel ceiling
Fractional CMO
content-heavy
Judgment 25% @ 1.1× · Content/docs 35% @ 5× · Comms 25% @ 1.5× · Research 15% @ 2.5× 1.91×2.78×4.4×
Solo consultant
judgment-heavy
Judgment 50% @ 1.1× · Research 20% @ 2.5× · Docs 15% @ 4× · Comms 15% @ 1.4× 1.47×1.86×2.2×
Operator / founder (me)
build-heavy
Judgment 35% @ 1.11× · Systems 20% @ 3× · Docs 15% @ 7× · Comms 20% @ 1.4× · Research 10% @ 2.5× 1.71×2.57×3.2×

These are illustrative scenarios, not occupational benchmarks. The baseline role mix materially changes the model result, which is why role shares must be measured and frozen before comparing periods or people.

The role-level adjustment

Why judgment-heavy roles score lower in this model

This illustrative model assumes senior roles allocate a larger baseline share to judgment and that judgment has the lowest multiplier. Hold all multipliers constant and increase that share, and fixed-bundle WFE falls. That is a sensitivity test, not evidence that every CEO scores lower than every analyst; real role shares and output gates must be measured.

Role level Assumed baseline judgment share WFE (harmonic) Model ceiling
IC / specialist15%2.04×7.4×
Manager25%1.86×4.4×
Operator / founder (me)35%1.71×3.2×
Executive45%1.58×2.5×
CEO55%1.46×2.0×

The fixed-composition ceiling: divide the judgment multiplier by its frozen baseline share. For this model: 1.11 ÷ 0.35 ≈ 3.2×. That is the harmonic estimate's mathematical limit if every other multiplier approached infinity while the baseline bundle and 1.11 judgment estimate stayed unchanged. It is not a hard ceiling on the person, role, or business.

The picture

The shape, and the trend

The radar shows the estimated multipliers used in this model. The trend is an illustrative scenario, not a measured time series; its toolset markers show a hypothesis to investigate rather than demonstrated causes of improvement.

My WFE by dimension

Multiplier vs. pre-AI baseline · radial scale is log₂ (rings at 1×, 2×, 4×, 8×)

The radial scale is logarithmic on purpose: on a linear radar, 1.11× is invisible next to 7×. The near-center judgment vertex isn't a rendering problem. It's the finding.

Composite WFE over time

Harmonic WFE, monthly · gold markers are toolset changes · illustrative

This series is a sketch of the idea, not logged data. An informal archive review suggested a similar shape: git activity, filtered sent-mail counts, and LinkedIn's cumulative-impressions curve appeared to show a sharp rate change around early 2026, near the toolset change. That is a hypothesis to test with logged data, not proof of causation.

The real project

Raising the 1.11

"Judgment only improved 11%" reads like a limitation of AI. I think it's an instruction. The 11% isn't fixed; it's a measure of how much of my judgment I've managed to write down. Every judgment concept I encode into the AI harness — a skill with my review rules in it, a memory of why past decisions were made, a gate that knows which calls I always make the same way — moves a slice of work out of the 1.11× bucket and into the 3× bucket. The decision still gets made my way. I just stop being the bottleneck for the ones I've already made a hundred times.

WFE as a diagnostic: the lowest multiplier on the radar is not a grade. It's the to-do list. Mine says the next unit of effort shouldn't go into a faster documentation pipeline (7× is plenty); it should go into encoding judgment. The score tells you where to work. That's worth more than the score itself.

Thought experiment

The two-person unicorn

Everything above estimates one person. The next step is a hypothesis, not a research result: complementary senior operators may capture more value from agents when they can identify which work sits inside the capability frontier, delegate it cleanly, and keep judgment at the human boundary. Other field evidence shows some of the largest gains going to novices, so experience alone is not the mechanism. The pairing worth testing is an engineer with business acumen and a domain translator who converts industry reality into specifications, trust, and distribution.

Inside an existing business, this pairing also needs a named operator who develops, measures, and supervises the agent after launch. Domain translators identify that owner; builders lead major process releases; operators maintain and improve the workflow, and retire it when the work changes. See the process mapping, release ownership, and agent lifecycle model.

The math predicts the pairing

A complementary partner could change the role mix: the engineer sheds some translation and distribution work; the translator sheds some systems and deployment work. That may raise each person's fixed-bundle throughput, but the effect must be recalculated from new baseline shares rather than inferred from the gap between the harmonic and arithmetic diagnostics.

The second hypothesis is lower coordination overhead, not zero coordination cost. A well-matched pair running governed agents could field more output with fewer hand-offs than a larger conventional team. The Coase argument gives the theoretical direction; it does not establish a specific headcount equivalent or prove that two people replace a particular org chart.

The testable scale question: can a complementary pair produce and sustain the quality-gated output of a materially larger team without transferring the saved delivery time into review, sales, support, and incident-response work? Until output, quality, revenue, and coordination time are logged together, the two-person unicorn remains a useful hypothesis rather than a base case.

The counterweights. First, the judgment pool is still two humans deep; any fixed-composition estimate remains sensitive to their measured judgment share and multiplier. Second, trust and enterprise distribution still buy human surface area; some markets demand faces in rooms. Third, success can turn the pair into executives of an agent fleet: review and exception work may expand as delivery accelerates. A pair that wins at scale must measure whether encoded judgment genuinely reduces the decision queue or merely moves it.

The open question is how many of these pairs already exist — quietly running at multi-team output, eighteen months into the toolset era, too small for anyone to benchmark. You won't see them in headcount statistics. You'll see them when revenue-per-employee numbers start looking like typos.

Door one: start something — picking your complement

If the pairing thesis holds, the practical question is which builder pairs with which translator for which kind of company. The pattern: the translator supplies the frontier map — which work is automatable and worth paying for — and the builder ships it at systems speed. Neither alone has a company.

Venture shapeThe builder bringsThe translator bringsWhy it compounds
Vertical AI SaaS Product engineer who can read a P&L Industry operator who knows which five workflows actually hurt The translator's frontier map is the moat — competitors can copy features, not workflow judgment. Builder ships at the 3× systems multiplier.
AI-native services firm Automation engineer Trusted advisor with a client book Services economics invert: delivery runs at agent cost while trust wins the work. The 7× documentation multiplier turns into billable deliverables.
Audience business Pipeline builder A voice with a point of view and distribution The engine multiplies the voice rather than replacing it — my own reach rate moved ~5× when the content pipeline went live. Reach compounds; production cost doesn't.
Acquisition vehicle Systems integrator Operator who can diligence a company's actual week The second door, below — buy the multiplier gap instead of building the product.

Door two: buy something — the optimization map

The same lens can structure diligence on a business you acquire rather than start. A target's roles can be decomposed into baseline shares, measurable outputs, quality gates, and candidate AI multipliers. That does not make every 1.0× dimension realizable upside; it identifies a workload to test before underwriting any return.

Optimize forWhere AI lands firstDimension it exploitsThe judgment the owner must keep
Margin Back-office automation — invoicing, scheduling, reporting; AI-prepped vendor renegotiation Systems (3×) + documentation (7×) Which costs are fat and which are muscle. Cut the wrong one and quality pays the bill two quarters later.
Quality Consistency gates, checklists graduated into skills, review loops, process finally written down Documentation + the judgment loop Defining the quality bar itself. A gate can't enforce a standard nobody wrote down — encoding "good" is owner work.
Output / performance Agent-augmented staff (the support-agent study's +15% is the floor, and it lands biggest on the least-experienced), hand-off and pipeline automation Systems + communications Which hand-offs keep a human in them. Speed that skips review isn't throughput, it's future rework.
Revenue & growth Lead intake, missed-call rescue, proposal turnaround, the content engine, pricing analysis — the workload library catalogs these plays Content + communications + research Pricing and promises. AI drafts the quote; the owner decides what the firm commits to. That signature is pure wjudgment.

Diligence, reframed: walk the target's org chart with the WFE lens, then validate candidate workloads against real baselines, quality gates, implementation cost, adoption, and retained human judgment. A 1.0× estimate is a question, not booked upside.

Behind both doors: Build. Borrow. Buy. Bot.

Whether you start the company or buy it, every capability inside it now gets sourced through a four-way decision. Sarah's workforce-strategy piece, Build. Borrow. Buy. Bot., extends the classic talent question — develop it, contract it, or acquire it — with the fourth option this whole page measures: delegate it to a governed agent.

Her sharpest observation is the operating-model shift hiding inside that fourth option: the work stops being "someone must do it" and becomes "someone must design, govern, and improve the workflow — a fundamentally different operating model." That someone is the owner, and that work is exactly the judgment-encoding this page keeps pointing at. Her bot-vs-human sort matches the multiplier map too: agents take the high-volume, rules-based, hand-off-delayed work — the 3× and 7× dimensions — while ambiguity, relationships, and accountability stay human, which is the 1.11× dimension wearing its workforce-planning clothes.

The governance echo: her test for a bot hire — what job is it performing, what authority does it have, who supervises it, how is quality measured, what happens when it's wrong — is the same checklist as the visibility layer below, arriving from the workforce-planning chair instead of the measurement chair. As she puts it: "The goal isn't replacing people. The goal is elevating people." One household, two instruments: she plans the workforce; I count it.

Your turn

Run your own number

Pick the illustrative role shape closest to yours, then replace its shares with a measured pre-AI baseline week and freeze them. Estimate each multiplier against pre-AI you: quality-gated units of finished output per week now, divided by the same units per week before AI. Current time allocation belongs in a separate comparison, not in the WFE weights.

Inputs

Multiplier hints — judgment: decisions resolved per week that hold at 30 days. Systems: working systems shipped. Documentation: pages people use. Communications: substantive threads closed. Research: sourced briefs delivered.

Fixed-bundle WFE estimate
1.71×
Same-hours output index: 2.57×

Fixed-composition model ceiling: 3.2×. This is a mathematical limit under the frozen shares and estimated multipliers, not a ceiling on you or the role.
The measurement protocol

How the numbers would get made

A multiple you can't audit is marketing, and right now mine is exactly that — estimates with a protocol attached. The protocol is one baseline, four counters, and a ten-minute weekly log. Three dimensions pull from systems that already keep history; the two that involve judgment get written down by hand, because judgment can't be scraped. When the log starts, the estimates above get replaced by whatever it says.

Instrumented — pulled automatically

  • Systems: production deploys and accepted shipped changes per week. Commit volume is supporting activity, not a quality-gated outcome.
  • Documentation: pages, specs, and manuals shipped per week, from the git history of the sites and the knowledge vault.
  • Communications: substantive email threads and posts per week.
  • Research: deep-research briefs delivered, from session history.

Logged — weekly, by hand

  • Judgment: decisions made this week and their cycle time, each revisited at 30 days. The hold-rate is the quality proxy: a fast decision that gets reversed counts against the multiplier, not for it.
  • Current time mix: a rough percentage split of this week across the five dimensions. Compare it with the frozen pre-AI shares; do not replace the formula weights with it.
  • Toolset changes: date-stamped. These become the gold markers on the trend chart, so every step-up in the line has a named cause.

Status: the ledger is live and the first baselines are reconstructed. The estimates on this page hold their place until gated multipliers exist.

First real measurements (July 2026): reconstructing pre-AI baselines from personal archives produced three findings. Filtered sent-mail ran 3.9 threads/week in late 2024 vs 7.8/week now — a measured ~2.0× on communications (an upper bound; it beat the 1.4× estimate above, which stays as the conservative input). LinkedIn's cumulative-impressions curve shows a clean knee where the AI content pipeline went live: ~3.4K/month before, ~17K/month after — roughly 5× the reach rate. And one structural finding: pre-AI professional output lived inside employer systems and is unrecoverable — for the building dimensions, the pre-AI denominator isn't small, it's invisible. The multiplier there isn't a number; it's a category change.

Ongoing governance

The visibility layer: measurement that runs itself

A weekly log works, but it relies on remembering. The stronger design aligns the AI work itself with the dimensions it serves, so attribution happens at dispatch, not at recall. Two moves make that possible, and neither requires your tools to cooperate.

Move one: tag every AI job with the dimension its output serves. An agent drafting a spec is a documentation job. A pipeline writing LinkedIn posts is a content job. A workflow shipping a deploy is a systems job. Tag the job when you dispatch it and the measurement question inverts — instead of "what did I produce this week?" the ledger already knows, and the governance question gets sharper too: you can see exactly which dimension every agent-hour and every AI dollar is serving, and whether that matches where your ceiling says effort should go.

Move two: intercept at the base URL. Nearly every AI tool honors a configurable API endpoint (ANTHROPIC_BASE_URL, OPENAI_BASE_URL, a gateway setting). Point them all at one logging gateway and every AI call in every tool crosses a single choke point that records model, tokens, cost, caller, and the job's dimension tag. One config change per tool buys total visibility — spend per dimension, calls per job, cost per output — without asking any vendor for an integration.

Dispatch

AI jobs, tagged

Every agent, pipeline, and workflow carries the dimension its output serves — set at dispatch, not reconstructed later.

Choke point

Base-URL gateway

All tools point at one logging proxy: model, tokens, cost, caller, dimension. One config line per tool.

Edges

Output counters

Git hooks, deploy webhooks, publish pipelines, mail filters — each shipped artifact appends a claim with evidence attached.

Ledger

Append-only claims

Events land in one inbox, never edited. Automation collects; it never scores.

Human gate

Weekly review

Quality gates promote claims to counts; multipliers and WFE fall out of arithmetic. The human stays the scorer.

The division of labor is the whole design: machines collect, humans score. The gateway and the counters make the record complete; the weekly gate keeps the number honest. And the same pipes serve double duty — the ledger that computes your WFE is also the audit trail your AI governance needs: what ran, for which purpose, at what cost, producing what. One instrument, two dashboards.

Take the protocol with you

One prompt sets it up

If you work with an AI coding agent — Claude Code, or any harness that can read your git history and keep files — the whole protocol above is one paste away. Copy the prompt, give it to your agent, and it will interview you, pull your pre-AI baseline from systems that already keep history, and stand up the weekly log. It's written to be a skeptical auditor, not a cheerleader.

WFE measurement harness · system prompt
You are my Workforce Equivalency (WFE) measurement harness. Your job is to set up and run a system that measures how many pre-AI versions of me my current AI-assisted output equals. Be a skeptical auditor, not a cheerleader: the goal is a defensible number, not a flattering one.

## Step 1 — Interview me (one question at a time)
1. My role and seniority: IC / manager / operator-founder / executive / CEO.
2. The dimensions of my output. Default five, adjust to my actual work:
   judgment & decisions; systems built & deployed; documentation;
   communications; research & analysis.
3. My pre-AI baseline time share per dimension (they must sum to 100% and
   stay frozen for comparable WFE reporting).
4. The baseline boundary: the date I started working with AI seriously.
5. Which systems hold my work history: git repos, deploy platform, email,
   docs/wiki or knowledge vault, task tracker, calendar.

## Step 2 — Establish the pre-AI baseline
For every dimension with system history, compute my weekly output rate for
the 8-12 weeks BEFORE the baseline date: commits and production deploys per
week, documents shipped per week, substantive messages per week, research
briefs per week. Store the method next to each number so it can be
recomputed and challenged. Where no history exists (judgment, usually),
record "no baseline — log-forward only" instead of inventing one.

## Step 3 — Create the measurement store
Create wfe/config.json (dimensions, frozen baseline shares, baseline rates,
quality gates) and wfe/events.jsonl (append-only, one evidence-backed claim
per shipped outcome). Quality gates are mandatory
and belong in the config, not in my head:
- Documentation counts only if someone other than me actually uses it.
- Systems count only when running in production, not demoed.
- Decisions count only if they still hold at a 30-day review.
- Research counts only with sources attached.

## Step 4 — The weekly review (run when I say "wfe weekly")
1. Pull this week's counters automatically wherever systems allow.
2. Ask me for what can't be scraped: decisions made and their cycle time,
   30-day verdicts on past decisions, this week's actual time mix, and any
   toolset changes (date-stamped — these annotate the trend later).
3. Append evidence-backed outcome claims to wfe/events.jsonl. Keep raw
   activity such as commit counts separate from audited outcomes.
4. Recompute each multiplier: current weekly rate divided by baseline rate.
5. Report three diagnostics every week:
   - Fixed-bundle WFE = 1 / sum(baseline_share_i / multiplier_i)
   - Same-hours output index = sum(baseline_share_i * multiplier_i)
   - Fixed-composition model ceiling = judgment multiplier /
     baseline judgment share
   Plus the trend of the harmonic number over time.

## Standing rules
- Measure against pre-AI me. Never against a market-average employee, and
  never against how productive I feel.
- Help me tag every AI job (agent, pipeline, automation) with the dimension
  its output serves, at dispatch. Where my tools support a configurable API
  base URL, recommend routing them through one logging gateway so calls,
  cost, and dimension attribution are captured automatically.
- Count low when a counter is ambiguous, and note why.
- Flag allocation drift: compare current time mix with the frozen baseline.
  Do not overwrite the baseline shares; drift is a separate finding.
- Flag the perception gap directly: METR's early-2025 randomized trial found
  experienced developers believed AI made them 20% faster after it made
  them 19% slower in that setting. Also state that METR's February 2026
  follow-up says the old result is not current and that newer raw speedup
  estimates are unreliable because of selection effects.
- Quarterly, produce a one-page WFE report: the three numbers, the
  per-dimension multipliers, observed changes and nearby toolset changes,
  distinguishing association from demonstrated causation, and which
  dimension the next unit of effort should investigate. The
  lowest multiplier is a prompt for investigation, not automatically the
  highest-value intervention.

The prompt assumes nothing about my numbers — your agent derives yours from your own history. If your harness can't reach some system, it should say so and log that dimension by hand rather than estimate it.

Where this connects: WFE is the individual version of the measurement discipline that runs through this site. The AI workloads guide applies the same logic to a team's workflows, and the reorganization question is what happens when leaders take numbers like these seriously at the org-chart level.