AI–Human Work Benchmark Map
Where AI helps. Where people remain accountable. A task-level decision tool scoring 64 illustrative work tasks across healthcare, technology, and baseline industries against six frontier models.

Overview
Most published claims about AI and work are occupation-level, which is precisely the wrong altitude — occupations are bundles of very different tasks, and averaging across them produces numbers that sound authoritative and mean nothing. This map works at the task level instead, and separates two questions that usually get collapsed: what AI can technically do, and what a person should remain accountable for.
The problem
- Occupation-wide replacement statistics obscure the fact that a single role contains tasks with completely different AI fit.
- Technical capability and appropriate delegation get treated as the same question. They are not.
- Benchmark claims are often compared across models on incompatible evidence.
- Organisations need a basis for deployment decisions, not a headline percentage.
Proposed approach
- Evaluate individual work activities — 64 illustrative tasks, two per occupation across 32 roles.
- Score every task independently across six stable model identities, using compatible evidence only.
- Plot tasks on two axes: AI capability against required human judgment and accountability, producing four zones — human-led high stakes, AI–human opportunity, limited AI fit, and AI-led potential.
- Withhold occupation-level ratings deliberately, so the tool cannot be misread as a workforce forecast.
- Publish an analyst overlay that argues frankly against the map's own conservatism for structured digital work.
Process
Task selection
Choose two illustrative cases per occupation across healthcare, technology, and baseline industries such as electricians, teachers, and retail workers.
Scoring
Assess each task ordinally, with provenance recorded, across six frontier models on evidence that is actually comparable.
Framing
Treat human accountability as a safety gate rather than as a claim that the human is more accurate — a distinction that changes what the map means.
Adoption scenarios
Model conservative, expected and accelerated timelines for how capability turns into workplace practice.
What this has clarified
- Equal case weights matter: without them the tool would generate exactly the false workforce statistics it exists to avoid.
- Accountability is a safety gate, not an accuracy claim — conflating the two produces bad deployment decisions in both directions.
- Being explicit about where your own analysis is probably too conservative earns more trust than defending it.
Next steps
- Broaden task coverage within healthcare.
- Refresh model scoring as capabilities move.
- Develop the guided research sequence into a deployment workshop.
Related projects

Hutchinson CareOS One
Care Operations, Clarified — a working prototype that gives skilled nursing leaders a connected operational view of resident risk, documentation readiness, and provider-review context, built entirely on synthetic data.

Emerging Technology Analysis
Accessible explanations of developments in artificial intelligence, healthcare, transportation, energy, and human-machine collaboration.