Methodology, one page

How we built the AI job map.

A tool that lets any worker find their job and see how much of it AI can do today, read task by task instead of inferred from abstract traits. The frontier sets how capable AI is at a kind of work. The job's real tasks are the evidence.

The core shift. Most tools score a job by its abstract traits. This one scores the actual tasks the job is made of, then checks the result against a peer-reviewed index built a different way. The task wording is what makes it work: "analyze financial data" and "analyze laboratory specimens" score very differently, and only tasks can tell them apart.

Step 1

1The benchmark source, and its gap.

We start from BenchmarkList.com, which tracks about 2,500 AI benchmarks (the exams researchers use to prove AI can do a task), each with its top model, score, and link.

About 550 of them, 22 percent, are tagged "Uncategorized." That is a data-hygiene gap, not a real class: each still carried an ability tag but arrived in batch imports and never got a top-level label. So we did not rely on the site's own categories, and built our own frontier grouping instead.

Step 2

2Seven frontiers, each with a maturity.

We made the benchmarks the spine and grouped AI capability into seven frontiers. Each frontier's maturity (0 to 100) is the median top-model score across its benchmarks, lowered for embodied and relational work where a simulation score overstates real readiness. Each maturity traces to specific, named tests.

Math65/100
MatureCalculation, quantitative and structured problem-solving.
Named tests: AIME (selects the U.S. Math Olympiad team)
Language60/100
MatureReading, writing, summarizing, explaining in words.
Named tests: MMLU-Pro (expert exam across 14 fields), MTEB (text search)
Reasoning58/100
MatureWorking through a problem in steps; structured judgment.
Named tests: MMLU-Pro and multi-step reasoning suites
Coding58/100
MatureWriting and checking software.
Named tests: SWE-Marathon, DeepSWE (real bug fixes engineers merged on GitHub)
Perception & vision45/100
Still emergingSeeing and interpreting images, documents, the physical scene.
Named tests: MMMU-Pro (college exams), DocVQA (scanned documents)
Social & relational30/100
Still emergingReading people, real-time back-and-forth, being with someone.
Named tests: relational benchmarks, lowered because a test score overstates real readiness
Embodied / robotics18/100
Still emergingPhysically doing things in the world with hands and tools.
Named tests: RoboDojo, SurgVLA-Bench (surgical robotics)
Step 3

3Splitting jobs into real tasks.

This is the shift from abilities. We pull O*NET 30.3 Detailed Work Activities, which break every U.S. occupation into its actual tasks: 2,083 activities across 923 jobs, each weighted by how central it is to the role. BLS OEWS (May 2024) supplies employment and wages. The tasks are observed job content, not inferred traits.

Step 4

4The frontier leads, task by task.

Each task is classified to the frontier it draws on, and that frontier's maturity says how far along AI is there. Each task is then scored for how much a current model with normal tools can do it:

Can do most

The task sits on mature frontiers. Example: "prepare financial reports."

Assists

AI can do part, but a person still owns it. Example: "diagnose and treat patients."

Stays human

The task needs an emerging frontier, physical presence, or being with a person. Example: "position a patient for a scan."

Coverage is the task-weighted share AI can do. And the wording of the task decides the score. "Analyze financial data" scores high; "analyze laboratory specimens" scores low. Abilities could not separate those. Tasks can.

Step 5

5Validation, stated plainly.

Coverage correlates 0.79 (Spearman) with the peer-reviewed AIOE index, and shows the expected wage pattern. That is lower than the ability version's 0.92, and the reason is structural: AIOE is itself ability-based, and the literature finds task-based and ability-based measures diverge within desk work. So a cross-method 0.79 is a different reading, not a lower-quality one.

The counseling check

Counseling now reads about 44, down from about 70 in the ability version. Its "counsel clients," "intervene in crisis," and "advocate" tasks read as stays human, while its documentation reads as can do most. You can see which tasks drive the number, and disagree with any one of them.

Counsel clients, stays humanIntervene in crisis, stays humanAdvocate, stays humanDocumentation, can do most
Step 6

6What changed from the ability version.

This tool replaced an earlier ability-based version. The crosswalk, so nothing is hidden:

Coverage: share on mature frontiers AI can do
% AI can do: task-weighted share of tasks AI can do
Targeting X% AI can do
~X% AI can do
Ability and skill profile
Detailed Work Activities (real tasks)
Distinctiveness-weighted abilities
Task-centrality weighting
Frontier breakdown
Frontier and task breakdown
(no per-item mark)
Per task: can do most / assists / stays human
Validated 0.92 vs AIOE
Validated 0.79 vs AIOE (task method)

One audit fix worth naming. Several tasks named for a "program" (program finances, program eligibility, program participants) were misfiled under Coding by the word alone. They were moved to Reasoning, Language, or Social. Coverage did not move, because it reads each task's own score, not its frontier label.

Step 7

7Outputs.

The web tool: find your job, read the coverage headline, open the frontiers and their tasks. A companion spreadsheet holds the full mapping, a Jobs sheet (coverage plus each frontier's share per occupation), a Task_Mapping sheet (all 2,083 activities, frontier, and AI-can-do score), a Job_Tasks sheet (pivot-ready), and a Frontier_Reference sheet (maturity, best models, example tests).

Honest limits

8Where it's weakest.

1

Exposure is not replacement. The most-exposed jobs have historically grown. Read coverage as overlap with the work, not as a prediction that anyone loses a job.

2

Classification is not perfect. The 2,083 tasks are sorted by an object-aware rubric. Most are right; a few edge cases mis-sort, and those are known.

3

The 0.79 is a floor, not a ceiling. Measured against an ability index by a different method, 0.79 is a floor on cross-method agreement, not a limit on quality.

Sources

What it's built from.

  • BenchmarkList.com. Frontier maturity and best models.
  • O*NET 30.3 Detailed Work Activities (USDOL/ETA). The tasks and jobs.
  • BLS OEWS, May 2024. Employment and wages.
  • Felten, Raj & Seamans (AIOE) and Eloundou et al., "GPTs are GPTs." The task-level method and the validation.

A task-level exposure and progress map, not a prediction of job loss.

See it for any job.

Open the tool