FIELD GUIDE · AGI, DEFINITIONS & HYPE

WHEN IS AI ACTUALLY AGI?

There is no AGI referee. Labs disagree about whether the finish line is breadth, autonomy, human-level performance, or replacing valuable work.

UPDATED 2026-09-21

NO PHD REQUIRED

ELI5

AGI is not a single scientific measurement with one universally accepted finish line. Different definitions emphasize different things: how many kinds of tasks a system can do, how well it performs compared with humans, how independently it can work, or whether it can replace people across a large share of the economy. That is why one spectacular benchmark result can be real without settling whether AGI has arrived.

GENERALITY

How broadly a system can transfer competence across different kinds of tasks and domains.

Being superhuman at chess or protein folding is impressive but narrow.

PERFORMANCE

How well the system performs relative to people or other systems on the tasks being measured.

Breadth and depth are separate: a system can be broad but mediocre, or narrow and extraordinary.

AUTONOMY

How long and how reliably the system can pursue goals without step-by-step human guidance.

Some AGI definitions treat autonomy as central; others treat it as a deployment property layered on top of capability.

ECONOMICALLY VALUABLE WORK

Tasks people and organizations pay humans to perform across the economy.

OpenAI's longstanding definition explicitly uses this economic benchmark rather than consciousness or humanlike personality.

WHY THIS BECOMES A FIGHT

WHY ARE PEOPLE FIGHTING ABOUT THIS?

Calling something AGI changes how people talk about safety, regulation, labor, investment, national power, and whether a technology milestone has been crossed. But the label can also compress several different questions into one dramatic word. A system might be broadly capable without being reliable, economically transformative without being autonomous, or superhuman on many tests without matching humans in ordinary life.

GET THESE OFF THE TABLE

THE STRAW MEN

The bad arguments first. Nobody gets to win by beating these.

THE ROBOT-GOD VERSION

“AGI MEANS A SENTIENT ROBOT GOD. IF IT ISN'T CONSCIOUS, IT DOESN'T COUNT.”

THE LEADERBOARD VERSION

“IT BEAT HUMANS ON THREE BENCHMARKS. CONGRATS, AGI ACHIEVED.”

NOW MAKE THE GOOD ARGUMENT

STEEL MAN THE CASE

Give the people you disagree with the version they would actually defend.

THE CAPABILITY MAP

GENERALITY SHOULD BE THE CORE

A useful AGI definition should distinguish breadth from depth. Google DeepMind's Levels of AGI framework explicitly separates how general a system is from how well it performs, while treating autonomy as a related but distinct deployment dimension. This makes AGI more like a multidimensional capability map than a single finish-line score.

A decathlete is defined by breadth across events, not by setting the world record in one sprint.

THE WORK TEST

ECONOMIC REPLACEMENT IS THE CONSEQUENTIAL THRESHOLD

OpenAI defines AGI as highly autonomous systems that outperform humans at most economically valuable work. This definition focuses on social and economic consequence: once systems can independently do most valuable work better than people, the distinction matters even if they do not think, feel, or behave like humans.

The economy does not care whether the machine feels human if it can do the job.

FOLLOW THE MONEY

WHO PAYS? WHO WINS?

LABS + INVESTORS

A looser AGI label can make progress sound historic and justify enormous capital spending. A stricter definition postpones the milestone and raises the evidentiary bar.

WORKERS + EMPLOYERS

Economic definitions matter because they tie AGI to substitution and productivity rather than philosophical resemblance to humans.

REGULATORS

Rules triggered by vague labels can be gamed or misunderstood. Capability thresholds are easier to govern when the measurable dimensions are explicit.

THE PUBLIC

People get a clearer picture when headlines separate 'superhuman at this task,' 'broadly capable,' 'highly autonomous,' and 'economically transformative' instead of calling all four AGI.

RECEIPTS, NOT VIBES

WHAT DO WE ACTUALLY KNOW?

THERE IS NO SINGLE AGREED DEFINITION

Google DeepMind's Levels of AGI paper begins by analyzing multiple existing definitions and proposes a framework built around performance and generality, with autonomy considered separately.

The disagreement is not merely media confusion; researchers and labs operationalize the term differently.

OPENAI USES AN ECONOMIC + AUTONOMY DEFINITION

OpenAI's Charter defines AGI as highly autonomous systems that outperform humans at most economically valuable work.

That definition does not require consciousness, a humanoid body, or perfect performance at every possible task.

REAL-WORLD WORK EVALS ARE TRYING TO MAKE THE ECONOMIC VERSION MEASURABLE

OpenAI's GDPval evaluates models on economically valuable real-world tasks drawn from 44 occupations.

A benchmark like GDPval can measure part of the economic-capability question without by itself establishing AGI.

ONE AMAZING RESULT DOES NOT ESTABLISH GENERALITY

DeepMind's framework explicitly distinguishes narrow superhuman capability from broader general performance.

The relevant question is not only 'how high did it score?' but 'across how many genuinely different kinds of work?'

WHAT WOULD SETTLE SOME OF THIS?

WHAT WOULD CHANGE THE ARGUMENT?

TAKE THIS TO DINNER: Stop asking whether one score proves AGI. Ask what the system can do broadly, independently, reliably and for real money.

RECEIPTS

The guide is the map. These are the sources behind the substantive claims.

  1. Levels of AGI for Operationalizing Progress on the Path to AGIGoogle DeepMind / ICML
  2. OpenAI CharterOpenAI
  3. Measuring the performance of our models on real-world tasksOpenAI