What Is AGI (Artificial General Intelligence)? Explained

In December 2025, OpenAI's CEO Sam Altman said "we built AGIs" — and that AGI "kinda went whooshing by" with less societal impact than expected, so the field should move on to defining superintelligence. A few weeks later, in 2026, OpenAI's president Greg Brockman declared that the company's new model marked the beginning of the "AGI era." Same company. Same word. One says it already passed; the other says it just started.

That's not a PR stumble — that's the actual state of AGI. It's a goal without an agreed definition, which is how the same organization can announce, in the same season, that AGI is behind us and ahead of us at once. This piece explains why that keeps happening, what AGI actually refers to, and how to think about it without either the hype or the eye-rolling.

One disambiguation first, because it genuinely confuses searches: on US tax forms, AGI stands for adjusted gross income, and IRS pages explaining it sit right alongside AI articles in the search results. If you came looking for tax help, that's the wrong AGI. Here we're talking about artificial general intelligence.

The one-sentence definition, and the three terms around it

Strip away the noise and the core idea is short: AGI is a hypothetical AI system that matches or surpasses human capabilities across virtually all cognitive tasks — not just the ones it was trained for. The key moves are generalizing knowledge, transferring skills between domains, and solving novel problems without task-specific reprogramming, as Wikipedia's summary puts it. A chess engine can't do your taxes. A translation model can't learn to cook. An AGI, in principle, could walk into any of those situations and figure it out.

The term sits in a family of four, and most confusion comes from mixing them up:

TermWhat it meansStatus
AIThe umbrella term: any system doing tasks that need intelligenceEverywhere — search, recommendations, voice assistants
AIGC / generative AIThe branch that creates content: text, images, audio, codeCurrent wave: ChatGPT, Midjourney, and friends
AGIHuman-level ability across essentially all cognitive tasksHypothetical. No agreed test exists
ASIBeyond-human ability across all tasksMore hypothetical still

Notice the word "hypothetical" attached to AGI in nearly every careful definition, including Stanford HAI's. Nobody has built it, and — this is the part people miss — nobody has even agreed on what would count as building it.

Why nobody agrees on the definition

Here's a fact that surprised me when I first dug into it: the fight over what AGI means is older than ChatGPT, older than deep learning, and nearly as old as the field itself.

Start with a twist. When "artificial intelligence" was coined at Dartmouth in 1956, the goal was general intelligence — machines that could use language, form abstractions, solve kinds of problems reserved for humans. That was the original dream. Then decades of narrow, specialized systems happened, and at some point "AI" came to describe exactly those specialized systems. The term "AGI" emerged partly to reclaim the original meaning. As one researcher put it, AGI refers to the original AI at the inception of the field. The acronym is newer than the idea: Mark Gubrud used "artificial general intelligence" in 1997 writing about automated military production, and Shane Legg and Ben Goertzel popularized the term around 2002.

Then the arguing started. A well-known 2007 collection by Shane Legg and Marcus Hutter gathered dozens of published definitions of intelligence — professional researchers, all using the same word, meaning materially different things. A year earlier, the AGI Workshop's proceedings had already conceded it was "too early to conclude with any scientific definiteness which conception of 'intelligence' is the 'correct' one." Twenty years later, Stanford HAI's glossary still says there's "no universally accepted test," so claims are hard to verify. Even John McCarthy — the man who coined "AI" — wrote in 2007: "We cannot yet characterize in general what kinds of computational procedures we want to call intelligent."

The deepest split is worth understanding, because it explains most headlines. One camp measures intelligence by problem-solving: if the system performs at human level on the tasks, call it intelligent. Google DeepMind's levels-of-AGI framework (more on that below) lives here. The other camp insists intelligence is the learning capability — a system that can't learn isn't intelligent, however impressive its answers; a calculator solves arithmetic flawlessly and nobody calls it a genius. The definitional paper I kept returning to argues that typical machine learning systems are intelligent during training and not intelligent at test time, because once deployed they stop learning. Under that lens, a model that aces every benchmark but never updates itself is a very sophisticated recording, not a general intelligence.

Neither camp is wrong. They're answering different questions — "what can it do?" versus "how did it get the ability?" — and until the field agrees on which question matters, the definition stays open.

The AI effect: why every milestone gets ruled "not real intelligence"

There's a pattern in AI history with its own name: the AI effect. Each time machines crack a capability believed to require intelligence, the reaction is first amazement, then a quiet reclassification — "well, that turned out to be just search" or "just statistics" or "just pattern matching." Chess fell to Deep Blue in 1997, quiz shows to Watson, Go to AlphaGo in 2016. Each time, the goalposts moved.

The definitional paper cited earlier nails the mechanism: once a problem is solved, people look back and realize it was really the human developers who solved it — the machine just executed. No knowledge was acquired by the machine itself, so it "doesn't count." You can see the goalposts moving in real time: the moment models got good at conversation, "chatting" was retroactively demoted from an intelligence test to a parlor trick. Japan's Fifth Generation project in the 1980s listed "carry on a casual conversation" as an AGI-level goal. Today that same ability is considered trivial.

This loop is why I tell people they may never see an official AGI announcement that everyone accepts. Any candidate system will be examined, its methods dissected, and a sizable fraction of the field will rule that it's "just" whatever the method turned out to be. "Just" is the toll every AI milestone pays.

A ruler you can actually use: DeepMind's levels of AGI

Fine, definitions fight. Is there anything practical? Yes — the most useful artifact so far is Google DeepMind's 2023 framework, which trades "is it AGI?" for "what level is it?"

The idea: score systems on performance (emerging, competent, expert, virtuoso, superhuman) and on breadth of tasks. A "competent" system outperforms 50% of skilled adults across a wide range of non-physical tasks; "superhuman" means beating essentially 100% — which is where ASI begins. There's a second axis for autonomy, running from tool (fully human-controlled) through consultant and collaborator up to a fully autonomous agent — that last step is exactly the AI agent territory we cover separately.

The punchline that made the framework famous: DeepMind classified ChatGPT-class large language models as "emerging AGI" — performance comparable to unskilled humans, across a broad range of tasks. Not zero. Not "not AGI in any sense." Emerging.

The framework has real critics, and the criticism is instructive. If you define AGI purely by problem-solving, then gluing together a thousand specialized solvers on one computer would technically qualify — it solves everything, but nobody would call it general intelligence. The learning-versus-doing split from earlier shows up here too: measuring what a system scores says nothing about whether it could adapt to a genuinely new situation, which is the whole point of "general."

My own habit: when a lab declares something is or isn't AGI, I don't argue the label anymore. I ask which level, on which axis, judged by whom. It converts religious debates into table entries. You can do this at home with any product launch, and it takes about a minute.

Where we actually are: jagged, not general

So where do today's models stand? The honest one-word answer is "jagged." Current frontier models are superhuman on some slices (recall across millions of documents, certain math competition problems), decent-to-strong on many knowledge tasks, and startlingly weak on others — the failure mode gets called jagged intelligence: capabilities distributed extremely unevenly, brilliant here, common-sense-deficient there. The pattern isn't a footnote; it's the signature of how these systems learn.

Three specific gaps separate today's best models from anything you'd call general:

Reliability. The gap isn't that models "don't know" things — it's that they can be confidently wrong. That's the hallucination problem, and we have a whole piece on why models hallucinate; the short version is that a next-token predictor has no native mechanism for "I don't know." A system that's right 95% of the time and certain 100% of the time is not yet something you can hand a novel, open-ended job.

Long-horizon planning. Models can write a function or plan a trip. Running a months-long project — setting goals, decomposing, recovering from setbacks, re-planning when the world changes — is a different sport, and agents are only starting to attempt it.

Continual learning. Today's models are frozen at deployment. They don't learn from your conversation, remember it next month, or update themselves from experience. Training-time intelligence, test-time statue — the exact critique from the definitions debate, now felt as a product limitation.

This is also why "looks like AGI" ≠ "is AGI." A multimodal model that reads screenshots, writes code, and answers in your voice feels general in demos. Under the hood it's still one giant frozen function. Impressive, increasingly useful — and not what the researchers mean by general.

How long until AGI? A table, not a promise

Everyone asks for the AGI timeline. Here's the honest format for that answer — a table of who said what, because the spread is the answer:

WhoEstimate
Demis Hassabis (DeepMind CEO, 2023)Within a decade, possibly a few years
Jensen Huang (NVIDIA CEO, 2024)Within 5 years, AI passes any human test
Leopold Aschenbrenner (ex-OpenAI, 2024)2027 is "strikingly plausible"
Sam Altman (OpenAI CEO, Dec 2025)It already "went whooshing by"
Researcher surveys, 2012–2013 (median)50% confident by 2040–2050
Same surveys, mean2081
Same surveys16.5% answered "never"
Recent survey roundups (2025)Most say before 2100; current median around 2040

Notice the shape: people whose companies are building AGI cluster at "a few years." People who study it full-time cluster around 2040, with a meaningful tail saying never. And a 2012 meta-analysis of 95 predictions found something almost comedic — both historical and modern predictions skew toward "16 to 26 years from now." Whatever year you ask in, AGI tends to be 20 years away.

The field has been burned before, twice. In the early 1970s, funders realized researchers had grossly underestimated the problem and pivoted money to "applied AI." In the 1980s, Japan's Fifth Generation project promised AGI-adjacent goals on a ten-year timetable and collapsed with them. After two misses in twenty years, researchers stopped predicting altogether — they were tired of being called wild-eyed dreamers. Today's confident CEO timelines are exactly what the 1980s sounded like, and today's cautious survey medians are the scar tissue talking.

So my take: treat AGI as a direction, not a date. The capabilities are compounding on schedule; the moment of "arrival" will be argued about forever, because it's a definitional event, not a physical one.

Quick answers

What is AGI, in one sentence? A hypothetical AI system with human-level (or beyond) ability to learn, reason, and apply knowledge across essentially all cognitive tasks — not just the tasks it was built for.

AGI vs AI: what's the difference? "AI" is the umbrella, including every narrow system (recommendations, face recognition, chess engines). AGI is the hypothetical endpoint: one system that generalizes and transfers across domains instead of doing one thing well.

Is ChatGPT an AGI? By DeepMind's leveling, ChatGPT-class models are "emerging AGI" — roughly unskilled-human performance across many tasks. By the learning-capability school, they're not general intelligence at all, since they stop learning at deployment. Both answers are defensible; "it's settled" is not.

Is AGI the same as strong AI? Colloquially they're used interchangeably. In academic philosophy, "strong AI" specifically means a machine with actual consciousness or mind — a claim most AI researchers consider outside their scope. AGI, as most researchers use it, is about performance and generality, not consciousness.

When will AGI happen? Depends who you ask: industry leaders say within years, researcher surveys put the median around 2040, and 16.5% of surveyed experts said never. If someone gives you a single confident date, they're selling something.

If you keep one thing from this piece, keep the question, not the answer. Instead of "is it AGI yet?" — ask "what level, on which axis, and who's measuring?" That question never goes out of date, no matter how many times the goalposts move.