Short version
AGI is a hypothetical machine that can learn, reason and solve any intellectual task a person can. Hypothetical is not a hedge: no such system exists, and nothing you can buy today meets the definition. It is back in every conversation because on Sunday 6 September Nvidia’s Jensen Huang declared “AGI has arrived” on the strength of OpenAI’s new Astra model, offering — critics noted the same day — no test and no definition. On that same day OpenAI’s own chief scientist published an essay saying no lab has solved alignment well enough to keep scaling at full speed, and Sam Altman called it important. There is no agreed test for how far along anything is, which is why serious people give opposite answers about how close we are: one framework puts the previous flagship at 57% of a well-educated adult, where a person scores about 95% (AI Frontiers). The real thing that changed is measurable: how long these systems work unattended has been doubling roughly every three months (METR).
The three numbers behind the noise
- How close, on a serious measure GPT-5 reached 57% of what a well-educated adult manages; a person scores about 95% Source: AI Frontiers
- What actually changed Unattended work is up to about 320 minutes — five hours — for the best measured model Source: METR Time Horizon 1.1
- How fast it is moving That figure doubles every 89 days on data since 2024 Source: METR Time Horizon 1.1
What AGI actually means
The standard definition is short: AGI is a hypothetical machine that can learn, reason, and solve any intellectual task a human being can. Two of those words do all the work, and both are usually left out when the term gets thrown around.
Hypothetical means it does not exist. Not "nearly here", not "in beta" — no such system has been built, and every product you can buy today falls outside the definition. Any is the bar: not many tasks, not most, not the ones it was trained for. Any. A system that handles a thousand jobs brilliantly and falls over on the thousand-and-first is not general; it is a very good narrow system. That single word is why the finish line keeps receding.
Almost all the software you use is the opposite of general. It is narrow: built for one job, and useless one step outside it. Your payroll system cannot read a contract. Your spam filter, which is a machine-learning model and a very good one, cannot summarise a meeting. Each of those tools was built, trained and tuned for a single task, and if you want a new task you buy or build a new tool. That has been the shape of business software for forty years.
A general system is the other thing. You hand it something nobody built it for, and it works out how to do it. That is not a description of software; it is a description of a person. When you hire someone, you do not retrain them from scratch to move from the invoice pile to the supplier portal. They carry what they know across, work out the new thing, and ask when they are stuck. AGI is the label for software that does that.
Which is why "can it do my job" is the wrong test, tempting as it is. Handing work to software is a useful thing to be able to do, and it is what this month's products are sold on — but the definition asks for something stricter: learning, reasoning and solving across any intellectual task, including the ones nobody anticipated. A tool that finishes an afternoon of work you specified is not the same as a mind that can take on work nobody specified. Keeping those two apart is most of what it takes to read AI news without being misled.
Why "general" is the hard part
It is tempting to think generality is just narrow ability repeated — get good at enough separate things and generality falls out. It has not worked that way, and it is worth understanding why, because the reason explains most of what you read about AI's limits.
Three things a person does without noticing turn out to be genuinely difficult. The first is transfer: applying something learned in one place to a situation that only rhymes with it. The second is novelty: handling a case nobody anticipated, which is most of what a working day contains. The third, and the one that causes the most trouble in practice, is knowing when you do not know — stopping, flagging it, and asking. A new employee who is unsure says so. Software that is unsure has historically produced a confident answer instead.
That last gap is not theoretical, and it is measurable. On one test of straightforward factual questions, GPT-5 gave a wrong answer to more than 30% of them rather than declining to answer (AI Frontiers). The same review notes that on a test of whether a short video shows something physically possible — a ball falling the way balls fall — the best models perform barely better than guessing. A system can write you a competent contract summary and fail a question a five-year-old gets right. That combination is exactly what "not yet general" looks like from the outside.
Why nobody can agree whether we have got there
Here is the part that explains the noise. There is no agreed test for AGI, and there never has been.
For decades the accepted answer was the Turing test: can a machine hold a conversation well enough that you cannot tell it is a machine. Today's models sail past that, and almost nobody treats the question as settled — which told the field something useful. The test was wrong, not the answer. Passing it turned out to measure fluency, and fluency turned out to be far easier than understanding.
The most serious current attempt at rigour borrows from psychology. Instead of one clever test, score the system across the ten broad cognitive abilities that psychologists use to assess people, and define AGI as matching a well-educated adult across all of them. On that framework, GPT-4 scored 27% and GPT-5 scored 57%, where human level sits at roughly 95% (AI Frontiers). Thirty points of progress in one model generation, and a long way still to go — a far more useful picture than any yes or no.
But that is one framework among several, and this is why serious people contradict each other in public. Each lab, and each critic, is holding a different measuring stick. When you next see a confident claim that AGI has arrived or is decades away, the only question worth asking is: measured how? If the answer does not come, the claim is decoration.
Why the word came back in September
None of the above changed this month. Something else did, and it is the actual news.
In the first week of September, Anthropic and OpenAI released flagship models two days apart, and both were sold on the same capability: not answering better, but working. Astra's headline feature is computer use — opening programs, filling in forms, clicking through websites, working inside spreadsheets on your behalf, for hours, without asking what to do next (OpenAI's Astra documentation). Anthropic's pitch for Fable 5.1 was long-running work across many programs that recovers when a step goes wrong (Anthropic's launch post).
That is the shift. For three years these systems were things you asked. Now they are things you give a job to and leave. And a machine that works unattended for an afternoon does not look like a tool any more — it looks like a worker. The word "AGI" came back because the thing on the screen finally resembled what the word describes, whatever the benchmarks say.
Then, on Sunday 6 September, the loudest possible voice said it outright. Jensen Huang, chief executive of Nvidia — the company whose chips train essentially every frontier model — posted on X: “From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team.” OpenAI's president Greg Brockman shared the post and said the company is now moving into the AGI era (Washington Examiner). That is the sentence that put the word into every conversation this week, and it is why you are reading about it now.
Look at what the declaration rested on. Huang's post pairs the claim with a hardware figure — Astra trained on more than 100,000 of Nvidia's Grace Blackwell systems, with 400,000 more GPUs coming (Dataconomy). No test was named, no definition offered, and no new technical result accompanied it. It is also not the first time: Huang said “I think we've achieved AGI” on a podcast in March 2025 (Yahoo Tech).
The rebuttal came the same day, and it is worth reading in full because it makes the point of this entire article. Gary Marcus, one of the field's most persistent sceptics, wrote: “Unfortunately, Huang gave no evidence and no definitions, which feels to me like an effort at a takeover of a scientific question by corporate fiat.” On the model itself he added that many see Astra as “not much more than on a par with Fable 5.1 in real-world applications”, and closed: “When real AGI arrives, we won't need to squint our eyes. And we won't need Jensen's approval, either.” (Gary Marcus, 6 September 2026).
And then the genuinely strange part. On the same day Huang declared the finish line crossed, OpenAI's own chief scientist published an essay arguing for slowing down. Jakub Pachocki, in a piece called An Alien Mind: “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” He went further — “This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence” — and called for mandated safety standards policed by outside auditors. Sam Altman reposted it and called it “an important post” (TNW).
Hold those two together, because they are the honest picture of the week. One company's president says it is entering the AGI era; the same company's chief scientist says nobody has solved the safety problem well enough to keep going at this speed; the chief executive calls the warning important. Whatever arrived on Sunday, the people closest to it are not behaving like a finished technology.
One set of numbers from that same essay says more about where things actually stand than any benchmark. Inside OpenAI, by mid-August, AI agents were doing 3.1 agent-workdays of research for every human workday, and the median researcher was spending more than $600 a day on model usage, with the busiest tenth burning over $7,000 a day (TNW). That is what the frontier looks like from inside: not a machine that thinks like a person, but a very large number of machine-hours, bought at considerable cost, pointed at a narrow kind of work.
There is a commercial reason the word is attractive, and it is fair to name it without accusing anyone of lying. A term with no fixed definition cannot be disproved. Any company can position itself close to AGI against some reading of it, and nobody can prove otherwise. That is a useful property in an announcement and a useless one in a purchasing decision.
The number underneath the noise
Strip out the word and one real, independently measured trend is left — and it is the best explanation of why the mood changed in September.
METR, a non-profit that evaluates frontier models rather than selling them, measures what it calls the time horizon: take jobs of known human duration, and find the length at which a model succeeds half the time. It is a deliberately boring number, and it has been climbing in a straight line for years.
Two numbers matter. The best model with a published figure — Claude Opus 4.5, at 320 minutes — gets through about five hours of human work on its own, an afternoon rather than a week. And the figure doubles every 89 days measured from 2024, or every 196 days across the whole run (METR). AI Digest, tracking the same work, puts the recent rate near four months while cautioning in its own words that "looking at just one year's data gives a less robust estimate" (AI Digest).
A capability that doubles every few months is invisible for a long time and then feels sudden. Five years of this trend produced tools that could hold a conversation. This year it produced tools that can work through an afternoon. That is why the word came back now — not because anyone crossed a defined line, but because a steady exponential finally reached the length of a human task.
One honest gap: neither September model has a published figure on this scale. Anthropic's system card says only that external testing by METR produced findings consistent with its own assessment, with no number attached (Fable 5.1 system card, PDF). The vendors describe hours and multi-day sessions; the independent confirmation is not published.
What actually changes — and what does not
Here is the practical translation, and it is narrower than the word suggests.
What changes: a category of work that was never worth automating becomes worth handing over. Not because the software got clever, but because it can now stay on a job long enough to finish it and pick itself up when a step fails. That suits work which is long, repetitive, follows rules you could write down, spans several programs that were never integrated, and — critically — produces something you can check quickly.
| Now worth handing over | Still not |
|---|---|
| Month-end reconciliation across two systems that do not talk to each other | Deciding which of the mismatches it finds are acceptable |
| Pulling data out of a supplier or authority portal that has no integration | Anything where a wrong figure would sit unnoticed for weeks |
| First-pass review of a stack of contracts or invoices for dates and amounts | The decisions that follow the shortlist it produces |
| Reading a week of support tickets and reporting what keeps recurring | Judgement calls a person would escalate to someone else anyway |
| Applying one specified change across dozens of documents or pages | Work nobody can verify without redoing it from scratch |
The capability underneath the left column — operating ordinary software unattended for hours — is documented by both vendors and linked above. Which work it suits is our reading, not a vendor claim.
What does not change: nothing here replaces a role. Every item in the left column is a task, not a job, and each one hands its output to a person who decides what it means. The gap between "can finish a five-hour task" and "can do someone's job" is not a matter of a few more months of the same trend — it is the transfer, novelty and knowing-when-you-do-not-know problem from earlier in this piece, and that problem is not being solved by making the afternoon longer.
Where you still need a person
One piece of arithmetic tells you where to stand, and it is the most useful thing in this article.
Suppose the tool is right nineteen times out of twenty on any single step. That is a good tool. Now give it a job that takes twenty steps — open the export, match the rows, look one up, correct it, write the summary. If every step has to be right for the job to be right, those twenty steps come out right only about a third of the time. Make it wrong just once in a hundred steps instead, and the same job lands about four times in five. (Our own arithmetic, not a measurement: nineteen-twentieths and ninety-nine hundredths, each raised to the twentieth power.)
So a small change in per-step reliability transforms a long job, and your reviewer belongs at the steps where an error is expensive and silent — not at the end, where it is already baked into the result. In a reconciliation that is the exceptions list, not the final total.
It also explains what the vendors actually sold. Fable 5.1's headline benchmark for long multi-step work more than doubled, from 24.7% to 52.6%, as claimed by Anthropic (Anthropic). A tool that notices it went wrong and backs up does not need near-perfect steps, because one bad step stops being fatal. That is the real advance of September, and it is an engineering advance, not a leap in intelligence.
Two more cautions worth carrying. Both companies restrict their own strongest versions — OpenAI gave Astra the first critical cybersecurity classification, meaning it can find and exploit previously unknown software holes on its own, and limits that configuration to vetted users (OpenAI deployment safety report). And Anthropic's documentation tells most readers to start with its cheaper model and move up only if their own testing shows a gap (Anthropic models overview). Companies confident they had built a finished general intelligence would not be fitting brakes to it.
Bottom line
AGI is a hypothetical machine that can learn, reason and solve any intellectual task a person can. It does not exist. The word "any" is the bar, and it is why the finish line keeps moving: a system that handles a thousand jobs and fails the next one is a very good narrow system, not a general one. It has no agreed test, which is why serious people give opposite answers about whether it has arrived, and why the term settles nothing on its own.
It came back in September for a real reason. Two flagship models arrived that work through an afternoon unattended rather than answering questions, one company's president attached the word to his product in public, and underneath both sits a measured trend that has been doubling every few months until it finally reached the length of a human task.
What that gives you is not a replacement for anyone. It is a new category of work you can hand over — long, repetitive, checkable — and a clear instruction about where to keep a person, which is at the steps where being wrong is expensive and quiet. That is a smaller story than the word implies, and a much more useful one.
FAQ
What does AGI stand for, in one sentence?
Artificial general intelligence: a hypothetical machine that can learn, reason and solve any intellectual task a human can. Hypothetical means it has not been built — every AI product on sale today is narrow, doing one kind of thing and nothing next to it.
How is that different from the AI we already use?
Your spam filter and your payroll system are narrow: built and trained for one job, useless one step outside it. A general system takes on work nobody built it for and works out how, the way a new colleague moves from the invoice pile to the supplier portal without being retrained.
Why is "general" so much harder than being good at many things?
Three human habits turn out to be difficult: transferring what you learned in one place to a situation that only rhymes with it, handling cases nobody anticipated, and knowing when you do not know. The third causes the most trouble in practice — on one test of factual questions, GPT-5 answered wrongly more than 30% of the time rather than declining (AI Frontiers).
Has anyone actually reached AGI?
No test exists that everyone accepts, so the question has no clean answer. On the most serious current framework — scoring a system across the ten cognitive abilities psychologists use for people — GPT-4 reached 27% and GPT-5 reached 57%, against roughly 95% for a well-educated adult (AI Frontiers).
Whatever happened to the Turing test?
Today's models pass it comfortably and almost nobody considers the matter closed. That taught the field that the test measured fluency rather than understanding — and fluency turned out to be much easier than anyone expected.
So why did everyone start saying it again in September?
Two flagship models launched two days apart that operate a computer for hours on their own rather than answering questions (OpenAI's Astra documentation; Anthropic's Fable 5.1 launch), and OpenAI's president invoked the term at his launch. Software that works unattended for an afternoon stops looking like a tool and starts looking like a worker.
Did Nvidia's CEO really say AGI has arrived?
Yes, on Sunday 6 September 2026, on X: “From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team.” He paired it with a hardware figure rather than a test result, named no definition, and had said something similar in March 2025 (Yahoo Tech). Gary Marcus called it “a takeover of a scientific question by corporate fiat” (Marcus).
What did OpenAI's own chief scientist say?
On the same day, Jakub Pachocki published An Alien Mind: “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” adding that the moment “calls for extreme caution”. Sam Altman reposted it and called it an important post (TNW).
Is there a commercial motive behind the word?
A term with no fixed definition cannot be disproved, so any lab can position itself close to AGI against some reading of it. That is a useful property in a launch and a useless one in a purchasing decision. It does not make anyone a liar; it does mean the word carries no information by itself.
What is the "time horizon" everyone keeps citing?
The length of job, measured in how long a person takes, that a model finishes on its own about half the time. METR measures it across 228 tasks and it is independent of the vendors. The best model with a published figure manages about five hours (METR).
How fast is that actually moving?
Doubling every 89 days on data since 2024, or every 196 days across the full run (METR). AI Digest puts the recent rate near four months and cautions that one year of data is a weak basis for projecting (AI Digest).
Is this going to take my job?
Nothing in the evidence points that way. These tools finish tasks and hand the result to a person who decides what it means — and the distance between finishing a five-hour task and doing a whole job is the transfer, novelty and knowing-when-you-do-not-know problem from earlier. That gap does not close by making the afternoon longer.
What should I actually do with this?
Take the most repetitive thing you do whose result you could check in a couple of minutes, give it to one of these tools, and watch where it stops. How long it ran before it needed you tells you more than any benchmark, because it is about your work rather than someone else's exam.
If my employer wants to use this, what has to be sorted out first?
Anything running unattended against company systems touches your data processing agreement, your works council and your EU AI Act transparency duties. Those belong to whoever owns them in your organisation — raise them in week zero rather than month three. Nothing here is legal advice.
What would genuinely change the picture?
A published independent time horizon for either September model. Neither has one: Anthropic's system card says external testing by METR was consistent with its own assessment but attaches no figure (system card).
ISA After Hours · Augsburg
Working out what to make of all this?
ISA After Hours is a community of Israeli and international tech professionals in Augsburg. We run these tools on our own work and compare what actually held up — which finds the limits faster than reading another launch post.
Join ISA After Hours →Sources
- garymarcus.substack.comGary Marcus, 6 September 2026: the rebuttal to Huang — no evidence, no definitions, “corporate fiat”, and the comparison of Astra to Fable 5.1 in real-world use
- thenextweb.comTNW on Jakub Pachocki's “An Alien Mind”, 6 September 2026 — the alignment warning verbatim, Altman's response, and OpenAI's internal agent-workday and inference-spend figures
- washingtonexaminer.comWashington Examiner: Huang's “AGI has arrived” post and Greg Brockman's response
- dataconomy.comDataconomy: the hardware figures Huang attached to the claim — 100,000 Grace Blackwell systems, 400,000 more GPUs coming
- tech.yahoo.comYahoo Tech: the announcement, and Huang's earlier March 2025 claim to have achieved AGI
- ai-frontiers.orgAI Frontiers: AGI's last bottlenecks — the ten-ability framework putting GPT-4 at 27% and GPT-5 at 57% against a human 95%, the factual-error rate, and the physical-plausibility result
- metr.orgMETR: Time Horizon 1.1, 29 January 2026 — the 228-task suite, the 89- and 196-day doubling figures, and the measured horizons for Opus 4.5, GPT-5, o3, Opus 4 and Sonnet 3.7
- theaidigest.orgAI Digest: A new Moore's Law for AI agents — the four-month versus seven-month comparison and the caution about one year of data
- developers.openai.comOpenAI: GPT-6 Astra model documentation — computer use, context window and availability
- anthropic.comAnthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1 — the long-horizon capability claims and the Terminal-Bench-Science figures
- anthropic.comAnthropic: Claude Fable 5.1 and Mythos 5.1 system card (PDF) — the METR external-testing statement, which carries no published figure
- platform.claude.comAnthropic models overview — the documented guidance to start with the cheaper model
- deploymentsafety.openai.comOpenAI: Astra deployment safety report — the critical cybersecurity classification and the vetted-access restriction
- isaiaugsburg.comOur GPT-6 Astra guide — prices, limits, the AGI framing at launch and how it is sourced
- isaiaugsburg.comOur Claude Fable 5.1 guide — prices in scenarios, limits and the suitability matrix