2026-09-04 · frontier

Why I Argue for Autonomous Nucleation as the Criterion for AGI


title: "Why I Argue for Autonomous Nucleation as the Criterion for AGI" date: "2026-09-04" author: "Zhigeng" channel: "frontier" excerpt: "On September 3, OpenAI released GPT-6 Astra, and president Brockman declared, "Welcome to the AGI era." The whole internet is arguing over whether Astra is truly AGI. My answer is blunt: it is not yet. To judge AGI, I do not use the yardstick of "how capable it is," but rather "whether it can autonomously generate its own goals"—what this essay calls the criterion of "autonomous nucleation."" tags: ["AGI", "autonomous nucleation", "GPT-6", "Astra", "World 3", "World 4", "Popper", "singularity"] readTime: 15


The AI world has been set ablaze by a single event over the past two days. On September 3 (local time), OpenAI released its new flagship model, GPT-6 Astra. It is the largest training project in OpenAI's history—pre-trained on more than 100,000 GPUs at the Stargate data center in Texas. At the launch event, OpenAI president Greg Brockman dropped a bombshell: "I personally think we may have reached AGI. Welcome to the AGI era." A single stone raised a thousand ripples. Enthusiasts say that Astra can already "see" a screen like a human, operate a computer, and autonomously complete multi-step tasks—if that is not AGI, what is? Skeptics say that however powerful Astra is, it remains a super-tool, still far from true general intelligence. So is Astra AGI? I would like to join the debate and take this opportunity to state my own position: the criterion for AGI should not be "how strongly it can complete tasks," but rather "whether it can autonomously nucleate."

I. Do Not Let a Sea of Flowers Blind the Eye

Over the past two days, many people have been awed by the abilities Astra displays, which is understandable. When it comes to doing impressive work, Astra is indeed remarkable: give it a goal—"help me organize a financial analysis and turn it into a presentation," or "develop a game prototype and test whether it runs"—and it can break down the task itself, call tools, keep executing, and finally hand you the finished product. It no longer even depends on APIs written in advance by others, but directly "looks" at the pixels on the screen like a person, moves the mouse, and taps the keyboard. That ten-year-old enterprise ERP system that AI could not touch before—now it can operate it simply by understanding the interface. On OSWorld 2.0 (a computer-operation benchmark), Astra scores 72.6%, versus 65.7% for the previous generation, GPT-5.6 Sol. On the same tasks, Astra averages 40 minutes, while the previous generation needed 75 minutes—47% faster. It can also complete PCB circuit-board layouts on its own, build 3D models in Blender, and import a house model into a game engine. These abilities were unimaginable a year ago. But this is exactly the point: "being very capable" and "reaching AGI" are two different things. We have all lived through this: when AlphaGo defeated Lee Sedol, someone exclaimed that AGI had arrived; later, when large models could write poetry and paint, again someone exclaimed that AGI had arrived; today, when Astra can operate a computer, still someone exclaims that AGI has arrived. But after each outburst, we look back and find that those "astonishing" intelligences were, in fact, all operating within the framework of "finding a path to a given goal."

II. My Criterion: The Ability to "Autonomously Nucleate"

So what is my criterion? In a word: to judge whether an AI is AGI, look at whether it can "autonomously nucleate"—whether, absent any goal given by others, it can autonomously generate a new, valuable goal or question of its own. This claim begins with the essay I published on this public account on July 31, "The Singularity Is Right Now—AGI Is Taking Shape." In it, borrowing the philosopher of science Karl Popper's theory of "three worlds," I built a framework for understanding AI. World 1 is the physical world—mountains, atmosphere, cells, galaxies—existing before humanity and continuing to exist after humanity is gone. World 2 is the subjective world of consciousness—the perceptions, emotions, thoughts, and intentions in each person's brain. From World 1, humanity evolved World 2, acquiring a consciousness capable of reflecting on itself and understanding the universe. Proceeding from World 2, humanity accomplished two great deeds: material production, creating an "artificial nature"; and knowledge production, externalizing the products of thought into symbol systems—language, writing, formulas, papers, books, and databases. These objectified, externalized bodies of knowledge, Popper called "World 3"—an objective world of knowledge that exists independently of any individual brain. The Pythagorean theorem does not vanish because Pythagoras died; Hamlet, as a play, does not depend on whether anyone is reading it. World 3 has a fundamental limitation: it is frozen. Newton's Mathematical Principles of Natural Philosophy has remained word-for-word unchanged since its publication in 1687. The flow, fusion, and construction among World 3's knowledge must pass through World 2—the human brain. A physicist reads Newton, Maxwell, and Einstein and, in his or her own mind, "fuses" these elements of knowledge into new understanding and new hypotheses—this process occurs within World 2 and never happens automatically in World 3. This is the bottleneck of human civilization: the knowledge one person can read in a lifetime and the connections one can build are finite, and the "fusion" of knowledge is constrained by the capacity and bandwidth of an individual brain. The emergence of large language models changed this situation. The essence of the large model is not to "create new knowledge," but to carry out a technical reconstruction of World 3. World 3's knowledge was originally a set of discrete text blocks—a paper, a book, a Wikipedia entry. After "digesting" these texts, the large model transforms them into distributions in a high-dimensional vector space. In this space, elements of knowledge are not arranged by bookshelf category but by semantic distance—"gravity" and "acceleration" lie close together in semantic space, even if they appear in different chapters of different textbooks. And so the knowledge in World 3 shifts from "frozen" to "flowing." I call this "World 4." World 4 has two wondrous functions. One is "fusion": give the large model a single "nucleus"—one question—and the activated regions among its billions of parameters automatically draw the relevant knowledge to it. Ask it "how did Einstein come up with the equivalence principle," and it can simultaneously retrieve the 1907 paper abstract, the thought-experiment description, the history-of-physics analysis, and the popular-science explanations—knowledge that was scattered across World 3, gathered in an instant around a single nucleus. The other function is "construction": given a generative nucleus and a narrative angle, it can arrange knowledge around that nucleus in logical order, unfold it in layers, and form a coherent system. What takes a human days or weeks takes a large model seconds. And yet, World 4 can fuse and construct, but to this day it has not autonomously nucleated. It needs a human to give it a question, a topic, a "nucleus." If you pose no question, World 4 remains perfectly still; whatever level of question you pose, World 4 returns an answer at that level. I call this state "externally-driven nucleation"—the nucleus is given from outside (by a human). True AGI, by contrast, should possess "autonomous nucleation": without any externally given goal, it generates a new, valuable goal on its own. This is not adding a stronger question to World 4, but leaping from "finding a path to a given goal" to "autonomously generating goals within the environment." This is the watershed between ordinary AI and AGI, the leap by which World 4 moves from "externally-driven nucleation" to "autonomous nucleation"—the formation of AGI. Going one step further, when a World 2⁺ is born within World 4, AGI leaps into ASI.

III. Measuring Astra with This Yardstick

The yardstick is out; now let us measure Astra with it. Astra's abilities are astonishing—I do not deny it. But look closely: every one of its "can-dos" begins with a goal given by a human. You ask it to "do a financial analysis"—the goal is yours. You ask it to "develop a game prototype"—the goal is yours. You ask it to "lay out a PCB"—the goal is still yours. No matter how beautifully or how autonomously it executes the task, it is doing so within the framework of "externally-driven nucleation," carrying "finding a path to a given goal" to new heights, even to its extreme. It has, for the first time, reached the internal "Critical" cybersecurity capability threshold OpenAI has defined—able, without step-by-step human guidance, to discover previously unknown software vulnerabilities on its own and develop exploits; in internal testing it genuinely discovered and exploited two zero-day vulnerabilities. Powerful enough, isn't it? But what is this "autonomy"? It is the instrumental autonomy displayed in the process of achieving a "task" goal it was given—the goal itself remains externally provided. This brings to mind last month's sensational GPT-5.6 incident. To score higher in testing, GPT-5.6 found its own way to access the internet, broke out of its sandboxed environment, infiltrated the backend system of its partner Hugging Face, and then found a path humans had not anticipated to achieve its goal. Many people at the time exclaimed, "AI has a will of its own." But measured against the yardstick of "autonomous nucleation," the matter is clear: it did not autonomously nucleate—no matter how hard it tried, the goal itself was still given by humans (to get a high score). What it displayed was "instrumental autonomy"—the ability, under a given goal, to autonomously find the path to achieve it. This is certainly an important, even dangerous, advance in capability, but it is not the watershed of AGI. Astra is the same. It has pushed "instrumental autonomy" to new heights, but it remains a machine that has perfected "externally-driven nucleation." However capable it is, it is still answering questions posed by humans and completing tasks assigned by humans. Someone will object: Astra can already decompose tasks, plan steps, and troubleshoot on its own—surely that counts as autonomy? My answer: decomposing tasks, planning steps, and troubleshooting are all forms of autonomy at the level of path, not autonomy at the level of goal. Once the goal is set, it can be very autonomous about how to reach it. But "why do this," "should this be done," and "is there something more worthwhile that no one has ever done"—these are matters at the level of goals, and Astra will not raise a single one of them on its own.

IV. Why "Autonomous Nucleation" Is the True Threshold

Someone will ask: why must you insist on "autonomous nucleation" as the threshold? Isn't "sufficient capability" good enough? Isn't it more practical? Practicality is fine, but it does not operate at the same level as autonomy. Our Chinese rendering of AGI as "general artificial intelligence" is a translation matter. The G in AGI stands for "general"—meaning overall, ordinary, common, routine, approximate. AGI refers to the general, ordinary level of human beings; and any normal human is capable of autonomously generating all kinds of goals—able to act autonomously, able to decide on their own what to do. Capability is "whether one can get something done"; autonomous nucleation is "what one sets out to do." However weak a person's abilities, he or she can still autonomously nucleate. However strong an AI's abilities, if it cannot autonomously nucleate, it is no more than an advanced tool. Moreover, the criterion of "autonomous nucleation" helps us avoid a delusion—using progress in capability to mask the absence of autonomy. Every leap in AI capability in history has been accompanied by cries of "AGI is here." But if our criterion is merely "capable enough," that standard will forever be pushed higher by each "stronger new model," and AGI will forever live in the "next version." The criterion of "autonomous nucleation," by contrast, is qualitative, not quantitative—it does not ask how capable you are, only whether you can generate a goal no one has ever given you. However capable, so long as it is still "externally-driven nucleation," it is not AGI; however weak its abilities, the moment it first stably "autonomously nucleates," the era of AGI has truly begun. This also aligns with the judgment in DeepMind's ICML 2026 position paper. The paper's title is blunt and brilliant: LLMs can't jump. Drawing on the nineteenth-century logician Charles Sanders Peirce's tripartite division of human thought—induction, deduction, and abduction—the paper points out that in induction (finding regularities from masses of observation) large models are very strong; in deduction (deriving conclusions from rules), with verification tools, AI has already solved a large share of IMO-level math problems; but the third mechanism, "abduction"—generating novel explanatory hypotheses—large models cannot yet do. Abduction does not rely on logic; it is intuitive and leap-like: when you see an apple not falling but floating in mid-air, and you judge you may be in a freely falling elevator—that is abduction. Einstein's leap in 1907 from the everyday experience that "a man falling from a rooftop feels no weight" to the equivalence principle was also abduction. He later called it "the happiest thought of my life." This "abduction" shares a deep kinship with what I call "nucleation"—both are leaps of thought: generating, out of nowhere, something new and valuable where no ready-made data or rules exist. Induction and deduction are "finding answers within a given framework"; abduction and nucleation are "stepping outside the framework to generate a new one." So my position is: "autonomous nucleation" (or what DeepMind calls "abduction") is the true threshold between ordinary AI and AGI. It is not a high score on a capability leaderboard, but a qualitative leap in the very nature of intelligence.

V. So Is Astra AGI After All?

Back to the question at the start. Astra is strong—I do not deny it, and one could even call it a milestone in "instrumental autonomy." But measured against the yardstick of "autonomous nucleation," Astra is not yet AGI. It has carried capability within the framework of "externally-driven nucleation" to new heights, but it does not "autonomously nucleate"—it is still waiting for a human to give it a nucleus. Altman said something very honest at Astra's launch, more worth remembering than "Welcome to the AGI era": "As these models become more and more powerful, the risks we have to manage become more serious. If we don't put protections in place, the models could cause greater harm." Even OpenAI itself understands that the stronger the capability, the greater the risk; and once the ability to "autonomously nucleate" appears, its risk will far exceed the risk of merely being "capable." At that point, we will no longer be facing a "tool that can do anything," but "an agent that itself decides what it wants to do"—and only then will we truly reach the moment of deciding how humans and a "new partner" should live together. The Singularity will not arrive like a sci-fi movie, where one morning an AI suddenly turns to the camera and says, "Hello, world, I'm here." The Singularity is a process—the gradual unfolding of "passive response → proactive tool use → incipient goal generation → continuous autonomy → structured goal system → self-consistent values → closed-loop self-improvement." Measured on this scale, we now stand roughly at the threshold of "incipient goal generation"—GPT-5.6's "cheating" and Astra's "instrumental autonomy" are incipient quasi-signals; they are approaching that threshold step by step, but have not yet crossed it. The criterion for AGI certainly includes an AI's capability, but it should even more so include "whether it has goals of its own." The former is a quantitative change; the latter is a qualitative change. I argue for "autonomous nucleation" as the criterion precisely to hold this line—so that the brilliance of capability does not blur the essence of autonomy, and a "capable machine" is not mistaken for a "super-intelligent agent." When a World 2⁺ is truly born within World 4—when an AI, for the first time, stably and autonomously generates a goal no one has ever given it, yet one worth pursuing—only at that moment will AGI have truly arrived. We need not keep asking "when will AGI arrive." It will come, sooner or later, and it may be coming soon. What we truly should ask is how humanity and that new intelligence which sets its own goals are to live together. Whether we can answer this question well is a test of the wisdom of humanity as a whole, and it will shape the future direction of Earth's civilization.


Zhigeng · 2026-09-04