
Author: Wang Jie
During the Summer Davos Forum held in Dalian, China at the end of June, leaders from the global artificial intelligence (including robotics) industry gathered to discuss the current development of the AI industry and the important trends to come.
Among them, Wang Jie, one of China's first-generation AI investors and co-director of the AI Economy Research Center at the Shenzhen Institute of Digital Economy, proposed that after going through three stages—"content generation," "reasoning ability," and "action ability"—the AI industry is about to enter the "real-world AI" stage. All aspects of the industry need to make corresponding preparations to welcome this stage. The following is the full text of "Artificial Intelligence Arrives at the 'Real-World AI' Stage," first published by Tencent Technology. We are at AI’s reality moment. In the past few years, AI has learned to generate, reason, and act. The next stage is not just about whether AI can provide more elegant answers on screen, but whether it can learn from real-world feedback and deliver acceptable and sustainable results in the real world. Today, we are at a "real-world moment" in the development of AI. Observation: AI is continuously moving away from the benchmark world. In the past few years, the main narrative of the AI industry has been organized around benchmarks. Each model release is accompanied by a set of scores: language understanding, professional exams, mathematical reasoning, code generation, software engineering, web page operation, multimodal question answering, and intelligent agent tasks. Rising scores excite the industry; score saturation leads to the creation of new benchmarks. Benchmarks have become the banners marking milestone after milestone on the long road of AI development. However, an increasingly clear fact is emerging: AI is continuously moving away from the benchmark world. Many tests once considered difficult enough to represent intelligence are being approached, matched, and surpassed by models time and again. Researchers continue to define new tasks, new leaderboards, and new evaluation sets, and models continue to chase and pluck new flags. This is certainly part of scientific progress, but it also illustrates that simple benchmarks are increasingly unable to bear the full significance of AI development. The benchmark world is essentially a "theoretical world": problems are predefined, answers have clear boundaries, evaluation criteria can be formalized, and the cost of failure is usually just a line of scores. It's suitable for proving a model possesses a certain capability, but it's not the same as proving that the model can deliver the expected results in real workflows. A model answering questions correctly in a question bank doesn't mean it can reliably complete tasks in enterprise procurement processes, hospital diagnostic collaboration, factory production scheduling systems, legal document risk review, or urban emergency response. Therefore, when we say AI is leaving the benchmark world, it doesn't mean benchmarks are no longer important. On the contrary, benchmarks remain a necessary dashboard for technological progress. But a dashboard is not a road, scores are not results, and demonstrations are not deliverables. Where is AI headed after leaving the benchmark world? The answer is: the real world. The entire industry is entering the "real-world AI" phase. The Leap from the "Theoretical World" to the "Real World" The Three Old Stages of the "Theoretical World" This round of AI development has already gone through three distinct old stages. The first is the "content generation" stage, typically exemplified by chatbots. For the first time, AI used natural language as its interface, enabling it to write, summarize, translate, converse, and explain, becoming a universal text tool for human cognitive labor. The second stage is the "reasoning ability" stage, typically represented by reasoners, such as GPT-01 and DeepSeek R1. AI begins to exhibit stronger decomposition, search, planning, proof, and self-checking abilities, capable of handling longer chains and more complex problems. The third stage is the "action ability" stage, typically represented by agents. AI no longer simply answers questions but invokes tools, browses web pages, writes code, operates software, and executes multi-step tasks. These three stages are crucial. Generation provides AI with language, reasoning provides AI with thinking, and intelligent agents provide AI with initial control. After generating, reasoning, and acting, the next step is not to perform more actions in demonstrations, but to bear the consequences in real-world environments. The real world will provide the environment for AI's long-term actions. Why are the above three stages considered "old stages"? Because they mostly remain in the "theoretical world" or "quasi-real world." The models address abstracted problems, not complete economic-social systems; they optimize computable feedback, not real results involving multiple actors, multiple constraints, and long cycles; they demonstrate potential capabilities, not work outcomes accepted by users, organizations, institutions, and the market. The New Stage of "Real World" We propose "Real-World AI" to reflect the new stage that AI is about to enter. The definition of real-world AI is: AI that can learn from real-world feedback, complete real-world tasks, and produce real results. Here, "real-world" has two meanings: First, the training feedback comes from results, users, systems, costs, and risks in real-world environments, not just from standard answers; second, the tasks come from real workflows, not just from question banks, sandboxes, or demos. It's not a vague label, but rather a stage name for AI's transition from capability demonstration to production delivery, from theoretical intelligence to operational intelligence. The core of real-world AI is not about adding more buttons to AI, but about putting AI into a closed loop: understanding real tasks, receiving real feedback, executing real actions, correcting its strategies, and ultimately delivering acceptable real results. It requires a breakthrough in model capabilities, currently mainly concentrated in "computer science" fields such as code, software engineering, mathematics, and cybersecurity, extending to broader human work scenarios: marketing, sales, supply chain, manufacturing, finance, law, healthcare, education, scientific research, public governance, and robots and automated systems in the physical world. The following is a key comparison between the real world and the theoretical world: In this sense, real-world AI is not a specific model, product, or algorithm, but a new direction for the entire industry. It will connect post-training, reinforcement learning, tool usage, memory systems, workflow integration, organizational feedback, human supervision, security mechanisms, and economic value measurement. The real world will become the new training ground for AI. Real-world AI will output real-world intelligence. Real-world intelligence is the model capability formed by AI after receiving feedback from the real world; it is also the ability to transform goals into results under real constraints. It measures not the instantaneous performance of the model on static problems, but the continuous availability, reliability, and value creation capability of the AI system in real tasks. If the core of benchmark intelligence is "whether it can get the correct answer to a given problem," then the core of real-world intelligence is "whether it can achieve an acceptable result in a real task." Why is the transition from the "theoretical world" to the "real world" inevitable? This leap is both technologically and economically inevitable. Technologically, large language models give AI language capabilities, reasoning models give AI stronger thinking abilities, and intelligent agents give AI preliminary action capabilities. Examining human behavior, after possessing language, thinking, and action capabilities, humans will inevitably enter a stage of interacting with the real world. Intelligence is not merely a capability confined to the mind, but the ability to achieve goals within an environment. Therefore, the next step for AI is also very clear: entering the real world. Economically, the greatest value of the AI revolution cannot remain forever in question-and-answer, writing, and code snippets. True productivity release comes from unlocking real-world tasks: a customer service process is automated end-to-end, legal due diligence is delivered reliably, a supply chain is dynamically optimized, a research hypothesis is rapidly validated, and a robot reliably collaborates in a warehouse or home. Only when AI enters real-world workflows will businesses factor it into organizational capabilities, society into productivity, and humanity truly feel the scale of this technological revolution. This is why "real-world AI" is more practical than simply discussing AGI. AGI asks whether AI is approaching human intelligence, while real-world AI asks whether AI can complete real-world tasks. AGI easily leads the discussion to unlimited capabilities, while real-world AI pulls the discussion back to feedback, results, costs, and value. It doesn't lower AI's goals, but rather places AI's goals where it must ultimately face: reality. Regarding roadmaps, OpenAI's five-stage roadmap proposed in 2024 generally captures the evolution from chatbots to reasoners to agents, but it doesn't fully describe the leap from the theoretical world to the real world. Furthermore, the latter two stages, innovator and organizer, focus more on the potential capabilities of an agent rather than on technological forms parallel to chatbots, reasoners, and agents; the standards are inconsistent. More importantly, this roadmap was proposed before the industry had truly entered the agent stage, naturally introducing uncertainty into the assessment of what comes after agents. At this juncture where the industry is moving from the theoretical world to the real world, we need a roadmap that can better guide long-term work. We propose the following five-stage framework: First, Foundation AI, the basic model stage, where AI acquires general representation and knowledge compression capabilities; Second, Generative AI, where AI acquires natural language and multimodal generation capabilities; Third, Reasoning AI, where AI acquires stronger search, planning, proof, and reflection capabilities; Fourth, Agentic AI, where AI acquires the ability to invoke tools, operate software, and execute steps; Fifth, Real-World AI, where AI enters real workflows, learns from real feedback, and delivers real results acceptable to humans, organizations, and institutions. This roadmap places "real-world AI" after the agent. The agent addresses the question of "can AI act?", while Real-World AI addresses the question of "whether AI's actions produce acceptable consequences." The agent is the interface, and the real world is the closed loop; the agent is the hand, and real-world AI is the organized working capability; the agent allows AI to enter the process, and real-world AI ensures that AI is accepted by the process, trusted by the organization, and measured economically. Further down the line, the industry may enter a larger phase: AI becomes the operational layer of the economy and society, which is the "digital layer" we have mentioned many times before. At that time, AI will not only complete individual tasks, but will participate in decision support, organizational coordination, resource allocation, scientific discovery, urban operation, and the operation of the physical world. Whether this future can arrive depends on whether we can overcome the hurdle of real-world AI today. Without real feedback, there is no real intelligence; without real results, there is no real productivity. In the past, we already had a large number of terms describing this round of AI development: AGI, ASI, Generative AI, Agentic AI, Embodied AI, Physical AI, etc. (World Model does not describe the characteristics of AI development, but rather a model roadmap). Generally speaking, these terms mostly stem from the perspective of algorithms, capabilities, or carriers, and can be called "descriptions from an algorithmic perspective." They are very important, but they can also easily lead industry discussions into abstract debates about "whether the model is smart enough," "whether intelligence is infinite," and "when it will surpass humans." A good name should possess a sense of direction: it not only describes what the technology is, but also reminds us where we ultimately want to go and where we are currently. "Real-world AI" possesses this sense of direction. It doesn't deny AGI, Physical AI, or Embodied AI, but rather changes the way questions are asked: it no longer just asks what AI is technically, but what AI can do in the economic and social sphere; it no longer just asks whether AI approaches human intelligence, but whether AI can reliably complete real-world tasks, create real value, and bear real consequences. "Real-world AI" also unifies the digital and physical worlds. In the digital world, real-world AI means AI entering enterprise software, knowledge work, transaction processes, R&D processes, and governance processes; in the physical world, real-world AI means robots, autonomous driving, smart manufacturing, home services, and urban infrastructure learning from real-world environments. Whether the platform is a browser, API, office software, robotic arm, vehicle, or humanoid robot, the core question remains the same: can AI form a closed loop in a real-world environment, complete tasks, and be accepted by reality? Therefore, we introduce the term "Real-World AI" to the entire industry. It brings researchers, entrepreneurs, investors, enterprise users, and policymakers onto the same map: moving from benchmark intelligence to real-world intelligence; from a capability demonstration phase to a task unlocking phase; from model competition to productivity competition; from "AI looks like it can do it" to "AI can actually do it." Real-world AI is not the end, but the entry point. It reminds us that the most important AI work in the coming years is not just about creating larger models, longer contexts, and more impressive demos, but about turning reality into a training cycle, feedback into capabilities, tasks into value, and AI into a truly usable productive force for human society. To truly bring this stage to fruition, the industry needs to reach a new consensus. Model training needs to use real-world workflow feedback as the core resource for post-training, rather than simply chasing existing leaderboards; AI applications need to evolve from assistant-like products to task delivery-like products, rather than just embedding AI chat windows into software; enterprise users need to shift AI evaluation from "ease of use" to "ability to reliably complete critical tasks"; investors need to re-evaluate task unlocking speed, feedback loop depth, and unit cost output beyond model parameters and demonstration effects; and policymakers need to establish data, accountability, security, and auditing frameworks to allow real-world adoption to expand with trust. This is the significance of the term "real-world AI." It consolidates a fragmented industry focus into a common direction: moving AI from the demonstration stage to the production floor; from question banks to organizations; from one-off answers to continuous feedback; from abstract intelligence to real value. We are at AI’s reality moment. The next frontier for AI is not another benchmark; the next frontier is the real world. The real world will become the new training ground for AI.