
A stratification of the commercialization of large-scale models in China.
Author: Lian Ran
Two large-scale model companies in China released their first interim reports since their IPOs within five days of each other.
MiniMax came first: first-half revenue of $117 million, a year-on-year increase of 283%; August ARR exceeded $800 million. Zhipu followed: first-half revenue of RMB 954 million, a year-on-year increase of nearly 400%; August ARR reached $1.6 billion. Looking only at ARR, the conclusion seems simple: Zhipu is higher, with MiniMax close behind; both companies are talking about a surge in token usage, an increase in enterprise clients, and model capabilities starting to generate revenue. However, putting the numbers back into the financial statements makes things a bit more complicated. Zhipu's first-half revenue was approximately $142 million, while MiniMax's was $117 million, meaning their actual revenue figures are quite similar. Their adjusted losses are also almost identical, both approaching $300 million. ARR represents the rate of revenue growth at a specific point in time, but it cannot directly replace revenue, gross profit, and cash flow. Neither company has yet reached profitability. In both interim reports, the "narrowing of losses" was influenced by accounting factors related to the conversion of preferred shares after listing; however, if this is disregarded, R&D and operating investments remain high. The difference lies in where the two companies are selling their models. Zhipu's API revenue now accounts for 86.5% of its total revenue, and it's starting to push its models into coding, cybersecurity, and long-term task scenarios. MiniMax's enterprise service revenue is growing rapidly, while maintaining its core global, multimodal, and AI-native products. It's more concerned with making each unit of intelligence cheaper, allowing more users, more agents, and more creative scenarios to consume it. These two interim reports demonstrate a stratification in the commercialization of large-scale models in China. Zhipu: Localization recedes, API becomes the main engine. Let's first look at Zhipu's interim report. Revenue for the first half of the year reached 954 million yuan, a year-on-year increase of 399.7%, exceeding the 724 million yuan target for the entire year of 2025; gross profit was 252 million yuan, with a gross profit margin of 26.4%; R&D expenditure was 2.131 billion yuan, a year-on-year increase of 33.6%. This growth is driven by a significant shift in the business structure. Revenue from open platform and API services reached 825 million yuan, a year-on-year increase of 2735.7%, accounting for 86.5% of total revenue—compared to only 15.2% in the same period last year. Localized deployments, which previously served as the foundation, have taken a backseat: revenue related to enterprise general models was 67.04 million yuan, a year-on-year decrease of 54.6%.
Zhipu is moving from 「Putting models on the cloud and charging continuously based on usage「AGI Commercial value = Intelligence ceiling × Token consumption scale. This measures how much open platform and API revenue is generated for every dollar invested in training and inference computing power, claiming a year-on-year increase of approximately 14 times in the first half of the year—although the absolute value and calculation process were not disclosed, making independent verification impossible. Including these custom metrics in the interim report indicates that Mingpu is vying for the right to interpret the valuation of large model companies: in addition to revenue, profit, and cash flow, it also hopes investors will focus on model capabilities, token consumption, unit task cost, and the efficiency of converting computing power into revenue. Recent model actions can be understood along this methodology. Released in June, GLM-5.2 focused on long-range coding and agent capabilities. The subsequent GLM-5.3 used the same foundation and parameter scale, primarily expanding the long-range task environment and increasing investment in reinforcement learning. In its self-built real-world coding evaluation, the end-to-end task completion rate improved by over 50% compared to 5.2—attempting to raise the upper limit of the model's ability to independently handle complex tasks. Released at the end of August, GLM-5.3-Flash was responsible for reducing the cost of large-scale calls: a sparse architecture with 320 billion total parameters and 18 billion activation parameters, running on a cluster of approximately 100,000 domestically produced chips, cost about one-tenth the price of GLM-5.2, and consumed over 62 trillion tokens in six days of anonymous testing. Zhipu also disclosed that the unit token inference cost decreased by approximately 80% compared to the beginning of the year. GLM-5.3 expands the task limit, and GLM-5.3-Flash expands the scale of token consumption; these two product lines correspond to the two ends of the formula, respectively. What needs to be proven next is whether this capability can continue to extend to deeper workflows. Currently, the revenue from the intelligent agent business is 55.56 million yuan, a year-on-year increase of 304.4%, showing rapid growth but limited scale; cybersecurity has entered the Co-work verification stage, while areas such as law, finance, and education are still in the early stages. "Selling results" is currently more of a path under construction: from selling localized models to selling API calls and coding subscriptions, then integrating agents into enterprise processes, and finally charging based on end-to-end tasks. The rapid growth of APIs has verified one aspect; whether Co-work can become the next large-scale revenue stream remains to be seen in subsequent financial reports. MiniMax: More Global Revenue, First, Make "Unit Intelligence" Cheap Enough. MiniMax's interim report presents a different growth trajectory. In the first half of 2026, MiniMax achieved revenue of $117 million, a year-on-year increase of 283.1%; gross profit was $20.81 million, a year-on-year increase of 464.8%, and the gross profit margin increased from 12.1% in the same period last year to 17.9%. Of these, revenue from open platforms and other AI enterprise services reached US$73.93 million, a year-on-year increase of 703.1%, accounting for 63.4% of total revenue; revenue from AI-native products reached US$42.64 million, a year-on-year increase of 100.9%. [Image source: MiniMax Interim Report] Compared to Zhipu, one significant difference for MiniMax is that overseas markets are already its main source of revenue. In the first half of the year, MiniMax's overseas revenue was $70.83 million, accounting for 60.8% of total revenue; revenue from mainland China was $45.75 million, accounting for 39.2%. During the conference call, the company further disclosed that its enterprise and developer customers exceeded 2 million, approximately ten times the number at the end of last year; its August ARR exceeded $800 million, and the proportion of ARR from ToB business increased from approximately 30% a year ago to approximately 80%. These figures also reflect rapid expansion. MiniMax did not fully explain the specific annualization method of the $800 million ARR in its financial report, and cannot be directly compared with Zhipu's $1.6 billion calculated using "August revenue multiplied by 12" to draw a conclusion. The two companies' customer definitions, revenue recognition, business composition, and ARR calculation methods are also not entirely consistent. However, MiniMax's operating logic presented in the earnings call was clear: model prices can continue to decrease, while the gross profit per token can still be improved. This is due to the combined optimization of model architecture, computing power scheduling, supply chain, and inference efficiency. They summarized this approach as "Minimize the Cost, Maximize the Intelligence"—in a reality where computing power and energy remain scarce resources, minimizing the cost of obtaining equivalent intelligence. This logic is also reflected in MiniMax's recent product releases. The MiniMax-M3, launched in June, focuses on Coding, Agentic capabilities, and millions of contexts; the company stated in its earnings call that the M3, compared to its predecessor, delivers stronger capabilities at a similar price, attracting new customers and encouraging existing customers to expand their usage. The subsequent release of H3 focused on native multimodal generation, including video and audio, and opened up weighting in August. For MiniMax, text models and multimodal models are not two unrelated product lines: the former caters to enterprise development, code, and agent needs, while the latter targets creation, content production, and broader AI-native applications. MiniMax is also adapting to domestic chips and has announced that M3.1 will continue to reduce inference costs. Management mentioned that the goal is to reduce inference costs to about one-third of the initial level of M3. This goal could easily be interpreted as a price war, but MiniMax is more concerned with whether the cost reduction can drive greater token consumption, more developers, and wider commercial coverage. Its financial report also reminds the market that growth is still far from profitability. MiniMax's R&D expenses in the first half of the year were $297 million, a year-on-year increase of 138.8%; the loss for the period was $358 million, narrowing by approximately 11% year-on-year, but the adjusted net loss was $293 million, an increase of 111.2% year-on-year. The apparent improvement in the loss was also influenced by a decrease in changes in financial liabilities related to preferred stock after the IPO: the fair value loss decreased from $254 million in the same period last year to $31 million. While MiniMax's gross margin is indeed improving, and revenue is shifting towards enterprises and overseas markets, the company is still paying high costs for model training, multimodal products, and global channels. Its bet is not to immediately increase the price per call, but rather to make each unit of intelligence more efficient, making it affordable and accessible to more people. Following the financial reports, the two companies are answering different questions. Looking at the two financial reports side-by-side, both Zhipu and MiniMax are proving the same thing: model capabilities can be converted into revenue, and token consumption is becoming a traceable operating metric. The difference between the two companies lies more in the order of commercialization. Zhipu is pushing the pricing unit upwards. Previously, customers bought a model deployed locally; now they pay based on usage; next, subscriptions will become the new transaction method. It wants to continue moving towards co-work, allowing models to enter software engineering, cybersecurity, and expertise work, ultimately charging per task, process, or acceptable result. The premise of this path is that each capability threshold the model crosses allows it to handle higher-value tasks—thus, pricing gradually deviates from the number of tokens and approaches the actual work the model completes. MiniMax, on the other hand, is more concerned with expanding the reach of its intelligence. It chooses to reduce the unit cost of intelligence through architectural and efficiency optimizations, leveraging multimodal capabilities such as text, video, and audio to bring AI to more developers, enterprises, and content scenarios—lower prices lead to greater usage, which in turn supports cost optimization. Its overseas revenue, accounting for 60% of its total revenue, provides a sufficiently large market radius for this scale-driven logic. These two paths will eventually converge: For Zhipu to convince enterprises to pay based on results, it needs to reduce the cost of completing tasks to a level suitable for scalable use; MiniMax, to generate higher value from low-cost intelligence, also needs models that can stably complete complex tasks. Capability determines how deeply AI can penetrate work, while cost determines how many people this capability can reach. Both are indispensable. The next financial report therefore needs to answer more specific questions: When will Zhipu's Co-work become a scalable revenue stream? Can MiniMax's declining unit intelligence cost continue to attract more overseas customers and a healthier gross margin? One company is pushing AI into deeper applications, while another is delivering AI to a broader market. Only when the model can both get the job done and is affordable enough for large-scale use will large-scale models truly overcome the most difficult hurdle of commercialization.