China’s large AI models continued to dominate global usage last week, extending their lead over U.S. models for an 18th consecutive week.
According to the latest AI model usage data from OpenRouter, global AI model usage reached approximately 113 trillion tokens from August 24 to August 30, representing a 21.11% increase from the previous week.
Chinese AI models accounted for 55.16 trillion tokens during the same period, up 36.26% week over week. By comparison, U.S. AI models recorded around 17.07 trillion tokens, despite posting a much faster weekly growth rate of 76.71%.
Even with that surge from U.S. models, China maintained a substantial lead in total weekly usage and remained ahead for the 18th consecutive week.
Chinese Models Sweep the Top Three
The latest ranking also showed the strength of China’s AI ecosystem at the individual-model level.
The top three models by weekly usage were all from China, with an anonymous model known as Ox Alpha, or the “Niu Lai” model, moving into first place.
Ox Alpha recorded approximately 15.7 trillion tokens of usage during the week, representing a 36% increase compared with the previous week.
The model originally attracted attention because of its unusually strong performance and the limited information available about its developer.
Ox Alpha launched on August 20 with a context window of approximately 1.048 million tokens and a maximum output length of around 131,000 tokens. It also supports text, image and video inputs, giving it native multimodal capabilities. During its preview period, users could access the model free of charge.
Those specifications, combined with its rapidly increasing usage, quickly made Ox Alpha one of the most discussed anonymous AI models in the industry.
Zhipu Confirms Ox Alpha’s Identity
The mystery surrounding Ox Alpha lasted only a few days.
On the evening of August 26, Chinese AI company Zhipu publicly confirmed that Ox Alpha was actually GLM-5.3-Flash, a new member of its GLM family.
The announcement effectively ended speculation over who was behind the rapidly rising model.
Zhipu also revealed that GLM-5.3-Flash is the first natively multimodal model in the GLM-5 family. Unlike models that rely on separate processing components for different media types, its architecture is designed to handle text, images and video as part of its core multimodal capabilities.
Another major selling point is its pricing.
Zhipu positioned GLM-5.3-Flash at approximately one-tenth the regular price of GLM-5.3, with a limited promotional discount bringing the effective price down to roughly one-twentieth of the original GLM-5.3 pricing.
The combination of lower inference costs, multimodal capabilities and high performance appears to have helped the model gain traction quickly.

GLM-5.3-Flash Rises Rapidly
The growth of GLM-5.3-Flash was particularly notable because it entered the ranking only after its recent launch.
During its first week on the platform, the model climbed to sixth place, recording approximately 6.16 trillion tokens in weekly usage.
That performance is significant considering how little time the model had been available compared with more established competitors.
Its rise also highlights a broader change in the AI model market.
Usage is increasingly influenced not only by benchmark performance, but also by inference cost, context length, multimodal support and developer accessibility. A capable model that is inexpensive to call can potentially generate enormous usage even without having the strongest consumer brand recognition.
China’s AI Usage Continues to Expand
The latest figures also reveal a broader trend beyond the success of a single model.
Chinese AI models generated 55.16 trillion tokens of weekly usage, more than three times the volume recorded by U.S. models at 17.07 trillion tokens.
Although U.S. models grew faster on a week-over-week basis, China’s overall lead remains substantial.
The 18-week streak suggests that Chinese models have established a strong position in global API and developer usage, particularly as more models become available through platforms such as OpenRouter.
At the same time, the rapid growth rate of U.S. models shows that the competitive landscape remains highly dynamic.

What Ox Alpha’s Rise Means
Ox Alpha’s journey from an anonymous model to a publicly identified GLM release is arguably the most interesting part of the latest ranking.
Its rapid rise demonstrates how quickly a new AI model can attract developers when it combines competitive capabilities with aggressive pricing.
The model’s support for long contexts and multiple input formats also reflects where the AI industry is heading. This reflects the broader shift toward multimodal AI models, which can work with text, images and video within a unified workflow.
For Zhipu, the success of GLM-5.3-Flash could provide another opportunity to expand the global reach of the GLM family.
For the wider AI market, however, the numbers point to an increasingly intense competition between Chinese and U.S. models.
China’s 18-week lead in total weekly token usage is notable, but the rapidly changing individual-model rankings show that the race for AI usage is far from settled.








