Zhipu has launched GLM-5.3-FlashX, and its API and experience center are now open. According to PANews, the company said the model's maximum inference speed reaches 200 tokens per second.
The company said the speed increase is based on inference computing power provided by 100,000 domestic chips, along with additional infrastructure and inference optimization investment. Zhipu said its base model, GLM-5.3-Flash, was open-sourced on August 26, with 320 billion total parameters, 18 billion activated parameters, and a context window of 1 million tokens. It had previously been tested anonymously under the name "Ox Alpha."