GLM-5.3-Flash has become the most frequently used and most popular model on the B.AI platform, with cumulative token throughput surpassing 2.41 trillion. According to ChainCatcher, the model is the first native multimodal model in the GLM-5 series and has 320 billion total parameters and 18 billion active parameters.
It uses a hybrid architecture combining sparse and linear attention and supports a 1 million-token context window. Developers can still access the model for free on the B.AI platform for high-frequency APIs, coding, complex agents, and long-document processing.