Z.ai released GLM-5.3-Flash on August 26, 2026, positioning it as the first natively multimodal entry in its GLM-5 model family, capable of processing text, images, and video within a single 1-million-token context window. The model ships with 320 billion total parameters but activates only about 18 billion per forward pass, a mixture-of-experts design that keeps inference costs low while still allowing a very large overall knowledge base. It's released under the MIT license, meaning developers can self-host, fine-tune, and redistribute it commercially without the usage restrictions that accompany many competing frontier-adjacent models. Notably, the release was preceded by an unusual soft launch: an anonymous model had been quietly running on third-party inference platforms since August 20 with a free tier and the same 1-million-token context window, quickly becoming one of the most-called models on those services before the community traced it back to Z.ai and the company confirmed it as GLM-5.3-Flash on the 26th. Z.ai says the model outperforms its predecessor GLM-5.2 on coding and agentic benchmarks at roughly one-tenth the price, and claims it approaches Claude Opus 4.8 on its internal coding benchmark, though such vendor-reported comparisons should be treated cautiously until independently reproduced. Pricing is listed at $0.15 per million input tokens, $0.03 for cached tokens, and $0.50 per million output tokens, with a 50% promotional discount running through September 9, 2026. For developers building agents or multimodal tooling, this matters because it lowers the cost floor for large-context, multimodal inference while keeping the weights open enough to run outside Z.ai's own API, a combination that puts pressure on both closed frontier labs and other open-weight multimodal releases to compete on price and licensing terms simultaneously.