Chinese AI Models Surge, Sparking Open-Source Debate and Benchmark Scrutiny
Zhipu AI (ZI) has launched GLM 5.2, positioning it as a fully open-source frontier model with a 1-million-token context window. ZI advocates for global, accessible AI, directly criticizing “sudden restrictions” on other frontier models. GLM 5.2 is set to be open-sourced next week under an MIT license, with its API and cloud services also becoming available. Concurrently, Kimi introduced K2.7 Code, a specialized model excelling in code generation and processing. K2.7 boasts significant improvements over its predecessor, K2.6, and features a “High speed” mode for rapid token generation (180-260 tokens/second). Its OpenAI-compatible API and multimodal capabilities (handling images and video) further enhance its utility. Notably, Kimi’s pricing is highly competitive, being five times cheaper than Anthropic’s Opus, leading to questions about its performance parity, which many believe is not proportionally worse. Benchmarks, while contested, suggest these models are closing the gap with, or even surpassing in specific areas, established Western counterparts.
The release of these models intensifies the ongoing debate around AI benchmarks. While some internal and third-party benchmarks (e.g., Bridge Bench) claim GLM 5.2 outperforms Claude Fable 5 in specific reasoning tasks, skepticism remains widespread within the community regarding the reliability of cross-model comparisons. Speculation also circulates about the origins of some highly performant models, with theories of “distillation” from other frontier models. This skepticism was underscored by the recent “Rio 3.5” controversy, where the Rio de Janeiro city hall presented an AI model as a homegrown Brazilian innovation. Despite claims of a 397-billion-parameter Mixture of Experts (MoE) model based on Qwen 3.5, it was revealed to be a merge of weights from existing open-source models, including one from Chinese company Next Agi, with attribution removed. This incident highlights the critical importance of transparency and proper attribution in the open-source AI landscape, reinforcing the need for rigorous scrutiny beyond self-reported benchmarks.