As Google's Gemini 3.5 Pro sits stalled by performance problems, Chinese artificial intelligence companies have rolled out a string of new models closing in on the industry's frontier. Moonshot AI released Kimi K3, a 2.8-trillion-parameter model, while Alibaba unveiled a preview of Qwen 3.8 Max, a 2.4-trillion-parameter flagship.

Kimi K3 in particular has stunned foreign outlets by matching much of the performance of Anthropic's Claude Fable 5 and OpenAI's GPT-5.6, and on some benchmarks it outperformed Sol, the top model in OpenAI's GPT-5.6 lineup. Reuters quoted an analyst as saying, "The idea that Chinese models are still six months behind the American frontier is losing credibility."

Moonshot AI is expected to release Kimi K3's model weights on the 27th. On the Intelligence Index, a composite ranking from the independent evaluator Artificial Analysis, Kimi K3 scored 57 points to place third — behind Claude Fable 5 in first and GPT-5.6 in second — edging out Google's Gemini 3.0 Pro. It also outscored Claude Opus 4.8, GPT-5.5 (xhigh), Claude Sonnet 5 and GLM-5.2.

Kimi K3 also performed strongly on agentic and coding tasks, ranking fourth on Artificial Analysis's Coding Agent Index, trailing only Sol and Terra from GPT-5.6 and Claude Fable 5.

Moonshot AI's Kimi K3 placed third on a composite A.I. benchmarking index, trailing the leading frontier models. (Source: AA)
Moonshot AI's Kimi K3 placed third on a composite A.I. benchmarking index, trailing the leading frontier models. (Source: AA)

Its lower ranking on the Coding Agent Index, compared with its third-place finish on the broader Intelligence Index, is attributed to a weak showing on the SWE-Atlas Q&A benchmark.

Artificial Analysis's Coding Agent Index is calculated as an average of three benchmarks — Terminal-Bench v2, DeepSWE and SWE-Atlas Q&A. Kimi K3 scored 84 percent and 64 percent, respectively, on the first two, placing near the top, but managed only 23 percent on SWE-Atlas Q&A, dragging down its overall average.

Even so, Kimi K3 outscored Claude Fable 5 on Terminal-Bench v2, and on DeepSWE it beat xAI's latest model, Grok 4.5, as well as Anthropic's Claude Opus 4.8. On the Frontend Code Arena, it topped GPT-5.6 to take first place.

Its agentic capabilities also drew strong marks. On the τ³-Banking test, Kimi K3 narrowly edged out GPT-5.6's Sol for first place. The benchmark measures a model's practical customer-service skills, simulating multiturn interactions in a retail banking environment to gauge how effectively it uses tools.

Kimi K3 also placed first on AutomationBench-AA, which evaluates SaaS workflow automation, and SpreadsheetBench 2, which tests spreadsheet tasks, outperforming frontier models on both. On BrowseComp, a benchmark for web search and information retrieval, it set a new record with a score of 91.2, surpassing the previous mark of 90.4 set by GPT-5.6's Sol.

On Artificial Analysis's Coding Agent Index, Kimi K3 ranked just behind Anthropic's Claude Fable 5. (Source: AA)
On Artificial Analysis's Coding Agent Index, Kimi K3 ranked just behind Anthropic's Claude Fable 5. (Source: AA)
Kimi K3 posted the next-best results after the GPT-5.6 lineup. The Terminal-Bench 2.1 benchmark measures how well A.I. agents perform complex, long-horizon tasks in a real command-line interface environment. (Source: AA)
Kimi K3 posted the next-best results after the GPT-5.6 lineup. The Terminal-Bench 2.1 benchmark measures how well A.I. agents perform complex, long-horizon tasks in a real command-line interface environment. (Source: AA)

It also finished second on AA-Briefcase, which evaluates long-horizon knowledge work by A.I. agents; CharXiv (RQ) w/Tool, a test of visual chart comprehension; and Zerobench w/Tool, a visual reasoning benchmark. On GDPval-AA v2, which assesses practical task performance across 44 professions, it ranked third.

Kimi K3 did rank lower, between sixth and eighth place, on several of the most demanding tests. On Humanity's Last Exam, it placed sixth, behind Google's Gemini 3.1 Pro Preview. On CritPt, which measures reasoning on unpublished physics research, it also placed sixth — behind four GPT models and Claude Fable 5 in fifth.

By Artificial Analysis's cost-per-task metric, however, Kimi K3 is notably expensive to run. That is partly because its weights have yet to be released, meaning it can currently only be accessed through Moonshot AI's own A.P.I., and partly because there is no way to turn off its reasoning mode — only a "max" setting is available, so users are billed even for the lengthy reasoning tokens generated before an answer.

Artificial Analysis's measurements also show that Kimi K3 tends toward verbosity. Across the Intelligence Index evaluation, it generated 130 million output tokens in total, more than double the field average of 63 million — meaning it produces far lengthier answers than other models even on identical problems, driving up costs.

Kimi K3 placed first on the τ³-Banking benchmark test, which simulates a retail banking environment to evaluate a model's practical customer-service skills. (Source: AA)
Kimi K3 placed first on the τ³-Banking benchmark test, which simulates a retail banking environment to evaluate a model's practical customer-service skills. (Source: AA)

Once Kimi K3 becomes fully open-weight and third-party hosts begin offering it, competition is likely to push prices down somewhat, though its 2.8 trillion parameters make it unlikely to become as cheap as DeepSeek, analysts say. The model uses a mixture-of-experts architecture that activates only 16 of its 896 experts at a time, meaning inference costs do not scale directly with its parameter count — but the G.P.U. memory and computing resources needed to run the service remain substantial.

◇ Alibaba unveils Qwen 3.8 Max preview, claiming performance "just behind Claude Fable 5"

The preview version of Alibaba's Qwen 3.8 Max, unveiled on July 19 at the World Artificial Intelligence Conference in Shanghai, has 2.4 trillion parameters. It is a vastly expanded multimodal model compared with its predecessor, capable of processing text, images, video and documents within a single system. Alibaba said it also plans to release Qwen 3.8 Max's model weights as open source in the near future.

Alibaba described Qwen 3.8 Max as a flagship model built for coding and A.I. agent tasks, claiming it ranks among the world's best in general performance — trailing only Claude Fable 5. The company emphasized the model's coding capabilities in particular, saying it made major gains in agentic coding, software development, front-end development and function calling. It also said the model's agent capabilities were optimized for tool use, complex workflows and multistep task automation.

Alongside the preview launch, Alibaba made the model available through its Token Plan, Qoder and QoderWork services. Because the announcement came on July 19, independent benchmarking sites such as Artificial Analysis have yet to publish formal evaluations. But community testing on platforms including the Aider Leaderboard, LiveBench and WebDev Arena suggests the model performs competitively with GPT-5.6 and Kimi K3 on front-end coding, agentic coding and web U.I. generation.

A performance comparison between Alibaba's Qwen 3.8 Max (preview) and Kimi K3. (Source: Stackpuff)
A performance comparison between Alibaba's Qwen 3.8 Max (preview) and Kimi K3. (Source: Stackpuff)

Expectations for the new version are also being fueled by an earlier assessment from Weights & Biases, an MLOps platform for developing, testing and evaluating A.I. models, which found that the previous version, Qwen 3.7 Max, ranked among the best performers on spreadsheet reasoning and workflow automation.

Qwen 3.8 Max also uses a mixture-of-experts architecture, though the number of active parameters has not been disclosed — a figure that is central to determining service costs, making it difficult to predict its eventual price. The previous version, 3.7 Max, was priced at $2.50 per million input tokens and $7.50 per million output tokens; with 3.8 Max's far larger scale, analysts expect prices to rise.

◇ A wave of powerful rival models heaps pressure on Google

The stream of new releases from China comes as reports emerge that Google's Gemini 3.5 Pro has been delayed by coding performance shortfalls, drawing comparisons between the two.

Gemini 3.5 Pro was originally expected to launch in June, but according to Bloomberg and other outlets, its performance on code generation and debugging fell short of Google's internal targets, potentially pushing the release back by several months. Google reportedly fine-tuned the model last month with additional coding data to address the shortfall, but the results were unsatisfactory, prompting the company to restart training from scratch, the reports said.

The pressure on Google has intensified since OpenAI's recent release of GPT-5.6, which sharply improved coding and cost efficiency, and as Meta and xAI rolled out their own flagship models — Muse Spark 1.1 and Grok 4.5, respectively — emphasizing coding and agentic capabilities. OpenAI, Anthropic and Google are typically regarded as the industry's three frontier labs, and if Google's next model fails to outperform those from Meta or xAI, the company risks losing its claim to frontier status.

With Chinese models also closing the gap, Google is seen as choosing to polish the model further rather than ship something unfinished. Google's shares fell nearly 4 percent following Bloomberg's report. A Google spokesperson said, "We're shipping models quickly while balancing performance and cost efficiency, and we continue to test Gemini 3.5 Pro and an upgraded Flash model with our partners."

◇ 58 percent of U.S. companies now run on Chinese A.I. models

Separately, traffic data from OpenRouter, an A.I. model routing platform, shows that Chinese models accounted for a record 58 percent of the A.I. tokens used by American companies in the first week of July, overtaking American models for the first time.

Bloomberg's analysis of OpenRouter's public traffic data found that the share of Chinese models reached 58 percent overall and peaked at 63 percent in a single week. In weekly token throughput, Chinese models processed roughly 18 trillion tokens, more than three times the roughly 5.5 trillion processed by American models.

DeepSeek was the most widely used model. Bloomberg, The Beijinger and other foreign outlets noted that the figure was below 10 percent as recently as early 2025 — meaning the balance has flipped in roughly a year.

OpenRouter offers unified access to A.I. models through a single A.P.I. key, allowing users to call more than 400 models from dozens of companies — including OpenAI, Anthropic, Google, Meta, DeepSeek and Alibaba — through one interface. Because the platform transparently tracks the weekly token consumption of the global A.I. startups and engineers who rely on it to build real systems, its traffic data is widely viewed as a genuine indicator of market-share trends.

Foreign media outlets have cited the shift as an example of how a stark price war has reshaped global A.I. usage in practice, with developers increasingly opting for Chinese open-weight models — often 10 to 50 times cheaper — over costlier American frontier models for coding automation and agentic workflows.

SNS 기사보내기
Hyun-seon Park
저작권자 © 디잇플러스 무단전재 및 재배포 금지