China’s AI competition is increasingly moving beyond model size and benchmark performance, with Alibaba and DeepSeek putting greater emphasis on the cost of running advanced systems.
Alibaba’s Qwen3.8-Max is its largest AI model so far, while DeepSeek’s V4-Flash is attracting attention for significantly lower inference prices than many competing models.
The developments highlight a growing focus across the AI industry: building models that are not only capable, but also affordable to operate at scale.
Alibaba expands Qwen’s scale
Qwen3.8-Max contains around 2.4 trillion parameters and uses a mixture-of-experts architecture that activates only a portion of the model for each request.
Alibaba says approximately 95 billion parameters are active at a time. By avoiding activation of the entire model for every request, the architecture can reduce computing requirements and response times.
DeepSeek uses a similar sparse approach with a considerably smaller model. Its V4-Flash model contains 284 billion total parameters, with around 13 billion active during inference.
Moonshot AI’s Kimi K3 sits closer to Alibaba’s model in overall size, containing 2.8 trillion parameters and approximately 104 billion active parameters.
Bigger models are not automatically more expensive
The number of parameters is only one factor in determining how much an AI model costs to operate.
Architecture, active parameter count, token consumption, and the number of times a model needs to be called to complete a task can all have a significant impact on the final cost.
Qwen3.8-Max supports text, image, and video inputs and offers a context window of up to one million tokens. Alibaba has also said the model completed a software engineering project over a 16-day period.
The model quickly became one of the highest-ranked Chinese text models on Arena.AI’s crowdsourced leaderboard, although several Anthropic models remained ahead in the overall rankings.
DeepSeek targets the cost of inference
DeepSeek is taking a different route with V4-Flash.
Instead of competing primarily through the largest possible model, the company is emphasizing inexpensive inference.
Artificial Analysis lists V4-Flash at $0.14 per million input tokens and $0.28 per million output tokens. The model has a one-million-token context window, 284 billion total parameters, and approximately 13 billion active parameters.
DeepSeek also offers substantially lower pricing for cached input. Artificial Analysis lists cache-hit pricing at $0.003 per million tokens for the Max Effort version of V4-Flash.
Cached input allows previously processed context to be reused, potentially reducing the amount of computation needed for repeated interactions.
Cost per task can matter more than API prices
Headline token prices do not tell the entire story.
A model with a low price per million tokens can still become expensive if it requires significantly more tokens or repeated interactions to complete a task.
Artificial Analysis estimated an average cost of around three cents per test for V4-Flash on one of its benchmark evaluations. By comparison, the estimated costs were 86 cents for Kimi K3, $1.86 for OpenAI’s GPT-5.6 Sol, and $3.15 for Anthropic’s Claude Fable 5.
The comparison accounts for the amount of input and output used by each model during testing.
This distinction is becoming increasingly important as AI moves toward agentic workloads, where models may make numerous calls, generate large amounts of output, and repeatedly interact with external tools.
Kimi K3 shows the other side of the equation
Moonshot AI’s Kimi K3 provides an example of why model capability and inference efficiency need to be evaluated together.
The model is listed at $3 per million input tokens and $15 per million output tokens, with cached input priced at $0.30 per million tokens.
On Artificial Analysis’ AA-Briefcase benchmark for agentic knowledge work, Kimi K3 averaged $10.57 per task. The model generated roughly 120,000 output tokens and used an average of 83 turns per task.
Those numbers demonstrate how quickly costs can accumulate when an AI system performs lengthy, multi-step workloads.
Kimi K3 nevertheless recorded one of the highest scores on the benchmark, showing that businesses may accept higher operating costs when a model provides stronger performance on demanding tasks.
Open weights change the deployment equation
Pricing is only one part of the competition.
Alibaba, DeepSeek, and Moonshot AI have also embraced open-weight models, giving developers the option of downloading model weights and deploying them through their own infrastructure or third-party providers.
DeepSeek V4-Flash is available as an open-weight model under an MIT licence, while Kimi K3 is also available with downloadable weights under Moonshot AI’s licence.
Open weights can reduce reliance on a single provider’s hosted API and give organizations greater control over deployment.
However, running a large model independently still requires significant computing infrastructure. The economic advantage therefore depends on factors including hardware costs, utilization rates, engineering resources, and the efficiency of the deployment.
Businesses may not need the biggest model
The shift toward cheaper inference also reflects changing enterprise requirements.
For many organizations, the most important question is not whether an AI model is the strongest available system, but whether it is capable enough to complete a specific business task at an acceptable cost.
That creates an opportunity for models that offer a balance of performance, affordability, transparency, and accessibility.
Open-weight systems can be particularly attractive in this environment because businesses have greater flexibility over where and how models are deployed.
China’s AI competition is becoming more cost-focused
Alibaba and DeepSeek illustrate two different strategies emerging within China’s AI ecosystem.
Alibaba is competing with an extremely large model that uses sparse activation to control the cost of inference. DeepSeek is pushing aggressively on low token prices while using a substantially smaller active parameter footprint.
Moonshot AI is pursuing another path with Kimi K3, combining a very large model with strong performance and open-weight availability.
Together, these approaches suggest that the next phase of AI competition may not be determined solely by who can build the largest or most capable model.
As companies deploy AI across increasingly large workloads, the cost of producing each useful result could become just as important as benchmark scores. The ability to deliver sufficient intelligence at a sustainable price may ultimately determine which models gain the widest adoption.
Source: https://www.artificialintelligence-news.com/news/china-ai-model-race-alibaba-deepseek-costs/


