
Alibaba.com says its Accio AI agent completed 107 e-commerce tasks at over 50% lower cost than OpenAI's Codex and Anthropic's Claude Code, as the two firms clash over AI security.
Alibaba.com says its commerce-focused AI agent, Accio, can complete a broad range of e-commerce tasks at more than half the estimated cost of general-purpose rivals from OpenAI and Anthropic — a claim the Chinese platform is using to court small and medium-sized businesses worldwide.
The company announced the results at CoCreate, its annual event for entrepreneurs and small businesses, on September 9. In a 107-task benchmark evaluation, Alibaba.com said the estimated total cost of completing the full task set was $3.69 with Accio, compared with $9.27 for OpenAI's Codex and $9.51 for Anthropic's Claude Code. Alibaba.com described completion quality as comparable.
How the cost gap is achieved
Alibaba.com attributes the difference to optimizing the full system around real commerce work rather than relying on a single general-purpose model. The company says Accio uses commerce-specific data and workflows to post-train lightweight models for clearly defined tasks, keeping more capable models in reserve for deeper reasoning.
The system also routes each step of a complex request to the resources it needs, balancing quality, speed, cost and data requirements instead of defaulting to the cheapest or most powerful model. Cache reuse, context compression and coordinated agent execution further reduce repeated processing and redundant token consumption, according to the company.
"For small businesses, unaffordable AI is useless," said Kuo Zhang, President of Alibaba.com. "Our goal is not simply to make AI more powerful. It is to make commerce AI practical and affordable for even a one-person company."
A single workspace, and an open benchmark
Alibaba.com also said Accio has evolved into a unified workspace for global e-commerce operations, letting small businesses research markets, identify product opportunities, develop products, evaluate suppliers and manage daily operations. Through connections with Amazon, Shopify, eBay, TikTok Shop and Walmart, sellers can access supported storefront workflows from one place.
The company has open-sourced its Commerce Agent Bench on GitHub. Rather than synthetic exercises, the benchmark draws on real merchant activity — 10 million active SMB users, 1.6 million conversations and 200,000 execution traces — distilled into 107 end-to-end commerce tasks across seven categories and four levels of autonomy. The tasks include reviewing hundreds of unstructured emails, spotting payment fraud, calculating landed costs and booking multi-carrier shipping routes.
Alibaba.com noted that no single model led across the board, which it says reinforces the case for task-level routing — the idea that different tasks are best handled by different AI models rather than one model for everything.
Context: a wider dispute with Anthropic
The benchmark arrives amid a public and legal dispute between the two firms. In June 2026, Anthropic accused Alibaba of illicitly extracting its Claude AI model capabilities, describing it as the largest known attack of its kind on the company, and sought congressional help to crack down on the alleged activity. Alibaba has since barred employees from using Anthropic's Claude Code for work, placing it on a high-risk software list, after reports that the tool contained features that could identify users linked to China.
Against that backdrop, Alibaba.com's cost benchmark doubles as a competitive argument: that a commerce-specialized agent can undercut general-purpose coding tools on price while matching their output on industry tasks. The results are based on Alibaba.com's own evaluation, and the company has published the underlying benchmark to allow outside scrutiny.
Launched in 1999, Alibaba.com is a global business-to-business e-commerce platform serving buyers and suppliers from over 200 countries and regions. It is part of Alibaba International Digital Commerce Group.