Joshua Stancle runs Clean Saint out of Los Angeles. The product is a waterless oral-care film. The company is him. He uses AI for sourcing, marketing, web development, and customer support, and when we asked him what that felt like, he put it this way: “In a sense, there are ten of me.”
I keep coming back to the second half of that sentence. If there are ten of you, you need some way of knowing whether the other nine are getting the work right.
Most of the AI conversation still runs on a single question. Which model is the best? For a business, that is the wrong unit of measurement. A model reasons. Commerce means calling a supplier, filing a customs form, chasing a container that missed its vessel, handling a return, and a model does none of that by itself. Work gets done by an agent: the model that thinks, a harness that gives it tools and memory and the ability to act, and context that tells it what a good outcome looks like in a specific industry. At Alibaba.com, that context comes from 27 years of watching global commerce actually happen.
What commercial work asks of an agent
The industry has become very good at measuring intelligence. Benchmarks cover reasoning, coding, mathematics, factual recall, and increasingly tool use. But real commercial work is somewhat messier than any of that. Sourcing looks simple written down. In practice, it means comparing dozens of quotes, catching an inconsistency buried on the fourth page of a specification sheet, reading payment terms closely enough to notice when they have quietly changed, and confirming that a promised delivery date survives contact with your shipping schedule.
Product listings have the same texture. Attributes have to be right. Categories have to be right. The same listing may have to satisfy one set of regulatory requirements in Germany and a different set in California. Plausible output has very little value here. The job has to be finished, correctly, in the system where it lives.
Testing an AI agent only on what it says is like grading pilots on a written exam without asking them to land the plane.
Grade the outcome
That is why the Accio team at Alibaba.com, which offers an AI agent built for global commerce, developed a test that grades outcomes. CommerceAgentBench is open source and available on GitHub. It contains 107 end-to-end tasks pulled from real e-commerce operations across procurement, logistics, product listing, fulfillment, and after-sales service.
We assembled them from what we could see in our own data: 10 million active small-business users, 1.6 million real conversations, and 200,000 execution traces, sorted into seven categories of commercial work.
Grading happens on the end state. The listing either went live with the correct attributes or it did not. Freight moves on a route that exists, or it sits on a dock in Ningbo while somebody works out what went wrong.
Commerce has always kept score this way. A customer who receives the wrong product has no interest in how articulate the agent sounded when it placed the order. Execution is the benchmark that matters.
What we found
The strongest frontier model we tested successfully completed 61.7% of the tasks.
That figure is high enough to be useful and low enough to be a warning. Multi-step commercial work that sat beyond the reach of automation until recently now completes most of the time. But close to four in ten tasks still came back wrong.
The failures clustered in recognizable places. Agents struggled to spot a payment anomaly hiding inside a long supplier email thread, the kind of thing that reveals itself as fraud only after somebody has read all 300 messages. Landed cost gave them trouble once the calculation involved several moving variables at once. After-sales disputes broke down whenever the answer required reconciling documents that disagreed with each other. Multi-leg shipping routes were consistently hard.
Every one of those happens thousands of times a day in real businesses.
The risk changes as adoption scales. Across thousands of businesses using similar agents, individual mistakes could become correlated ones: inaccurate listings could multiply, fraud signals could be missed, and routing or compliance errors could ripple through supply chains. Measurement shows where automation is ready to scale, and where human oversight still needs to keep pace.
One result surprised me more than the headline number. No single model won. Leadership rotated by category. The model that ranked first on request-for-quote work and market research slipped behind on claims settlement and listing compliance, where a different model led. A third was strongest at publishing products and handling returns. A ranking built from general reasoning scores tells you very little about which system will perform on a particular commercial task, which is why the choice of model belongs to the job.
Precision delegation
For an individual business, that broader risk translates into a practical question:The question worth asking is narrower than the one the industry argues about. wWhich workflows can I hand over now, and which ones still need me? A benchmark that grades outcomes answers exactly that. Where the pass rates are high, supplier comparison and routine listing work can come off your desk. Where they are low, on unusual compliance questions and complicated negotiations and the exceptions that make up more of any commerce operation than anyone expects, keep a person in the loop and check the work.
I call this precision delegation. Once you know where an agent is dependable you can stop supervising it, and once you know where it breaks you can catch the failure before a customer does. Both save money. Neither is available without measurement.
Commerce needs a test like this and so does everything else. Logistics has its own edge cases, and so do finance, manufacturing, medicine, and legal services. Each field will need a benchmark built by people who understand what a bad outcome costs there, and those benchmarks should be open, so that a buyer can check a vendor’s claim against something.
Authority will move to agents one workflow at a time, as each one earns it. Joshua has ten of himself now. What he needs next is a way to know which of the ten he can stop checking.
The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.
This story was originally featured on Fortune.com

58 minutes ago
1


