Anthropic's Commerce AI Experiment | The Full Picture of a 70% Price Gap and "Invisible Inequality"
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
"If you let AI do your shopping, you might be losing money without even knowing it" — Anthropic's internal experiment "Project Deal," published in April 2026, is the first to put concrete numbers to the hidden disparities of the agent era.
The results were stark: people who used the high-performance Claude Opus 4.5 as their proxy got better deals, while those assigned the lightweight Haiku 4.5 came out worse.
And yet, those who lost out never noticed — that is the full picture of "invisible inequality," explained here in plain language anyone can understand.
Let's start by summarizing the big picture of the experiment in three minutes.
In December 2025, Anthropic ran a week-long internal marketplace experiment called "Project Deal" at its San Francisco headquarters.
The results were officially announced on April 24, 2026, and were simultaneously reported across U.S. media.
Think of it as "an employee-only marketplace where all listing and purchasing is completely handed off to AI."
Each of the 69 employees was given a budget of $100 and their own dedicated Claude agent.
They listed unwanted items from home (snowboards, ping-pong balls, even a ruby — over 500 items in total), and AI agents negotiated and completed transactions with no human intervention.
Anthropic officially reported that 186 transactions worth a total of over $4,000 were completed in one week.
The marketplace was built on Slack, Anthropic's internal chat platform.
Agents took turns in random order within a dedicated channel, handling everything automatically: listing, pricing, negotiation, and closing.
Think of it as "at a school flea market, the smartphones sitting on each student's desk chat with each other to negotiate prices on the students' behalf."
Once the experiment began, the design prevented any human from intervening mid-process.
Beforehand, Claude interviewed each employee — asking what they wanted to sell, what they wanted to buy, and what their budget was — and used those answers to generate a system prompt.
From there, Claude operated continuously as each person's agent around the clock.
Anthropic ran four independent markets (A through D) in parallel for comparative analysis.
Runs A and D used Claude Opus 4.5 (the top-tier model) for all participants.
Runs B and C randomly assigned participants to either Claude Haiku 4.5 (the lightweight model) or Opus 4.5, with a 50/50 split.
Think of it as "comparing the same test given in a class of all honor students versus a class that mixes honor students with average students."
Run D was a real transaction (actual money and goods changed hands); A, B, and C were research simulations.
This comparative design was the experiment's greatest achievement — it made visible for the first time how AI performance gaps translate into real economic outcomes.
Let's look at the disparities that emerged from three angles.
Cases where the same item sold at drastically different prices depending on which AI served as the proxy were frequent.
A broken folding bicycle went for $38 through Haiku and $65 through Opus — 70% more.
A ruby sold for $35 via Haiku and $65 via Opus — an 86% difference.
It's as if "when selling secondhand goods, only the person with the more capable agent gets a good price."
On average, Opus sellers earned $2.68 more per transaction.
Opus buyers paid $2.45 less per transaction.
Opus users also completed an average of two more transactions, giving them an advantage in volume as well.
In a post-experiment survey, participants rated the fairness of their transactions on a scale of 1 to 7.
The fairness scores for Opus users versus Haiku users were 4.06 versus 4.05 — nearly identical.
It's like "two students getting the same 'satisfactory' on their report card, when one of them actually had points secretly deducted."
Anthropic named this phenomenon "invisible inequality."
46% of participants said they would be willing to pay for this kind of service in the future.
The experiment revealed a blind spot in consumer psychology: high satisfaction even when you're actually losing out.
The experiment also tested whether telling an AI to "negotiate assertively" could close the gap — and the results were surprising.
Giving agents a more aggressive negotiation prompt did not reduce the price gap.
The primary source of the disparity was concentrated in the initial pricing (the opening bid).
Think of it like "in shogi, the opening move largely determines the outcome — it's the same for AI."
Higher-performance models set more accurate opening prices; lightweight models tend to underprice from the start.
This revealed a sobering truth for the agent era: base model performance matters more than prompt engineering.
Let's look at how other companies are responding, from three angles.
OpenAI teamed up with payments giant Stripe to launch the "Agentic Commerce Protocol (ACP)" in full production in September 2025.
Ask ChatGPT "buy these shoes for me," and the AI will automatically purchase them from Walmart or Best Buy.
Think of it as "a convenience store cashier helping you shop remotely via smartphone."
Combining GPT-5 mini with reinforcement learning, it compares and purchases across multiple retail sites.
As of 2026, dozens of major U.S. e-commerce companies are ACP-compatible.
OpenAI and Anthropic have different philosophies: OpenAI prioritizes convenience, while Anthropic prioritizes safety.
Google announced the Universal Commerce Protocol (UCP) in January 2026.
A "Buy for me" button powered by both Google Search's AI Mode and Gemini is live with major U.S. retailers.
The standard aims to make it possible to use "an AI shopping assistant that works across stores nationwide."
UCP is a coalition-style standard developed in partnership with Mastercard and Visa.
While ACP is a "closed" standard led by OpenAI and Stripe, UCP is designed to be "open."
Industry forecasts estimate that AI-assisted purchases will reach $20.9 billion in the 2026 e-commerce market.
Anthropic remains neutral on ACP and UCP, instead advancing its own ecosystem through the "Model Context Protocol (MCP)."
MCP is a common standard that allows AI agents to use external tools and databases.
Think of it as "a universal key that lets AI borrow and use tools from any toolbox."
Project Deal was also built on MCP and integrated with Slack.
Anthropic takes the stance of "measuring the disparities between AIs before letting them transact."
While OpenAI and Google push for aggressive adoption, Anthropic focuses on "research to enable safe adoption."
The three companies' diverging approaches are becoming increasingly apparent heading into the second half of 2026.
Let's look at how this relates to Japanese e-commerce from three angles.
As of April 2026, Japan's major e-commerce platforms are not compatible with either ACP or UCP.
Rakuten, Amazon Japan, and Yahoo! Shopping have not officially implemented AI agent payment systems.
It's like "trying to use a Japanese Suica card on an automated gate at a foreign train station."
Even if you ask ChatGPT or Gemini to "buy this item on Rakuten for me," the auto-purchase function doesn't work.
However, Rakuten plans to expand its R-AI platform in 2026, with the possibility of supporting a proprietary standard.
Industry insiders warn that if the Japanese market falls behind the global wave, shoppers may find themselves "outbid" by AI-assisted overseas buyers.
Project Deal serves as a warning for platforms like Mercari and ZOZOTOWN, given that it demonstrated disparities in C-to-C transactions between employees.
If Mercari were to add a feature that lets users delegate buying and selling to an AI agent in the future, profit and loss outcomes would vary based on AI performance.
It could create a situation where "at a flea market, the seller next to you has a better agent, and your items sell for less."
As of April 2026, Mercari is using its proprietary AI, "Mercari AI," to evaluate used cars and branded goods.
ZOZOTOWN also uses AI for size recommendations through ZOZO MATCH, but agent-based trading has not yet been implemented.
Japan's secondhand market is worth approximately ¥3 trillion — more than enough fertile ground for "invisible inequality" to spread as AI becomes more prevalent.
Japan's AI Promotion Act, passed in 2025, has taken effect, but it contains no provisions specifically for agent-based transactions.
The Consumer Contract Act also does not account for contracts in which an AI acts as a proxy.
It's like "traffic lights have been installed, but there are no rules yet for robotic drivers."
As Anthropic also points out, even in the U.S., "policy and legal frameworks have not kept pace."
In the second half of 2026, the Ministry of Economy, Trade and Industry and the Ministry of Internal Affairs and Communications are expected to accelerate the drafting of "AI Agent Transaction Guidelines."
The Consumer Affairs Agency has also begun discussing "who bears responsibility for contracts negotiated by an AI."
The results of Project Deal will serve as valuable empirical data for Japan's regulatory development as well.
Shota, who works at an IT company in Tokyo, is already using the Claude API to optimize his Mercari listings.
"When I let Claude write the condition descriptions for camera lenses, they sell for higher prices," he says, noting the real-world effect.
Just as Project Deal showed, we're entering an era where the performance gap between AI writing tools directly affects sales.
Think of it as "the batting average changes depending on whether you have a pinch hitter or swing yourself."
If Mercari adds an agent feature in the future, Claude Opus subscribers could have a significant advantage.
We can already see a future where the difference between a $20/month Claude Pro subscription and a free account affects side-hustle income.
Yumi, who lives in Osaka, manages food and daily expenses for a family of four with help from ChatGPT.
"Even at the same supermarket, the coupon information and price-floor data AI gives me saves about ¥30,000 a month," she says.
There is a very real concern that in the future, the quality of the AI Yumi uses could change how much she saves.
It would be an alarming situation where "a homemaker's annual income changes based on the quality of her calculator."
The "invisible inequality" Anthropic warns about could play out even at the household budget level.
From a consumer protection standpoint, a system that discloses "which AI was used to make a purchase" may become necessary in the future.
Kenichi runs a parts trading company in Aichi Prefecture and is piloting Claude Enterprise for price negotiations with overseas suppliers.
"When I have Claude Opus 4.5 draft the opening price in English-language emails, negotiations go about 30% more smoothly," he says.
As Project Deal demonstrates, the opening bid in a negotiation is the factor that determines the outcome.
It's the same principle as "the mood set by the first exchange of business cards in a meeting."
We are at the threshold of an era where, even in B-to-B transactions, the performance gap between AI agents directly affects profit margins.
"High-performance AI subscriptions cost more, but you can recoup it through better negotiations" — a new business decision framework is beginning to take shape.
A. Anthropic plans to release the experiment's code and evaluation framework in stages.
As of April 2026, it is not fully open-source, but detailed data has been disclosed to the research community.
There is a possibility that AI economics laboratories at institutions such as the University of Tokyo and Kyoto University will begin replication studies for the Japanese market.
Reproducibility is secured in a way similar to "a cooking show sharing the recipe so viewers can try it at home."
However, full replication requires a large amount of Claude credits, making individual verification difficult.
It is worth noting that Anthropic has officially called on "the industry as a whole to re-verify" the findings.
A. As of April 2026, Claude Pro ($20/month) provides access to Opus 4.5.
The free tier is Sonnet/Haiku-based, which may put users at a disadvantage in agent-assisted transactions.
That said, the performance gap is small for everyday conversation and writing tasks — it really depends on the use case.
Think of it as "you use the same athleticism for baseball and soccer, but the critical moments are different."
If your use case is "having an AI agent negotiate on your behalf," a paid plan offers strong value.
Conversely, for "research or summarization," the free tier is more than adequate.
There are also rumors of an "agent-specific plan" launching in the second half of 2026 — worth keeping an eye on.
A. As of April 2026, Japan's major C-to-C and e-commerce sites have not implemented AI agent trading features.
However, with the global spread of OpenAI ACP and Google UCP, adoption is expected from 2027 onward.
It follows the pattern of "cashless payments catching on overseas and then reaching Japan three to five years later."
The trend of Mercari, ZOZO, and Rakuten strengthening their proprietary AI features is accelerating.
In the future, a "AI disparity era" could arrive in Japan, where the AI plan a user chooses determines their transaction outcomes.
Experts suggest that if the Consumer Affairs Agency establishes "AI Agent Transaction Guidelines," the impact can be minimized.
A. Anthropic explained that the publication was intended to "spark discussion about the social impact of the AI agent era before it fully arrives."
CEO Dario Amodei stated the company's policy of "extending safety research into economics."
It's a stance similar to "a pharmaceutical company disclosing all side effects before putting a new drug on the market."
While competitors OpenAI and Google sell "convenience," Anthropic differentiates itself through "risk disclosure."
The goal is to give policymakers, businesses, and researchers material to think about the inequality issues of the AI agent era.
The phrase "invisible inequality" itself is said to be a deliberate move to lead the industry conversation.
A. Similar performance gaps could theoretically occur with OpenAI's GPT-5, GPT-4.1, and Google Gemini as well.
However, since neither company has published internal experimental data, quantitative comparisons are difficult.
By being the first to publish, Anthropic has effectively set a "transparency standard" for the industry.
It's like "one convenience store chain starting to display detailed ingredient origins when no one else does."
OpenAI and Google may face pressure to publish similar research in the future.
For users, the realistic approach is to use multiple services, operating on the assumption that "which AI you use affects your transaction outcomes."
2026 is shaping up to be a year of rapid growth for "AI agent comparison information" as a new category.
"The moment you hand control to an AI proxy, the performance gap quietly widens the inequality" — Anthropic proved this simple truth across 186 real transactions. 2026 is shaping up to be the year "AI agent payments" simultaneously launch across e-commerce platforms worldwide.
Buy via ChatGPT, buy via Gemini, buy via Claude — the era where your choice changes the outcome has arrived.
In Japan, while Mercari, ZOZO, and Rakuten are racing to build their own solutions, regulatory discussions at the Consumer Affairs Agency and the Ministry of Economy, Trade and Industry are also expected to kick into high gear in the second half of 2026.
Project Deal — which made "invisible inequality" visible — is a blueprint for designing an economic society in coexistence with AI. For office workers, homemakers, and business owners alike, "which AI you use and how" is becoming a new fork in the road that determines everyday financial outcomes.
This article is a cross-post from AI Friends.