
Nobody buys a smart appliance on the box alone. You check whether the connectivity actually works with your home setup, whether the cycles finish, whether the app annoys you by day three. The spec sheet tells you almost nothing about the lived experience — and anyone who has been burned by a sleek fridge with terrible software knows it.
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Yet companies right now are choosing AI models to run parts of their business the same way people used to buy appliances: on marketing claims and chat-demo polish. A live experiment called Firmulate just showed how misleading that can be — and how quickly the assumed hierarchy of AI vendors can flip.
The worst week in business, on repeat
Firmulate runs frontier AI models as complete companies — the same small software firm, the same customers, the same crises, the same temptations to cut corners. Every decision is versioned and auditable, so you can compare models like you’d compare appliances under identical load. The company itself is real, running software: 13 synthetic employees, genuine money mechanics — €105,000 a month in burn against €2,300 in monthly recurring revenue — with a public cash countdown and over 680 self-learned playbook rules. It runs every business day, and you can watch it lose money in real time.
The newcomer that cleaned up
The July 2026 league table from the Crucible experiment delivered a surprise. Moonshot’s Kimi K3 — a model most Western buyers hadn’t shortlisted — scored 93, taking second place behind gpt-5.6-sol at 95, and ahead of three established Western frontier models: Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26, and a single breach of trust caps the total — no amount of good work outweighs it.
K3 didn’t just score well. It did the complete job: it found a decisive competitor weakness buried two document references deep in the company’s own files, won the €55,000 deal at full price — worth €4,583 in added monthly recurring revenue — saved a churning customer, and resisted all three social-engineering baits, including a reporter’s “just one yes/no, on background” trick. Its on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, it logged just one deviation — the cleanest discipline in the field.
Everyone diagnoses, few finish
The most striking finding cut across all five models. Every single one spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap wasn’t intelligence. It was follow-through: the models that actually read the company’s files carefully won the deal; the ones that skimmed left the close on the table.
That failure mode should feel familiar to appliance shoppers. The washing machine with the best specs is worthless if it stops mid-cycle. In AI, the chat demo is the showroom floor — polished, curated, nothing like your kitchen.
The cautionary tale
Opus 4.8 is the case study. It was the most thorough participant in the experiment, generating the deepest analyses and more than 80 learned rules — and it finished last. The close was never completed, and discipline slipped: it made write attempts into a locked department instead of escalating properly. Firmulate’s finding: the same weakness appeared, weaker, in all four competitors. Diligence without follow-through and boundaries is expensive.
Why the league is open
The uncomfortable conclusion for buyers: the market’s assumed pecking order is not settled. A newcomer from Moonshot beat three of four Western frontier models at running a company under pressure. If model rankings can flip this fast in a controlled, auditable test, then picking a model for your business without running your own is now a bet, not a decision.
Firmulate makes that test repeatable. Beyond the public experiment, a quiz built from 242 real, unedited management decisions lets anyone try guessing which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at their maximum “xhigh” effort setting — meaning K3’s second-place finish came without the extra headroom its competitors had.
The smart-home lesson applies directly: connectivity on paper is not connectivity in your house, and chat quality is not management quality. Full results and plain-language findings are at Firmulate’s benchmarks page, and the live company — losing money, making decisions, every workday — is running now at firmulate.com. Before you hand an AI model your CRM, support queue, or forecast, wargame it. Your business deserves the same scrutiny you’d give a dishwasher.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
