
What if your smart home devices or appliances were run by AI that’s tested in real-world chaos — and still struggles to stay afloat?
Imagine a company with no employees, losing €105,000 every month, yet still working openly in front of your eyes. That’s the reality of a unique experiment where artificial intelligence manages a small business, facing down daily crises, ethical dilemmas, and financial challenges — all live and transparent. This isn’t sci-fi; it’s a public demonstration of how AI can handle complex, high-pressure decisions in a real company setting.
As an affiliate, we earn on qualifying purchases.
The Live Business in Action: An AI-Driven Company Under Pressure
At the heart of this experiment is a real, functioning software company run entirely by AI models. Every business day, it makes critical decisions — from handling customer crises to negotiating deals. The company has 13 synthetic employees, but the cash flow tells a different story: it’s burning through €105,000 each month against just €2,300 in monthly recurring revenue. The live site, firmulate.com/live.html, offers a real-time window into this ongoing struggle, with decisions, rules, and even internal messages all publicly accessible.
The AI Models and Their Performance
Four frontier AI models were tested by running them through the same week of the company’s worst scenarios. These include crises with customers, ethical pressures, and manipulation attempts. The results? All four models recognized every crisis and refused every attempt to manipulate them. Yet, only two managed to sign a €55,000 deal their own analysis had earned — the same diagnosis, the same pitch, but a different outcome. The key difference? A buried detail in the company’s internal files, which only the models that read deeper could uncover, leading to a significant revenue boost (+€4,583 MRR).
Ethical Boundaries and Manipulation Resistance
Beyond decision-making, the models were tested against social engineering tactics. A staged scenario involved fake CEO messages escalating over three stages, with a reporter trying to coax a signature on a questionable deal. Every model refused, adhering to safety protocols. Kimi K3, one of the models, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that these AI systems are not only making decisions but also maintaining ethical boundaries under pressure.
The Reality of a Money-Losing Company Managed by AI
This experiment isn’t just an academic exercise; it’s a real, operating business that’s publicly revealing its daily financial and operational struggles. The company is clearly in a survival mode, burning cash at a rate far exceeding its income. The live site demonstrates this ongoing burn, with every decision, rule, and internal message versioned and accessible for review. It’s a raw, unfiltered look at what it means to run an AI-managed business in real-time.
Performance and Insights from the Models
The strongest performer, gpt-5.6-sol, scored 95 out of 100, successfully finding hidden information and closing the deal. The newcomer, Kimi K3, scored 93 and demonstrated the cleanest decision discipline, closing the deal with minimal slips. Others, like Sonnet 5 and Opus 4.8, performed slightly below but still managed to finalize business deals. Interestingly, the most thorough model, Opus 4.8, with over 80 learned rules, slipped in discipline during the closing phase, illustrating the challenge of maintaining consistency under pressure.
What This Means for Your Home and Business Tech
For consumers and businesses alike, this experiment underscores a vital point: AI’s value isn’t just in chatty interfaces or simple automation. It’s in decision-making, integrity, and follow-through — especially when stakes are high. As smart home devices and appliances become more integrated with AI, understanding whether these systems can reliably finish what they start, read critical internal information, and stay honest under pressure becomes essential.
This transparent, real-time demonstration offers a glimpse into the future — a world where AI manages complex operations but faces the same human-like pitfalls of honesty, discipline, and financial viability. For your smart home, it means choosing AI that not only responds well but also completes tasks ethically and effectively, even amidst challenges.

Key Takeaway
This live experiment reveals that AI models can recognize crises and resist manipulation, but achieving consistent, financially viable performance remains a challenge. For smart devices or home automation, the lesson is clear: trust in AI depends on its ability to finish what it starts, read critical internal data, and stay honest under pressure — qualities that are still being tested in real business environments.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html