firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In the world of smart home management and connected appliances, reliability and trust are everything. Yet, as artificial intelligence (AI) models grow more sophisticated, a surprising pattern emerges: diligence alone doesn’t ensure success. Recent experiments reveal that even the most thorough AI can falter when it comes to closing the deal, highlighting the importance of prioritization over volume of effort.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Testing AI in a High-Stakes Business Simulation

Firmulate recently conducted a groundbreaking experiment: testing four leading AI models by running them through the worst week a small software company might face. This simulated environment included real crises, customer interactions, and the temptation to cut corners. Each model was tasked with managing the company, making decisions, and ultimately closing a €55,000 deal based on their analysis.

The Benchmarks and Results

  • gpt-5.6-sol scored the highest at 95, successfully uncovering critical information buried deep within company files and closing the deal.
  • Kimi K3, a newcomer, scored 93—demonstrating the cleanest discipline in refusing manipulation attempts and signing the deal.
  • Sonnet 5 scored 88, while a second Sonnet model scored 77, both closing the deal but with more slips in process discipline.
  • The baseline, a do-nothing approach, scored 26—highlighting how partial progress still significantly outperforms inaction.
Amazon

smart home device logs management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Diligence Isn’t Enough

Despite the detailed analyses and strict rule adherence, the most thorough AI model, Opus 4.8, finished last. It learned over 80 rules and performed deep analyses but ultimately left the close on the table. The reason? a lapse in discipline—decisions that should have escalated into proper procedures instead were written into a locked department, missing the final step.

Interestingly, the critical weakness was not in the obvious crisis moments but in reading and interpreting company documents. The models that examined internal files, rather than just external signals, secured the full deal at a premium of over €4,500 in monthly recurring revenue (MRR). This underscores an essential insight: in business, understanding the underlying context can outweigh sheer effort or volume of effort.

Resistance to Manipulation and Ethical Testing

Social engineering attempts—fake CEO messages escalating in stages and a reporter trick—were tested across all models. All refused manipulation attempts, with Kimi K3 explicitly treating suspicious requests as potential impersonation. This demonstrates that AI models are increasingly capable of withstanding unethical pressures, a critical factor for smart home systems where trust is paramount.

Amazon

AI-powered home automation system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Implication for Smart Devices

While these experiments focus on a fictional software company, the lessons resonate strongly for smart home technology. Whether managing complex appliances or automating routines, AI systems must do more than just respond well—they must finish what they start, read relevant internal information, and stay honest under pressure.

In practical terms, this means prioritizing the reading and understanding of internal data—like device logs or usage history—over simply responding quickly or covering more ground. For consumers and manufacturers alike, the question isn’t just about AI’s conversational abilities, but whether it can reliably execute tasks and uphold trustworthiness when it counts.

Amazon

smart home security system with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Diligence Alone Isn’t Enough

The experiment clearly shows that a model’s thoroughness doesn’t guarantee success. Opus 4.8, with its extensive rules and analysis, still failed to close the deal due to a discipline lapse. Similarly, in a home setting, an AI that meticulously responds to every prompt might still miss critical internal signals or neglect to escalate issues properly.

For companies deploying AI, the key takeaway is the importance of strategic prioritization—focusing on reading, understanding, and executing critical tasks—rather than just increasing effort or rule complexity. Trustworthiness and impact depend on these nuanced decisions, not volume alone.

Amazon

connected appliances with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watching AI in Action at Firmulate

Interested in seeing this in real time? Firmulate’s live platform allows you to observe AI models managing a simulated business environment, complete with crises, internal data, and decision-making processes. You can even run your own enterprise against a read-only export of your operations, testing AI’s readiness before deployment—no real systems are impacted.

Visit firmulate.com/live to watch the experiment unfold and explore how AI’s diligence and prioritization strategies impact results.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The recent experiment at Firmulate reveals that diligence, while admirable, does not guarantee success in AI decision-making. Prioritization, understanding internal data, and maintaining discipline are critical for trustworthy AI performance—less volume, more focus.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Dyson Cordless Vacuums for Cars (2026) — Guide 15

Discover the top Dyson cordless vacuums perfect for car cleaning in 2026. Our roundup highlights the best models, features, and tips for choosing the ideal one.

How to Refresh a Cedar Chest Smell Without Over‑Sanding

A simple, gentle approach can restore your cedar chest’s fresh scent without over-sanding, ensuring lasting freshness—discover how inside.

Trunk Packing for Airport Runs: Fit More Bags Safely

Inefficient trunk packing can waste space and compromise safety—discover expert tips to optimize every inch for your airport runs.

Pet Travel in the Trunk Area: Safety Rules You Can’t Ignore

I’m here to help you ensure your pet’s safety in the trunk—discover essential rules you can’t ignore for a secure, stress-free journey.