Designing AI Systems That Fail Honestly
The most dangerous AI systems do not crash. They keep responding while their evidence is stale, policies conflict, retries multiply, tool calls become expensive, or the system can no longer justify the commitments it is making.
In this hands-on workshop, participants redesign a failing AI-commerce workflow that appears technically healthy but is behaviorally unsafe. They work through conflicting inventory information, an ambiguous payment timeout, a stale return policy, and an exhausted AI tool-call budget. The objective is to design a system that remains useful without making promises it cannot defend.
Traditional reliability asks whether a service is up, fast, and error-free. AI systems introduce a higher bar: whether the system is still safe to trust. A dashboard can be green while customers receive incorrect answers, duplicate actions, unfair denials, or decisions nobody can explain later.
This workshop gives teams an architecture method for controlling that risk. Instead of treating the system as simply “working” or “down,” participants design modes that progressively reduce autonomy: answer normally, become conservative, defer the decision, freeze liability actions, or escalate to human review.
Session details:
Participants work in teams on a realistic commerce-agent scenario. The agent can answer customers, check inventory, interpret return policy, initiate refunds, and call external tools. Then the workshop injects failures:
Inventory data is eight minutes stale.
A payment call times out after the customer may already have been charged.
A retired return policy outranks the current policy.
Retries create duplicate side-effect risk.
The tool-call and cost budget is nearly exhausted.
APIs return successful responses, but the evidence is no longer reliable enough for a customer commitment.
For each failure, teams answer:
What promise has the business made to the customer?
What can the AI still safely show or say?
What evidence is admissible for this decision?
What must the system stop doing?
When should it defer, refuse, or escalate?
What proof is needed before automation can be re-enabled?
Participants use an architecture canvas to make the hidden controls explicit: customer promise, commitment boundary, admissible evidence, freshness rules, churn budgets, degraded modes, human-review triggers, and decision receipts.
About Rohit Bhardwaj
Rohit Bhardwaj is a Director of AI & Data Architecture at Salesforce, where he focuses on enterprise AI, agentic systems, cloud-native architecture, distributed systems, data platforms, security, and large-scale transformation.
Over his career, Rohit has designed and led complex enterprise platforms across AWS, Google Cloud, microservices, real-time data, API ecosystems, resilient distributed systems, and AI-enabled architectures. His work increasingly focuses on the challenges enterprises face as software evolves from deterministic services to AI-native and agentic systems—particularly around reliability, governance, evidence, security, observability, cost, and safe autonomy.
Rohit is the author of System Design with AI Interview Guide: Designing Scalable, Agentic, and Defensible Systems, published by Apress. The book presents a modern approach to system design covering scalability, distributed systems, AI architecture primitives, security, reliability, economics, agentic systems, and real-world architectures including e-commerce, ride sharing, payments, fraud detection, messaging, video streaming, file storage, and search. (Springer Link)
Book:
Amazon: https://a.co/d/09Zs1twa
Publisher / Springer Nature: https://link.springer.com/book/10.1007/979-8-8688-2782-2
O'Reilly: https://learning.oreilly.com/library/view/system-design-with/9798868827822/
Rohit is also an O’Reilly instructor and a frequent speaker at technology conferences including No Fluff Just Stuff, UberConf, GIDS, and other international events. His talks focus on practical architecture lessons from building and operating complex systems, including AI control planes, trusted agents, inference at scale, evidence-first RAG, AI security, distributed-system failure, and AI-era software architecture.
As a trusted advisor and architecture leader, Rohit works at the intersection of business strategy and deep technical architecture—helping teams translate complex business problems into scalable, resilient, secure, and economically sustainable systems.
Rohit holds an MBA in Corporate Entrepreneurship from Babson College and graduate-level education in Computer Science from Boston University and Harvard University.
Connect with Rohit:
LinkedIn: http://linkedin.com/in/rohit-bhardwaj-cloud
X / Twitter: @rbhardwaj1