Scaling APIs for Millions of AI-Driven Calls
AI agents are becoming a new class of API consumers. Unlike human users, agents can create bursty traffic, retry aggressively, call multiple tools in parallel, and accidentally amplify downstream failures. A single user request can become a large chain of API calls, model calls, vector searches, database lookups, and workflow events.
This talk explains how to design APIs for this new reality.
We will cover agent-aware rate limiting, budget-aware throttling, backpressure, load shedding, idempotency, deduplication, deterministic caching, async workflows, event-driven APIs, tail-latency SLOs, and cost observability.
Participants will learn how to tag and trace agent traffic, control runaway tool calls, prevent retry amplification, design graceful degradation, and build runbooks for cache storms, retry storms, dependency brownouts, and cost spikes.
The core message:
APIs exposed to AI agents must be contract-safe, retry-safe, cost-aware, observable, and degradation-ready.
Classic API scaling assumed relatively predictable traffic.
AI-driven API traffic is different because:
- One prompt can create many downstream API calls.
- Agents can retry, loop, and fan out.
- Tool-calling creates bursty and non-human traffic patterns.
- Cost grows with requests, retries, context size, model calls, and downstream work.
- Failures can amplify quickly across gateways, SDKs, queues, databases, and model APIs.
Agenda
- Why AI Changes API Scaling
Human traffic versus agent traffic, tool chains, fan-out, retries, and burst patterns. - New Failure Modes
Retry storms, cache-miss storms, malformed tool calls, version drift, DB saturation, and cost spikes. - Traffic Control for AI Agents
Agent-aware rate limits, per-tenant budgets, per-tool quotas, fair queuing, and adaptive backpressure. - Resilience Patterns
Idempotency keys, deduplication, bounded retries, circuit breakers, bulkheads, timeouts, and load shedding. - Caching for AI Workloads
Deterministic-result caching, semantic-aware caching, stale-while-revalidate, negative caching, and cache warming. - Async and Event-Driven APIs
Queue-first design, workflows, webhooks, streaming responses, outbox patterns, and dead-letter handling. - Observability and Cost Governance
Chain IDs, tool IDs, agent IDs, tail-latency SLOs, per-agent cost attribution, anomaly detection, and loop detection. - Runbooks and Readiness
Playbooks for retry storms, cache storms, provider brownouts, cost spikes, and safe degradation.
About Rohit Bhardwaj
Rohit Bhardwaj is a Director of AI & Data Architecture at Salesforce, where he focuses on enterprise AI, agentic systems, cloud-native architecture, distributed systems, data platforms, security, and large-scale transformation.
Over his career, Rohit has designed and led complex enterprise platforms across AWS, Google Cloud, microservices, real-time data, API ecosystems, resilient distributed systems, and AI-enabled architectures. His work increasingly focuses on the challenges enterprises face as software evolves from deterministic services to AI-native and agentic systems—particularly around reliability, governance, evidence, security, observability, cost, and safe autonomy.
Rohit is the author of System Design with AI Interview Guide: Designing Scalable, Agentic, and Defensible Systems, published by Apress. The book presents a modern approach to system design covering scalability, distributed systems, AI architecture primitives, security, reliability, economics, agentic systems, and real-world architectures including e-commerce, ride sharing, payments, fraud detection, messaging, video streaming, file storage, and search. (Springer Link)
Book:
Amazon: https://a.co/d/09Zs1twa
Publisher / Springer Nature: https://link.springer.com/book/10.1007/979-8-8688-2782-2
O'Reilly: https://learning.oreilly.com/library/view/system-design-with/9798868827822/
Rohit is also an O’Reilly instructor and a frequent speaker at technology conferences including No Fluff Just Stuff, UberConf, GIDS, and other international events. His talks focus on practical architecture lessons from building and operating complex systems, including AI control planes, trusted agents, inference at scale, evidence-first RAG, AI security, distributed-system failure, and AI-era software architecture.
As a trusted advisor and architecture leader, Rohit works at the intersection of business strategy and deep technical architecture—helping teams translate complex business problems into scalable, resilient, secure, and economically sustainable systems.
Rohit holds an MBA in Corporate Entrepreneurship from Babson College and graduate-level education in Computer Science from Boston University and Harvard University.
Connect with Rohit:
LinkedIn: http://linkedin.com/in/rohit-bhardwaj-cloud
X / Twitter: @rbhardwaj1