Mumbai: Operational complexity is emerging as a critical roadblock to scaling AI (artificial intelligence), even as enterprises accelerate AI adoption, according to a new study.
Nearly 1 in 20 AI requests fail in production as capacity limits become the primary bottleneck to scaling AI reliably, reveals new data from Datadog, the AI-powered observability and security platform. This implies that its not the model intelligence but operational complexity, hindering reliable AI at scale.
Datadog’s State of AI Engineering 2026 report, based on real-world data from thousands of organisations running AI in production, highlights a compounding complexity challenge as AI systems scale.
Nearly seven in ten companies (69%) now use three or more models alongside increasingly complex agent workflows. Around 5% of AI model requests fail in production, with nearly 60% of those failures caused by capacity limits – leading to slowdowns, errors, and broken experiences in AI-powered applications.
“AI is starting to look a lot like the early days of cloud,” said Yanbing Li, CPO – Datadog.
“The cloud made systems programmable but much more complex to manage. AI is now doing the same thing to the application layer. The companies that win won’t just build better models – they’ll build operational control around them. In this new era, AI observability becomes as essential as cloud observability was a decade ago,” added Li.
Speed Requires Control
Competitive pressure is accelerating AI deployment across startups and large enterprises alike. But as systems scale, speed without control creates risk. Failures are increasingly driven by system design, including fragmented workflows, excessive retries, and inefficient routing.
“The next wave of agent failures won’t be about what agents can’t do but what teams can’t observe,” said Guillermo Rauch, CEO – Vercel, the company behind Next.js and a platform for building AI-powered web applications.
“We built agentic infrastructure at Vercel because agents need the same production feedback loops as great software. Unlike traditional software, agents have control flow driven by the LLM itself, making observability not just useful, but essential,” added Rauch.
“Innovation alone isn’t enough,” added Li. “To scale AI with confidence, organisations need real-time visibility across the entire stack – from GPU utiliSation to model behaviour to agent workflows. Visibility and operational control are what allow teams to move fast without sacrificing reliability or governance. At scale, how you operate AI may matter more than the models you choose.”
Additional key findings:
- Multi-model is now the norm: OpenAI remains the most widely used provider at 63% share, alongside rising adoption of Google Gemini and Anthropic Claude which grew by 20 and 23 percentage points, respectively.
- Agent framework adoption doubled year-over-year, accelerating development but also introducing more moving parts into production systems.
- The amount of data sent to AI models per request is also rising: the average number of tokens more than doubled for ‘median use’ teams (50th percentile of usage volume) and quadrupled for heavy users (90th percentile).
