All posts
AI/ML October 9, 2026

The 3 AM Test: Why Agentic AI's Survival Depends on Infrastructure, Not Just Intelligence

TK

Thomas Kunnumpurath

VP Systems Engineering · Solace

The 3 AM Test: Why Agentic AI's Survival Depends on Infrastructure, Not Just Intelligence

A few months ago, I was speaking with a customer, a large airline, about their pilot agentic AI program. They had a sophisticated agent designed to optimize flight crew scheduling in real-time, reacting to delays, cancellations, and staffing changes. The demo was breathtaking. It showed the agent autonomously re-routing, re-assigning, and notifying. Then the head of operations leaned forward and asked, “Can it do that at 3 AM on a Tuesday, when our primary data center is having network issues, without waking anyone up?”

That question cuts right to the core of why the massive forecasts for agentic AI spending are clashing so brutally with the reality on the ground: only a minority of enterprises are actually scaling agents beyond the pilot phase. Everyone’s focused on model quality, prompt engineering, and fine-tuning. But the real bottleneck isn’t the agent’s intelligence; it’s its operational resilience. Most agents simply aren’t built to run autonomously at 3 AM.

The Unseen Architecture Tax of Agentic AI

I spent nearly a decade at Deutsche Bank, leading the migration from TIBCO Rendezvous to Solace for mission-critical trading systems. We were dealing with millions of messages per second, sub-millisecond latency, and systems that absolutely could not fail. The biggest fear wasn’t the cutover itself, but whether the new system would actually run reliably, day-in, day-out, without human intervention. We had to build for that 3 AM moment – the unexpected network partition, the overloaded consumer, the need to replay a day’s worth of trades for audit. This experience forged my conviction: for any system to truly scale and deliver value, its underlying infrastructure is paramount.

Now, as we work with enterprises deploying Solace Agent Mesh across diverse industries—from airlines to banking to manufacturing—I see history repeating itself. The allure of AI’s cognitive capabilities often overshadows the foundational engineering principles required for production-grade reliability. An LLM can be brilliant, but if its messages aren’t delivered, if it can’t find its peers dynamically, or if it overwhelms a downstream system, that brilliance is moot. The agents that are truly delivering value in production today aren’t necessarily the ‘smartest’ from an LLM perspective; they’re the ones built upon an infrastructure that can withstand the harsh realities of the enterprise.

Beyond Model Quality: The Five Pillars of 3 AM Survival

For an agent to pass the 3 AM test, it needs more than just a good brain. It needs a robust nervous system that handles the mundane, yet mission-critical, aspects of distributed computing:

  1. Guaranteed Delivery: Imagine an agent deciding to execute a critical trade or reschedule a flight. What happens if the message confirming that action gets lost? In real production, best-effort messaging isn’t an option. Agents need infrastructure that guarantees message delivery, even across unreliable networks or system failures. We solved this at Deutsche Bank with persistent messaging, ensuring no critical event was ever dropped, regardless of component status.
  2. Dynamic Discovery: Agents are inherently dynamic. They might spin up, migrate, or fail. How do they find each other? How do they discover the specific data streams or services they need? Hardcoding endpoints is a recipe for operational disaster. A robust agent infrastructure needs dynamic, topic-based routing and service discovery, allowing agents to find and communicate with relevant peers or data sources without manual configuration or downtime.
  3. Backpressure Mechanisms: An agent can generate decisions or requests at an incredible rate. What happens if the downstream system it’s communicating with – a legacy database, a human workflow, another agent – gets overwhelmed? Without backpressure, you get cascading failures. Production agents need mechanisms to slow down producers when consumers are struggling, preventing system overloads and maintaining stability. This was crucial for our real-time credit card controls at Capital One; we couldn’t just flood external systems.
  4. Replay and Auditability: The AI Act, like MiFID II before it, demands auditability. If an agent makes a critical decision, you need to understand its context, its inputs, and its chain of reasoning. This requires the ability to replay historical event streams. If an agent fails or acts unexpectedly, the ability to ‘rewind’ and inspect the exact sequence of events that led to its state is indispensable for debugging, compliance, and recovery.
  5. Observability: Beyond just logging LLM outputs, true observability for an agentic system means understanding its entire lifecycle: message flows, latency, resource consumption, and error states across all interactions. My experience building low-latency trading support tooling with KDB+ taught me that visibility into every single event is the only way to diagnose complex, high-throughput systems. You need to see the agent’s infrastructure health, not just its internal thought process. This is why when I built my ESP32 IoT demo with AI, I integrated full event tracing; the intelligence was only as good as the reliability of its sensor data stream.

The Real Work Starts Now

The exploration phase for agentic AI is over. The demos are impressive. The next challenge is scaling. It’s about moving from a cool PoC to a system that can run at 3 AM without a human getting paged. This shift has almost nothing to do with the quality of your LLM and almost everything to do with the robustness of the infrastructure underneath it – the delivery guarantees, the discovery, the backpressure, the replay, and the observability.

My advice? When evaluating agentic AI solutions, don’t just ask about the model. Ask about the messaging. Ask about the failure scenarios. Ask about how it handles network partitions, consumer slowdowns, and the need for compliance-grade audit trails. The agents that survive and thrive in production won’t be the ones with the highest benchmark scores on an isolated task, but the ones built on a foundation that can reliably deliver on its promises, day and night. That’s the real measure of production readiness.

TK

Thomas Kunnumpurath

VP of Systems Engineering at Solace

Share