Technical Publications
Articles & Notes
Notes on building reliable AI agents, cloud reliability, cost optimization, and automated systems.
Why My Production AI Agents Kept Breaking at 3 AM (And How We Fixed It)
A practical guide on why autonomous AI agents fail in real-world production systems and how we built SRE reliability patterns, guardrails, and cost-efficient memory architectures to keep them running smoothly.
Key Engineering Takeaways
Unchecked agent memory was causing runaway latency and huge token costs. We fixed this with sliding-window memory and structured state checkpoints.
Agents would hallucinate API arguments and get stuck in infinite retry loops. We implemented strict Pydantic schemas and circuit breakers.
Integrated SRE health checks and ServiceNow automated alert routing to catch degraded agent sessions before users notice.
Reduced LLM token waste and API costs by over 40% with intelligent caching and local fallback routing.
Upcoming Articles & Case Studies
Designing Cost-Efficient Cloud Architectures for SRE Teams
How we analyze cloud utilization, right-size Azure and GCP workloads, and use automated scripts to prevent surprise cloud bills.
AgentPhased: Building Modular Rust and Python Agent Runtimes
A deep dive into AgentID, AgenTool, and AgentMem for long-term state retention and sandboxed tool calling.
AI-Powered ServiceNow Automation for Fast Incident Resolution
Using LLM pipelines to triage tickets, suggest remediation steps, and speed up incident resolution times.