
AI Agents built on the JVM look like microservices from the outside, but behave unlike anything you have operated before. Long-running LLM streams pin carrier threads, complex multi-step reasoning creates spiky heap allocation profiles, and high-frequency tool invocations expose latency vulnerabilities in standard thread pools and HTTP client configurations.
This talk provides a production survival guide for engineers running AI agent workloads on the JVM:
Attendees will leave with practical tuning parameters and monitoring strategies to keep JVM-based AI agents performant and reliable at scale.