
Managing a Python sidecar just to run an LLM feature alongside your Java service creates extra infrastructure complexity, security overhead, and failure points.
In this talk, learn how to ditch the translation layer entirely! Using modern OpenJDK features (Foreign Function & Memory API, Vector API) and libraries like jlama and TornadoVM, we can execute LLM inference directly on the JVM.
What we will cover:
Learn how to build self-contained, air-gapped, and ultra-fast on-device AI applications exclusively in Java.