Most Spring Boot services run on Kubernetes, and most Kubernetes incidents with Spring Boot services come from the same handful of settings: probes that restart healthy pods, shutdowns that drop in-flight requests, and memory limits the JVM doesn't respect. None of it is exotic. All of it shows up in support tickets as "random 502s during deployments" or "pods keep restarting."
Here's the checklist I'd apply to every Spring Boot service before it goes to production on Kubernetes.
1. Probes: liveness is not readiness
Spring Boot exposes dedicated probe endpoints — /actuator/health/liveness and /actuator/health/readiness — and in Spring Boot 4 they're enabled by default. Use them for what they mean:
Liveness answers "is this process broken beyond repair?" If it fails, Kubernetes restarts the container. It should never check the database or other services — a database blip would restart every pod at once and turn a small incident into an outage. Readiness answers "should this pod receive traffic right now?" It can include dependencies the service can't work without, because failing it just removes the pod from the load balancer.
Add a startup probe for the JVM's warm-up time, so liveness checks don't kill a pod that's still starting.
2. Graceful shutdown that actually drains traffic
Graceful shutdown is enabled by default in recent Spring Boot versions: on SIGTERM, the server stops accepting new requests and waits for in-flight ones, up to spring.lifecycle.timeout-per-shutdown-phase (30 seconds by default). But Kubernetes removes the pod from Service endpoints asynchronously — for a few seconds after SIGTERM, traffic may still arrive. A short preStop sleep covers that gap.
3. Memory: let the JVM size itself from the container
Modern JVMs read the container's memory limit. Use -XX:MaxRAMPercentage rather than a hard-coded -Xmx, and leave headroom: the heap isn't the only memory a JVM uses — metaspace, thread stacks, direct buffers and the code cache all live outside it. Around 70–75% for the heap is a sensible starting point. Setting memory requests equal to limits avoids surprise OOM kills when nodes come under pressure.
The configuration, all together
The grace period must exceed the preStop sleep plus Spring's shutdown phase, or Kubernetes will kill the process mid-drain. The built-in sleep action also works with minimal, shell-less images, where exec: sleep would fail.
For resizing a running pod's CPU and memory without a restart — and why the JVM makes that tricky — see in-place pod resize and the JVM. For the logging side, see correlation IDs and structured logging in Spring Boot.
Frequently asked questions
What should a Spring Boot liveness probe check?
Only whether the application process itself is healthy. It should not check databases or downstream services, because a dependency outage would cause Kubernetes to restart every pod. Dependency checks belong in the readiness probe.
How do I avoid dropped requests when a Spring Boot pod shuts down?
Enable graceful shutdown, add a short preStop sleep so Kubernetes can remove the pod from Service endpoints, and set terminationGracePeriodSeconds longer than the preStop sleep plus Spring's shutdown timeout.
How much memory should the JVM heap use in a Kubernetes container?
Use -XX:MaxRAMPercentage instead of a fixed -Xmx. Around 70–75% of the container limit is a common starting point, leaving room for metaspace, thread stacks, direct buffers and the code cache.