Most Spring Boot services run on Kubernetes, and most Kubernetes incidents with Spring Boot services come from the same handful of settings: probes that restart healthy pods, shutdowns that drop in-flight requests, and memory limits the JVM doesn't respect. None of it is exotic. All of it shows up in support tickets as "random 502s during deployments" or "pods keep restarting."

Here's the checklist I'd apply to every Spring Boot service before it goes to production on Kubernetes.

1. Probes: liveness is not readiness

Spring Boot exposes dedicated probe endpoints — /actuator/health/liveness and /actuator/health/readiness — and in Spring Boot 4 they're enabled by default. Use them for what they mean:

Liveness answers "is this process broken beyond repair?" If it fails, Kubernetes restarts the container. It should never check the database or other services — a database blip would restart every pod at once and turn a small incident into an outage. Readiness answers "should this pod receive traffic right now?" It can include dependencies the service can't work without, because failing it just removes the pod from the load balancer.

Add a startup probe for the JVM's warm-up time, so liveness checks don't kill a pod that's still starting.

2. Graceful shutdown that actually drains traffic

Graceful shutdown is enabled by default in recent Spring Boot versions: on SIGTERM, the server stops accepting new requests and waits for in-flight ones, up to spring.lifecycle.timeout-per-shutdown-phase (30 seconds by default). But Kubernetes removes the pod from Service endpoints asynchronously — for a few seconds after SIGTERM, traffic may still arrive. A short preStop sleep covers that gap.

3. Memory: let the JVM size itself from the container

Modern JVMs read the container's memory limit. Use -XX:MaxRAMPercentage rather than a hard-coded -Xmx, and leave headroom: the heap isn't the only memory a JVM uses — metaspace, thread stacks, direct buffers and the code cache all live outside it. Around 70–75% for the heap is a sensible starting point. Setting memory requests equal to limits avoids surprise OOM kills when nodes come under pressure.

The configuration, all together

deployment.yaml (container excerpt)
containers: - name: loan-api image: registry.example.com/loan-api:1.8.0 env: - name: JAVA_TOOL_OPTIONS value: "-XX:MaxRAMPercentage=75" resources: requests: { cpu: "500m", memory: "1Gi" } limits: { memory: "1Gi" } startupProbe: httpGet: { path: /actuator/health/liveness, port: 8080 } periodSeconds: 5 failureThreshold: 30 # up to 150s to start livenessProbe: httpGet: { path: /actuator/health/liveness, port: 8080 } periodSeconds: 10 readinessProbe: httpGet: { path: /actuator/health/readiness, port: 8080 } periodSeconds: 5 lifecycle: preStop: sleep: { seconds: 10 } # let endpoint removal propagate terminationGracePeriodSeconds: 45 # > preStop + shutdown phase
application.properties
# Explicit, even where these match defaults, so intent is visible server.shutdown=graceful spring.lifecycle.timeout-per-shutdown-phase=30s management.endpoint.health.probes.enabled=true

The grace period must exceed the preStop sleep plus Spring's shutdown phase, or Kubernetes will kill the process mid-drain. The built-in sleep action also works with minimal, shell-less images, where exec: sleep would fail.

Four more things worth checking
CPU limits: a hard CPU limit can throttle the JVM badly during startup and GC; many teams set CPU requests but no CPU limit for Java services — decide deliberately. GC in small pods: since JDK 27, G1 is the default even in single-CPU containers, so re-check memory in small pods after upgrading. Actuator exposure: expose only health (and metrics on an internal port) — never the full actuator to the internet. Logs: write structured JSON to stdout with a correlation ID, so the platform can collect and search them.

For resizing a running pod's CPU and memory without a restart — and why the JVM makes that tricky — see in-place pod resize and the JVM. For the logging side, see correlation IDs and structured logging in Spring Boot.

Frequently asked questions

What should a Spring Boot liveness probe check?

Only whether the application process itself is healthy. It should not check databases or downstream services, because a dependency outage would cause Kubernetes to restart every pod. Dependency checks belong in the readiness probe.

How do I avoid dropped requests when a Spring Boot pod shuts down?

Enable graceful shutdown, add a short preStop sleep so Kubernetes can remove the pod from Service endpoints, and set terminationGracePeriodSeconds longer than the preStop sleep plus Spring's shutdown timeout.

How much memory should the JVM heap use in a Kubernetes container?

Use -XX:MaxRAMPercentage instead of a fixed -Xmx. Around 70–75% of the container limit is a common starting point, leaving room for metaspace, thread stacks, direct buffers and the code cache.

Liveness checks the process only; readiness can check dependencies.
Add a startup probe for JVM warm-up.
Graceful shutdown + preStop sleep + grace period = zero dropped requests on deploy.
Size the heap with MaxRAMPercentage, and keep memory requests equal to limits.