hello
发布于 2026-09-28 / 3 阅读
0
0

数据库一抖就重启?Spring Boot 的 Liveness 别这样配

数据库一抖就重启?Spring Boot 的 Liveness 别这样配

Kubernetes 的 Liveness 失败会触发容器重启,Readiness 失败则让 Pod 暂时不接收 Service 流量。把数据库、缓存和远端 API 的同一套检查直接挂到两种探针上,可能在外部依赖故障时重启所有副本。Spring Boot 官方文档特别提醒:Liveness 不应依赖外部系统,Readiness 是否检查共享外部系统也需要按业务判断。依据见 Spring Boot Actuator 文档和 Kubernetes 探针指南。

本文以 Spring Boot 4.1.1 的 Actuator 文档为准,假设应用在容器内监听 8080,已有可运行的 Kubernetes Deployment。下面的 YAML 是配置示例,未在目标集群实际部署验证。

先提供两个健康端点

按官方启用说明添加 spring-boot-starter-actuator。在 Kubernetes 环境中,Boot 会把应用可用性状态映射到 /actuator/health/liveness 与 /actuator/health/readiness;若要在非 Kubernetes 环境本地调试,可设置:

management.endpoint.health.probes.enabled=true

在本地启动后用 curl -i http://localhost:8080/actuator/health/liveness 和对应的 /readiness 检查路径与 HTTP 状态。若项目自定义了 management.server.port、端点暴露策略或安全过滤链,先检查实际访问端口和权限;不要假设默认路径一定从应用主端口可访问。

把探针接到 Deployment

以下片段应放在目标容器的配置下;端口需与应用实际监听端口一致:

ports:
  - name: http
    containerPort: 8080
livenessProbe:
  httpGet:
    path: /actuator/health/liveness
    port: http
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 3
readinessProbe:
  httpGet:
    path: /actuator/health/readiness
    port: http
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 3

这些时间值仅是起点,不是所有应用的推荐阈值。应结合启动耗时、GC 暂停、节点负载和发布策略调整。如果应用启动很慢,可按 Kubernetes 官方说明增加 startupProbe,让 Liveness 与 Readiness 在启动探针成功后再开始执行。

避免“探针绿了,主服务却挂了”

若 Actuator 运行在独立管理端口,它可能在主应用端口不可用时仍返回健康。Spring Boot 提供 management.endpoint.health.probes.add-additional-paths=true,把 /livez 和 /readyz 加到主端口;此时探针路径也要同步改为这两个路径。这个开关及路径来自 Boot 官方探针章节。

上线前分别验证:正常启动时 Readiness 何时变为成功;停止或滚动更新时流量何时停止;数据库短暂故障是否造成预期外的容器重启。不要因为“数据库挂了”就机械地把它纳入 Liveness。若把共享数据库加入 Readiness,也要接受它故障时所有副本可能同时退出流量池的后果。


评论