Developer Offshore research
Testing Kubernetes Probe State Transitions Before Offshore Release Support
· Research report
A controlled study separating startup, readiness, and liveness probe effects during slow start, dependency loss, overload, and recovery.
Use this report with the Research library and the related daily developer guides to turn evidence into a bounded work brief.
Key Stats
- 3 probe roles observed separately
- 5 controlled application states
- 1 synchronized pod-to-client timeline
Key Takeaways
- A probe result is a control signal, not direct proof of user availability.
- Test state transitions and recovery, not only a healthy steady state.
- Platform and application owners must agree on restart and traffic-removal semantics.
Decision and system boundary
The decision is whether one Kubernetes workload has probe behavior that supports a bounded release-support handoff. The study asks when startup, readiness, and liveness checks change Pod state; when endpoints receive or stop receiving traffic; whether a container restarts; and what a synthetic client experiences during slow start, dependency loss, overload, deadlock, and recovery. It does not ask whether probes are generally good. The unit is one versioned workload, manifest set, cluster release, service, traffic generator, and declared observation window.
This scope prevents a familiar category error: a successful probe is not the same as a successful user journey, and a failed probe is not always an application defect. Readiness can remove an endpoint without restarting it. Liveness can restart a container. Startup can postpone the other checks while initialization completes. Application owners define meaningful health semantics; platform owners approve thresholds and rollout behavior. An offshore developer can build the fixture, instrument transitions, and propose a correction without independently changing production.
Documented mechanisms and hypotheses
Kubernetes documentation distinguishes startup, readiness, and liveness probes and describes their effects. It also documents HTTP, TCP, command, and gRPC mechanisms plus timing and threshold fields. These are platform facts. They do not determine whether a particular endpoint checks the right thing or whether its timeout is appropriate. The primary hypothesis is that each configured probe corresponds to a deliberately named application state and produces the expected controller action without masking a separate user-facing failure.
A second hypothesis is that recovery is bounded and observable. Removing a Pod from ready endpoints should precede or accompany the sampled traffic response according to the declared objective, while a liveness restart should not create an endless crash loop. A third hypothesis is that external dependency loss should not automatically cause every replica to restart unless the application owner explicitly chooses that coupling. These are decision hypotheses to test; they are not universal recommendations for all workloads.
Cluster fixture and state controls
Use an isolated cluster with a three-replica synthetic service, a Service, pinned images, declared resource requests, and no customer traffic. Record node inventory, CNI and ingress path if used, deployment strategy, termination settings, topology, and every probe field. The application exposes separate test controls for initialization delay, event-loop or worker deadlock, local corruption, downstream unavailability, elevated latency, and recovery. Controls must be authenticated or available only inside the disposable fixture so they cannot become production attack surfaces.
Begin from a clean namespace for each scenario. Confirm image readiness and scheduling before introducing faults. Run no-probe and single-probe baselines before the combined configuration so effects can be attributed correctly. Add a deliberately wrong path or port as a seeded control that the evidence method must flag. Synchronize API server events, kubelet observations available through supported interfaces, application state changes, endpoint membership, restart counts, and client requests. Preserve manifests and image digests rather than relying on dashboard screenshots.
Transition matrix
For slow start, vary initialization within and beyond the declared startup allowance and observe whether liveness remains gated. For readiness, introduce a condition the application owner believes should stop new traffic, then measure endpoint removal and new versus established requests. For liveness, introduce a local unrecoverable stall and verify restart behavior, backoff, and return to readiness. For dependency loss and overload, compare the declared policy with actual readiness and restart outcomes rather than assuming one correct configuration.
Exercise recovery after every fault. A probe that detects failure but never returns the workload to service is incomplete evidence. Include rapid flapping near thresholds, one unhealthy replica, all replicas affected by the same downstream dependency, and a rollout occurring while one Pod is unready. Change one setting at a time and reset between runs. Define stop rules for sustained client failure, repeated restart backoff, loss of all ready endpoints, or a fault control that affects infrastructure outside the fixture.
Evidence and measurements
Retain fault injection time, probe request and result where safely observable, Pod conditions, readiness-gate state if present, endpoint slices, container state, restart reason, Kubernetes events, rollout status, and per-request client result on one timeline. Measure detection, endpoint removal, restart, initialization, and recovery intervals as fixture observations. State clock source and sampling interval. Do not collapse them into a single availability percentage or claim a service-level objective from a small synthetic sample.
Separate established connections from new connections because endpoint removal does not necessarily terminate existing sessions. Separate application response from ingress or service-routing behavior. Record whether other replicas had capacity to receive traffic and whether readiness reflected that capacity. An expected denied request from a fault-control endpoint is not a service outage; an HTTP 200 from a shallow probe is not proof that the target user operation works. The evidence table must preserve these distinctions for reviewers.
Failure modes and correction choices
Failure modes include using the same shallow endpoint for every probe, liveness depending on a shared external service, startup thresholds shorter than legitimate initialization, readiness that remains true during inability to serve, expensive probes that worsen overload, timeout settings unsupported by the handler, and successful recovery that never restores endpoint membership. Another failure is interpreting a restart as repair when the underlying dependency or configuration remains broken. Identify the mechanism before changing thresholds.
A correction may split endpoints, make a check local, reduce probe work, adjust declared timing, expose a genuine initialization state, or change application recovery. Threshold padding alone is not a substantive repair unless the owner’s timing requirement and observed distribution support it. Do not force a rollout, delete Pods, or bypass safety controls to obtain a green result. The developer proposes the smallest reversible change; application and platform owners decide semantics, capacity implications, rollout, and emergency exceptions.
Review and operational handoff
The handoff includes cluster and tool versions, manifest and image hashes, application-state model, scenario matrix, raw events, endpoint observations, client series, timing method, seeded-control result, proposed diff, stop conditions, and recovery commands for the disposable environment. The reviewer reproduces one readiness transition and one liveness restart, then confirms recovery and client behavior. A final green Pod list without the transition record is insufficient because it can hide flapping, delayed removal, or a restart loop that happened earlier.
Access is limited to the isolated namespace and approved observability. The developer must not gain production cluster authority merely to prepare the study. Application owners review endpoint semantics, platform owners review kubelet and rollout interactions, security reviews any diagnostic surface, and release owners set go/no-go rules. The asynchronous record names open uncertainty and next ownership. If cluster-specific evidence is unavailable, say so and narrow the conclusion instead of substituting documentation or a local cluster result.
Limits and conclusion
The study covers pinned versions, manifests, images, nodes, traffic, and fault controls. Managed load balancers, service meshes, autoscalers, disruption budgets, topology, resource pressure, and client retries may change production behavior. A disposable cluster cannot forecast every scheduling delay or node failure. Probe success does not validate business correctness, data consistency, authentication, or downstream service health. A small client series cannot establish long-term reliability or a universal threshold.
Pass requires that each probe has an owner-approved meaning, injected states produce the intended transition, user-facing observations are separately recorded, recovery is bounded, seeded misconfiguration is detected, and stop rules are documented. Conditional pass names scenarios requiring different controls or more environment evidence. Fail identifies whether the defect is state definition, probe mechanism, timing, routing, restart behavior, or capacity. The outcome is a release-support brief, not authority to deploy or a claim that Kubernetes guarantees application availability.
Sources checked October 2, 2026
Kubernetes documentation, Configure Liveness, Readiness and Startup Probes: https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/. Kubernetes documentation, Pod Lifecycle: https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/. Kubernetes documentation, EndpointSlices: https://kubernetes.io/docs/concepts/services-networking/endpoint-slices/. These primary sources define controller and API concepts used in the protocol.
The sources do not prescribe application-specific health meaning or validate a company’s production thresholds. Those decisions require the workload evidence and named owners described here. Record the documentation and cluster versions together because features and defaults may differ.
Evidence table
| Signal | What to inspect | Owner |
|---|---|---|
| Outcome | Acceptance evidence for the bounded task | Task reviewer |
| Control | Access, test, and approval boundary | Internal owner |
| Handoff | Open risks and next decision | Next owner |
Good distributed work is observable at the handoff: the result, evidence, limitations, and next owner are all explicit.
Frequently asked questions
Does a successful pilot authorize a production change?
No. It supports a bounded decision for the tested system and revision. The named internal owner still approves production access, rollout, exceptions, and accepted risk.
What should trigger a repeat?
Repeat the study when a relevant runtime, dependency, topology, policy, workload, integration, or operating assumption changes.