Overview
This post is a follow-up to Safe Pod Termination — Even an Alpaca Can Understand It.
Last time, we walked through what happens when a Pod terminates, based on how the various Kubernetes components work. A Deployment rolling update involves restarting Pods, and we found that the following two countermeasures are effective for handling requests without loss during that process:
- Sleeping in the preStop hook
- Implementing graceful shutdown in the application
In this post, we’ll actually perform a Deployment rolling update while live traffic is flowing, and check whether those countermeasures really work.
📘【note】 This post is day 18 of Kubernetes2 Advent Calendar 2020. Yesterday’s post was @south37’s 手を動かして学ぶコンテナ標準 - Container Runtime 編.
Two countermeasures for safe Pod termination
As mentioned above, there are broadly two things you can do to make Pod termination safe. Here’s an overview of each. For details on what’s happening under the hood, see the previous post.
Sleeping in the preStop hook
The preStop hook lets you define processing (a hook) that runs before a container is stopped.
If you sleep for a set duration in the preStop hook, the container waits that long before being stopped.
This lets you make sure that termination handling only begins after traffic has actually stopped being routed to the Pod (i.e., after it’s been removed from service).
The thing to watch out for here: there’s no dependency between service removal and the preStop hook — you can’t have the preStop hook wait to confirm service removal before it exits. Because of this, there’s no guarantee that container termination happens after service removal.
Application graceful shutdown
If you implement graceful shutdown in your application, you can guarantee that when the application starts shutting down, it finishes handling whatever requests it was already accepting before the process exits.
Note, though, that once the application starts shutting down, it can no longer accept new requests.
So, let’s run some experiments to see whether these two countermeasures actually help suppress errors during a rolling update.
Running the experiment!
Experiment flow
We deploy a test application to a Kubernetes cluster ahead of time and send it a steady stream of traffic.
While requests are being sent, we restart the Deployment (kubectl rollout restart).
In an actual production setting you’d more likely see a rolling update rather than a restart, but for the purposes of this experiment all we need is for Pods to actually stop and start, so we’re using kubectl rollout restart as a stand-in.
The application
I’ve prepared a sample application for this experiment. Here’s a rundown of it:
- https://github.com/hhiroshell/cowweb-go/tree/v1.1.1
- A sample application written in Go
- A startup flag lets you tune how much CPU load handling a single request generates1
- A startup flag lets you specify whether to perform graceful shutdown on exit
- The preStop hook is written into the Deployment manifest at deploy time
How CPU load is tuned
CPU load is generated by repeatedly generating random values.
You can set the loop count via a startup flag, e.g. l=640.
// c.load is the value given by the l flag
for i := 0; i < c.load; i++ {
for j := 0; j < c.load; j++ {
rand.Intn(len(cows))
}
}
How graceful shutdown is implemented
Graceful shutdown is implemented using Go’s standard http.Server.Shutdown().
This is also gated by a startup flag that specifies whether to shut down gracefully.
sig := make(chan os.Signal)
defer close(sig)
signal.Notify(sig, syscall.SIGTERM, os.Interrupt)
<-sig
if *shutdownGracefully {
if err := server.Shutdown(context.Background()); err != nil {
log.Print(err)
}
}
How the preStop sleep is configured
The preStop hook sleep is written into the Deployment manifest.
Here’s an example that runs sleep 5 as the preStop hook:
apiVersion: apps/v1
kind: Deployment
metadata:
name: cowweb
spec:
template:
spec:
containers:
- name: cowweb
lifecycle:
preStop:
exec:
command: ["sh", "-c", "sleep 5"]
Let’s try it!
Now let’s actually verify the effects of the preStop sleep and graceful shutdown.
1. Checking the effect of the preStop sleep
Let’s compare results across three conditions. The preStop sleep duration is varied across three values (0s, 5s, 8s), with everything else held constant.
| # | Condition 1-a (0s) | Condition 1-b (5s) | Condition 1-c (8s) |
|---|---|---|---|
| preStop sleep (s) | 0 | 5 | 8 |
| Graceful shutdown | true | true | true |
| Replica count | 8 | 8 | 8 |
| CPU load flag | l=640 | l=640 | l=640 |
| Max requests/sec (rps) | 200 | 200 | 200 |
Here are the results:
| # | Condition 1-a (0s) | Condition 1-b (5s) | Condition 1-c (8s) |
|---|---|---|---|
| Total requests | 24060 | 24060 | 24060 |
| 2xx responses | 23008 | 23900 | 24060 |
| 5xx responses | 1052 | 160 | 0 |
| Error rate | 4% | 1% | 0% |
| Avg response time (ms) | 23 | 29 | 29 |
With no preStop sleep (0s), 4% of requests errored out; with a 5s sleep that dropped to 1%, and with 8s it dropped to 0%. Looking at this alone, the preStop hook seems to suppress errors — but what’s actually happening on the Kubernetes and application side?
The diagram below depicts what happens to a Pod during termination as part of a rolling update, for a case where the preStop sleep isn’t long enough and errors occur as a result.
The key thing to notice: the preStop sleep finishes and SIGTERM is sent to the container before kube-proxy has updated iptables. That means the application starts shutting down while traffic is still being routed to it — i.e., before it’s been removed from service.
The Go application used in this experiment (and, presumably, many production applications) can’t accept new requests once shutdown processing has started, and returns an error as the response. Giving the preStop sleep enough time ensures that container shutdown only begins after service removal, which — as we saw with the 8s sleep — suppresses the errors.
2. Checking the effect of graceful shutdown
Next, to check the effect of graceful shutdown, let’s run an experiment with the following two conditions. Everything is held constant except whether graceful shutdown is enabled (the CPU load flag and max request rate differ from the previous experiment — more on why below).
| # | Condition 2-f | Condition 2-t |
|---|---|---|
| preStop sleep (s) | 8 | 8 |
| Graceful shutdown | false | true |
| Replica count | 8 | 8 |
| CPU load flag | l=1024 | l=1024 |
| Max requests/sec (rps) | 100 | 100 |
Here are the results:
| # | Condition 2-f | Condition 2-t |
|---|---|---|
| Total requests | 12060 | 12060 |
| 2xx responses | 11416 | 12060 |
| 5xx responses | 644 | 0 |
| Error rate | 5% | 0% |
| Avg response time (ms) | 47 | 48 |
Without graceful shutdown, about 5% of requests errored; with it enabled, that dropped to 0%. It looks like graceful shutdown also suppresses errors during a rolling update — so what’s happening in this case?
The diagram below shows what happens to a Pod at termination time in the case where graceful shutdown successfully suppresses errors.
Here, container termination only begins after kube-proxy has updated iptables (i.e., after service removal). At a glance this looks like there’s no problem at all, but that’s not quite right. Even though the Pod has been removed from service, requests it accepted before that point might still be in progress. If a non-graceful termination happens while such requests are still outstanding, their processing is forcibly cut off and the response becomes an error.
With graceful shutdown, the application process only exits after in-flight requests finish being handled, which suppresses errors as shown in the results above.
Discussion — applying this to a production environment
Through these experiments, we’ve seen that both the preStop sleep and graceful shutdown are effective at suppressing errors during a rolling update. So what else needs to be considered before using these in an actual production environment?
How long should the preStop sleep be?
For the preStop sleep duration, you need to consider how quickly service removal actually completes in practice. Service removal involves the endpoints-controller’s reconciliation process and kube-proxy’s iptables updates, and how long these take presumably depends on the scale of the cluster (number of Service resources, number of Pods under them, number of Nodes in the cluster, etc.).
Measure how long service removal actually takes in each of your environments, and set the preStop sleep to exceed that.
Is graceful shutdown necessary?
Whether you need graceful shutdown depends on the characteristics of your application. For example, if interrupting request processing could cause data inconsistency, you should implement graceful shutdown properly.
Choosing not to do graceful shutdown probably only makes sense in cases where it’s clearly fine for processing to be interrupted, or where there’s a specific reason you want the Pod to restart as quickly as possible.
📘【note】 You might think you could just make the preStop sleep long enough to wait until there are no in-flight requests left before shutting down — but I don’t think that’s reliable. Applications that depend on external services or a database can be affected by those external components in unpredictable ways. Request handling could end up taking longer than expected, and shutdown could begin before the response is returned.
Summary
In this post, we confirmed that the following two countermeasures are effective at preventing request loss during the Pod restarts caused by a Deployment rolling update:
- Sleeping in the preStop hook
- Implementing graceful shutdown in the application
To actually apply these in a production environment, you need to consider: for 1. the preStop sleep, the actual scale of your cluster; and for 2. graceful shutdown, the characteristics of your application.
That’s all — thanks for reading to the end!
This just changes the number of loop iterations internally, so it’s not something you can dial in with units like millicores. ↩︎