A customer opened a support ticket reporting an unreachable control plane.

Jared Rodriguez and I investigated and found in the audit logs that an application was LISTing Secrets. Every time it listed Secrets, the API server would stop responding. Leader leases would be lost, pods would roll, and the whole problem would start over with a new API server replica.

Our first thought was that the customer’s software was hammering the API server with thousands of LIST requests per second. We thought we were looking for an application that was querying too often. But there was another part of the workload to account for: what were those requests reading?

The Secrets had been accumulating through automated deployments of a bad Helm chart. Each failed install and rollback cycle left another release Secret, and the customer had configured Helm to retain the last 200 releases. They were doing this in several namespaces. That gave us a collection of release data to investigate along with the application listing it. We needed to reproduce the failure to find out how much of it depended on the request rate and how much depended on the data being read.

Reproducing the failure

We reproduced the problem in a CAPA cluster in AWS with 200 objects containing 1 MB of data each. Secrets or ConfigMaps, either would do. Listing them made the API server stop responding. Overlapping requests made it worse, but with enough data a single request was enough.

So much for needing thousands of requests per second. We had been focused on how often the application was asking, but we needed to look at how much work each request was causing.

At this point we could reproduce the failure without Helm. It didn’t matter what was in the Secrets, and we could do the same thing with ConfigMaps. We needed to look at what the API server was doing when it read those objects from etcd.

200 objects shouldn’t be enough to take down a control plane. Once those objects were in the cache, requests that could use it wouldn’t need to read them from etcd again. The problem was getting them into the cache in the first place. At 1 MB each, that’s 200 MB of data to read, and those reads were timing out. The API server doesn’t de-duplicate these requests, so another list request arriving while the cache was still loading would start another read from etcd. And another. Each request made the reads slower, causing more timeouts, and the data never made it into the cache. The cache that should have relieved the load couldn’t finish loading because of that same load.

Restarting didn’t help. The data was still in etcd, and the new process had to read it before its cache would be useful. We couldn’t even use the API to remove the data because the API was what we’d just broken.

Following a LIST request to etcd

Digging into the code, we followed the LIST request from the API server to etcd to see why each request was starting another read instead of waiting for the cache to finish loading.

The generic registry’s ListPredicate passes the selectors, limit, and continuation token to storage. For a list of Secrets in one namespace, it reads under that namespace’s Secret prefix. A list across namespaces covers the resource’s whole prefix. A GET for one named Secret takes a different path; we’re interested in the collection read.

The next stop is Cacher.List. An empty resourceVersion sends the request straight to storage. Specifying resourceVersion="0" allows a cached response, but only if the cache is ready. Here’s what happens when it isn’t, with the comments omitted:

if listRV == 0 && !c.ready.check() {
    return c.storage.List(ctx, key, opts, listObj)
}

There’s no waiting for the initial load here. Each caller goes through that same check and starts its own storage request. While the cache is unavailable, adding readers adds work for etcd.

In the etcd backend, store.List builds the key range and calls KV.Get. It can limit the read:

if s.pagingEnabled && pred.Limit > 0 {
    paging = true
    options = append(options, clientv3.WithLimit(pred.Limit))
}

So pagination does reach etcd. But Limit counts objects, not bytes. If your page size is 500 and you have 200 large Secrets, you’ve still asked for all of them in one response. Without a positive limit, this path doesn’t supply an object-count limit to etcd at all.

A label selector doesn’t spare this storage path the read. Once the data comes back, appendListItem decodes each object before checking whether it matches. You might only get a few Secrets back, but the API server still had to read and decode the other candidates to find them.

Loading the cache

The API server loads its watch cache in pages of 10,000 objects:

storageWatchListPageSize = int64(10000)

Our 200 objects don’t even come close. All that data fits into the first page.

The cache constructor sets its reflector’s page size to that value. The reflector calls cacherListerWatcher.List, which requests everything under the resource prefix, passing the limit and continuation token through to storage. It needs to populate the cache, so it can’t just load the Secrets one application happens to want.

The client-go ListAndWatch code first gets the list and uses it to populate the local store. Then it watches for changes. The ordinary pager defaults to 500 objects, which would also fit all 200 of our objects. Even when the list spans several pages, List collects them before returning. Smaller pages reduce individual response sizes, but the consumer still ends up holding the collection.

Initial informer requests use resourceVersion="0". If the API server’s cache is ready, great. If it isn’t, they take the storage fallback above. And startCaching doesn’t mark that cache ready until the initial list has succeeded and replaced its contents. Until then, incoming requests can keep piling more work onto the same storage path.

Timeouts and recovery

Kubeadm configures controller-manager and scheduler to use the local API endpoint. Each talks to the API server on its own control-plane node. That matters when leadership moves between nodes.

While all this is happening, other control-plane components still need the API. Controller-manager, for example, has to renew its leader lock. If it can’t renew within its renewal deadline, it stops leading. Its callback is pretty final:

OnStoppedLeading: func() {
    klog.Fatalf("leaderelection lost")
},

When a controller-manager on another node takes over, its controllers and informers start loading through that node’s API server. If those reads hit the same problem, that API server can become unresponsive too. Each change in controller leadership can shift the load to another node’s API server and repeat the failure there.

The API server also has to answer the kubelet’s health probes. Kubeadm configures /readyz for readiness and /livez for liveness. Failing readiness marks the container unready; it doesn’t restart it. If the server stops answering long enough to fail its liveness probe, the kubelet restarts the container. Now that API server has an empty cache and has to read the data from etcd again.

Timing out a request doesn’t undo the work already spent reading and decoding its data. Cancellation can stop work that observes it, but other reads may still be running. Now the replacement processes need to initialize too. The normal recovery behavior puts more load on the service that’s failing.

Reporting it to the security team

We submitted what we considered a denial-of-service vulnerability to the Kubernetes security team. They declined to classify it as a CVE.

I disagree with that decision. A user doesn’t need cluster-admin privileges to cause this. Permission to create sufficiently large objects of a single resource type lets them build up a collection like the one we used. When an application or controller lists those objects, the read can take down the control plane for the whole cluster. The user doesn’t even have to issue the LIST themselves if an existing reader does it for them.

We did this with only 200 objects. The failure isn’t specific to Helm or to Secrets; ConfigMaps worked too. Allowing a user to create resources shouldn’t give them the ability to deny service to everyone else using the cluster.

Avoiding the same problem

Don’t use Secrets or ConfigMaps to store large amounts of data. Just because Kubernetes accepts an object doesn’t mean you can safely fill hundreds of them with megabytes of data. Splitting the data across more objects doesn’t solve this either; a list or cache initialization still has to load the collection. Keep the bulk data in storage intended for it, with the access controls it needs, and keep only the configuration or references your workload needs in Kubernetes.

If you’re automating Helm deployments, look at how much release history you’re keeping and what happens when a deployment keeps failing. In this customer’s case, the automation kept adding Secrets across several namespaces. Keeping less history and stopping the repeated install and rollback cycle would have limited the buildup. Also look at applications that repeatedly list Secrets. Backing off when requests time out avoids adding more reads while the control plane is already struggling. Smaller page sizes also reduce the data in each response, though a cache still has to collect all the pages.

And when you’re testing your control plane, test whether it can load your data with an empty cache. Being able to serve requests from a running cache doesn’t tell you whether it can recover after a restart. That’s where this problem caught us.