Skip to main content

How We Cut 60% of Our Kubernetes Bill by Turning Clusters Off at Night

In 2023, our company asked the Platform Engineering team to investigate infrastructure waste across our Kubernetes platform and identify opportunities to reduce costs.

At first, I assumed this would be a fairly standard optimization exercise. I expected to find a few oversized workloads, forgotten namespaces, and maybe some inefficient services.

I was wrong.

kubectl top nodes

The deeper I looked, the clearer it became that we had underestimated the long-term impact of our Kubernetes adoption model.

Over the years, development teams had been given significant freedom. That helped teams move quickly, but it also created operational patterns that became extremely expensive at scale.

I started finding things that genuinely surprised me:

  • Development namespaces with one LoadBalancer per pod
  • Containers requesting absurd amounts of CPU and memory
  • Environments permanently sized for peak traffic
  • Non-production clusters running 24/7 despite being unused outside office hours

At one point, while reviewing resource allocations, I joked with a colleague:

“Are we secretly contributing to the next moon landing?”

Some workloads were so oversized that they completely defeated the purpose of moving to containers in the first place.


Trying the “Collaborative” Approach​

Our first instinct was to work directly with development teams.

For weeks, I helped produce optimization reports showing where resources were being wasted and how much money could potentially be saved.

We proposed architectural improvements, better resource requests, and cleaner scaling strategies.

But almost every conversation ended the same way.

Some teams said the engineering effort was too expensive compared to the savings.

Others explained that their systems had evolved over years and that changing them would require major rewrites.

A few admitted they intentionally overprovisioned everything because it made them feel safer during traffic spikes.

None of those answers were unreasonable.

But after enough meetings, I realized something important:

We were asking developers to prioritize infrastructure efficiency over product delivery. Naturally, infrastructure optimization would always lose.

That realization completely changed how I approached the problem.


The Question That Changed Everything​

One afternoon, while reviewing cluster activity, I noticed something obvious that we had somehow ignored for years:

Why are development environments running all night if nobody is using them?

Most developers worked roughly eight hours a day.

That meant our non-production environments sat idle for nearly 16 hours every weekday, plus entire weekends.

At that moment, the problem stopped looking complicated.

We did not need perfect resource optimization to achieve massive savings.

We simply needed to stop running infrastructure nobody was using.


Turning Clusters Off at Night​

We introduced a new internal policy:

  • Non-production namespaces would automatically scale down outside working hours
  • Production and core infrastructure namespaces would remain untouched
  • Teams could request temporary exemptions for releases or urgent fixes

I built a lightweight platform layer to manage schedules and exclusions through namespace annotations and regex-based policies.

The implementation itself was surprisingly straightforward.

The results were not.

Within a few weeks, our FinOps reports started showing dramatic reductions in infrastructure costs.

As workloads disappeared overnight, the cluster autoscaler naturally removed unused worker nodes.

That was the moment I realized how much waste had been hidden in plain sight.


Going One Step Further​

The success of nightly shutdowns led to another question:

Even during working hours, are developers really using every environment at the same time?

The answer was obviously no.

So we completely changed the workflow.

Instead of keeping development environments running by default, we turned them into on-demand resources.

Now developers explicitly start the environments they need through an internal platform portal.

Once activated, an environment stays online for eight hours before automatically shutting down again.

This reduced idle infrastructure even further while keeping the developer experience simple.

Ironically, the hardest part was not scaling workloads down.

It was scaling them back up correctly.

Many applications depended on shared services, APIs, workers, or databases, so starting one environment often required orchestrating several others automatically.

Over time, I expanded the platform to understand these dependencies and handle them transparently.


What I Learned​

This project became one of the most impactful optimizations I have worked on.

Not because the technology was particularly advanced.

But because it reminded me that platform engineering is often less about building complex systems and more about shaping operational behavior.

For years, we focused on optimizing workloads individually while ignoring the simplest question of all:

Should this environment even be running right now?

Sometimes the biggest infrastructure optimization is not smarter autoscaling or better resource tuning.

Sometimes it is simply turning things off when nobody is using them.


Open Source​

We originally implemented this workflow using KubeDownscaler, an open-source project created in 2017.

Unfortunately, the project was discontinued shortly before we started this initiative.

Together with other adopters, we helped keep the ecosystem alive through a community fork before eventually rewriting the project from scratch in Go and extending it for large enterprise environments.

That project later became GoKubeDownscaler, now fully open-source and freely available to the community.