The Hidden Performance Killer in Kubernetes: CPU Throttling and Latency Issues

The Hidden Performance Killer in Kubernetes: CPU Throttling and Latency Issues

Have you ever chased unexplained latency issues in Kubernetes?

You look at your dashboards and everything seems fine. CPU usage is hovering around 30–40%, the node has available capacity, and the pods appear healthy.

Yet your p95 and p99 latency values suddenly spike.

In cases like this, the problem may not be that your application does not have enough CPU.

The real issue may be that your application is not allowed to use CPU when it actually needs it.

One of the most common reasons behind this behavior is CPU throttling caused by Kubernetes CPU limits.

How Do CPU Limits Actually Work?

Let’s assume that you define the following resource configuration for a container:

resources:
  requests:
    cpu: "1"
  limits:
    cpu: "1"

Setting a 1 CPU limit does not necessarily mean that the container is permanently pinned to a single physical CPU core.

Under the hood, Linux uses cgroups and scheduler mechanisms to control how much CPU time a container is allowed to consume during a given period.

Let’s simplify the example.

Assume that the CPU quota period is 100 milliseconds.

If the container has a limit of 1 CPU, it can use roughly 100 ms of CPU time during that 100 ms period.

At first glance, this sounds perfectly reasonable.

However, things become more interesting when the application is highly multi-threaded.

Suppose you are running a Java service on the JVM and it uses 8 threads simultaneously.

If those threads consume CPU in parallel, they can theoretically burn through the available CPU quota very quickly.

For example:

CPU Limit        : 1 CPU
CPU Quota        : 100 ms
Active Threads   : 8

100 ms / 8 ≈ 12.5 ms

If the application consumes its CPU quota within the first portion of the scheduling period, the scheduler may prevent the container from using more CPU until the next quota period begins.

This is what we call CPU throttling.

This Is Where the Real Problem Begins

When a container exhausts its CPU quota, the application does not necessarily crash.

The pod may continue running.

Health checks may still pass.

CPU graphs may still look normal.

But because the application is periodically prevented from using CPU, incoming requests begin to wait.

As a result, you may start seeing:

  • Higher API response times
  • Growing request queues
  • Increased p95 and p99 latency
  • JVM threads waiting longer
  • Longer garbage collection pauses
  • Unexpected performance spikes

The most misleading part is that the average CPU utilization can still appear relatively low.

For example, Grafana may show:

CPU Usage: 35%

while the same service experiences:

p99 Latency: 2.5 seconds

The issue is not that the application is using all available CPU.

The issue is that it cannot access CPU at the moment it actually needs it.

How Can You Detect CPU Throttling?

If you are using Prometheus, one of the most useful metrics is:

container_cpu_cfs_throttled_periods_total

You should also look at:

container_cpu_cfs_periods_total

which represents the total number of CPU scheduling periods.

A simplified throttling ratio can be calculated as:

Throttling Ratio =
container_cpu_cfs_throttled_periods_total
/
container_cpu_cfs_periods_total

A practical PromQL example would be:

rate(container_cpu_cfs_throttled_periods_total[5m])
/
rate(container_cpu_cfs_periods_total[5m])

If this ratio starts increasing, it is a strong indication that your container is regularly being throttled because of its CPU limit.

However, relying on a single fixed threshold is not always the best approach.

For example, 1–2% throttling may be irrelevant for one application but may cause noticeable problems in a latency-sensitive API service.

That is why throttling should always be correlated with other metrics such as:

CPU Throttling
        +
p95 / p99 Latency
        +
Request Rate
        +
CPU Usage
        +
Application Thread Count
        +
GC Duration

If CPU throttling rises at the same time as p99 latency, CPU limits should definitely be part of your investigation.

Is Removing the CPU Limit the Solution?

For latency-sensitive workloads, one commonly used approach is to remove the CPU limit and configure only an appropriate CPU request.

For example:

resources:
  requests:
    cpu: "2"
    memory: "4Gi"
  limits:
    memory: "4Gi"

In this configuration, Kubernetes still knows that the pod requires 2 CPUs for scheduling purposes.

However, the application is not artificially capped at that exact CPU level.

If the node has spare CPU capacity, the application can temporarily burst above its request.

This can be especially useful for bursty workloads.

For example:

Normal CPU Usage
        ↓
       30%
        ↓
Traffic Spike
        ↓
Application temporarily uses 150–200% CPU
        ↓
Request queue is processed quickly
        ↓
CPU usage returns to normal

With a strict CPU limit in place, this temporary burst may never be possible.

But Removing CPU Limits Introduces Another Risk

There is an important caveat, especially for Java applications.

The JVM sizes several internal components based on the amount of CPU it believes is available.

These may include:

  • Garbage Collector threads
  • ForkJoinPool
  • Parallel Streams
  • Framework-specific thread pools

If the CPU limit is removed, the JVM may see far more processors than the application actually needs.

Imagine a small service running on a Kubernetes node with 64 CPU cores.

The service may realistically need only:

2 CPU

But if the JVM detects a much larger processor count, it may create more GC and worker threads than necessary.

This can increase CPU contention with other containers running on the same node.

In other words, while solving one throttling problem, you may accidentally create a noisy neighbor problem.

Using Active

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *