Token Bucket: A Rate Limiting Algorithm
A limited resource is consumed in a controlled manner to keep the system stable and scalable.
We all have studied the Leaky Bucket technique while learning networking or system design. It’s one of the classic approaches used for traffic shaping and rate limiting.
But in real-world distributed systems, another algorithm is used far more frequently:
Token Bucket
Interestingly, the idea of “tokens” is something we hear a lot today in AI systems as well. LLMs process and charge based on tokens, and even AI APIs internally use token-based rate limiting to control traffic and prevent overload.
While the meaning of a token is different in networking and AI, the core idea remains surprisingly similar:
A limited resource is consumed in a controlled manner to keep the system stable and scalable.
And that’s exactly where the Token Bucket algorithm becomes so interesting.
The Problem With Rate Limiting
Imagine you own an API. Without rate limiting, a single user could Spam requests, Overload servers, Exhaust database connections, Cause cascading failures. So we need a mechanism to control request flow.
That’s where rate limiting algorithms come in.
Revisiting Leaky Bucket
The Leaky Bucket algorithm is based on a simple idea, Imagine a bucket with a tiny hole at the bottom. Water enters the bucket at different speeds, but leaves at a constant rate.
How It Maps to Requests
Incoming requests = water entering bucket Processed requests = water leaking out Overflow = rejected requests
The output traffic becomes smooth and constant.
And it is good for traffic shaping, smooth network flow, constant processing rate
The Main Limitation of Leaky Bucket
Leaky Bucket is too strict during bursts.
Suppose:
Your system can safely handle short spikes
A user suddenly sends 20 requests together
The average traffic is still reasonable
Leaky Bucket may still reject many requests because the outflow rate is fixed.
It smooths traffic aggressively, even when the system could temporarily tolerate bursts.
That’s where Token Bucket becomes much smarter.
Enter Token Bucket
Instead of controlling outgoing flow, Token Bucket controls permission using tokens. The idea is simple:
Tokens are added to a bucket at a fixed rate
Every request consumes one token
If tokens are available → request is allowed
If bucket is empty → request is rejected
Visualizing It
Bucket Capacity = 10 tokens
Refill Rate = 5 tokens/sec
If the user is idle for some time, the bucket fills up. Now suddenly the user sends 10 requests together. Since tokens were accumulated earlier, all 10 requests can be served instantly
This is the key difference.
Why Token Bucket Is Better
1. Allows Controlled Bursts
This is the biggest advantage. Real systems rarely receive perfectly smooth traffic. Traffic is naturally bursty:
Users refresh pages
Mobile apps sync data
APIs receive sudden spikes
Token Bucket handles this gracefully.
2. Better User Experience
Instead of rejecting temporary spikes immediately, the system tolerates short bursts. That means:
Fewer unnecessary failures
Smoother APIs
Better latency experience
3. More Flexible
You can tune Bucket capacity and Refill rate Separately.
Example:
Capacity = 100
Refill = 10/sec
This means:
Short bursts up to 100 allowed
Long-term rate capped at 10/sec
Very practical.
4. Easy to Scale Using Redis
Requests may hit multiple API servers simultaneously. So the available token count needs to be shared across all servers consistently. That’s where Redis helps.
A typical implementation stores:
available tokens
last refill timestamp
inside Redis for each user or API key.
Whenever a request arrives:
The server checks Redis
Calculates how many tokens should be refilled
Updates the token count
Consumes one token if available
Real-World Usage
Token Bucket is heavily used in:
API gateways
Cloud platforms
Routers
Distributed systems
Because modern systems care about:
burst handling
fairness
scalability
And Token Bucket balances all of them well.

