When you talk about system performance, two terms come up often.
- Latency
- Throughput
Both relate to "speed," so they are easy to confuse at first.
In practice, though, they measure different things.
Put very roughly:
Latency = how long a single operation takes
Throughput = how much can be processed in a given amount of time
Once you can tell the two apart, discussions about performance improvement become much easier to organize.
What Is Latency
Latency is the time from when an operation starts until the result comes back.
For an API, for example, the flow looks like this:
Request -> Server -> Response
The time from sending the Request to receiving the Response is the latency.
Suppose you have:
GET /users/1
Latency: 100ms
That means the API returns a response in about 100ms.
So it helps to think of latency as:
how long you wait for a single operation
In a web app, it also relates to the time from when a user presses a button until the result appears on screen.
The higher the latency, the more likely users are to feel that "this is slow."
What Is Throughput
Throughput, on the other hand, measures:
how many operations can be completed in a given amount of time
If an API server can handle 1,000 requests per second, then:
Throughput: 1,000 requests/sec
The same applies beyond APIs.
For message processing:
50,000 messages/sec
For image processing:
500 images/sec
So throughput is about:
how much volume the system as a whole can handle
Latency vs. Throughput
Summarized, the difference is quite simple.
| Metric | What it measures |
|---|---|
| Latency | Time taken by a single operation |
| Throughput | Amount processed per unit of time |
Suppose an API server has:
Latency: 100ms
Throughput: 1,000 req/s
This means:
- A single request returns in about 100ms
- The system as a whole can process about 1,000 requests per second
Both are about performance, but they look in different directions.
Does Low Latency Mean High Throughput?
This is a slightly tricky point.
Intuitively, you might think:
If each operation is fast, shouldn't it be able to process a lot?
Sometimes that is true.
However,
Low latency = high throughput
does not always hold.
For example, suppose a server can process one request in 100ms.
If it can only handle one request at a time, then simple arithmetic says it can process only about 10 requests per second.
Latency: 100ms
Concurrency: 1
Throughput: about 10 req/s
On the other hand, with the same 100ms per operation, if it can process 100 requests concurrently, things change.
Latency: 100ms
Concurrency: 100
Throughput: about 1,000 req/s
In other words, even with the same latency, throughput changes depending on concurrency.
Roughly speaking:
- Latency: how long one operation takes
- Concurrency: how many operations can be processed at once
- Throughput: how many operations end up being processed per unit of time
So you cannot conclude from latency alone that "this system holds up well under a large number of requests."
In Web APIs, Latency Tends to Matter
Consider opening a product page on an e-commerce site.
Browser -> GET /products/123 -> API -> Response
There is a big difference in how it feels when this API responds in 100ms versus 3 seconds.
For operations where the user is directly waiting, latency tends to be important.
Of course, a service with heavy traffic needs throughput too.
But if you have:
Throughput: 10,000 req/s
p95 Latency: 5 seconds
it may be quite hard to use.
Being able to process a large volume and each individual user having a comfortable experience are separate matters.
In Batch Processing, Throughput Can Matter More
Conversely, for operations where nobody is waiting for the result on the spot, throughput can matter more.
For example, consider the task:
Convert 1 million images
Even if the time per image goes from:
100ms -> 90ms
if only a few can be processed at once overall, it may not be a big improvement.
Instead, increasing the amount processed per unit of time, such as:
100 images/sec -> 1,000 images/sec
can have a bigger effect.
Picture adding more workers to process in parallel.
Job Queue -> Worker A
Job Queue -> Worker B
Job Queue -> Worker C
Job Queue -> Worker D
Even if the latency per item barely changes, overall throughput can go up.
It Varies by Operation in IoT, Too
The same is true in IoT.
Consider a system that receives telemetry from a large number of devices.
Device -> MQTT Broker -> Consumer -> Database
If tens of thousands of devices send data periodically, then rather than:
whether one message is processed in 20ms or 30ms
what matters more can be:
how many tens of thousands of messages can be processed per second
That is throughput.
Now what about sending an emergency stop command to a drone or robot?
Server -> Command -> Device
In this case, rather than:
how many thousands of emergency stop commands can be processed per second
you care about:
how quickly the command arrives after being sent
Here, latency matters.
So even within the same system:
- Telemetry reception is about throughput
- Real-time control is about latency
The important metric changes depending on the operation.
Rather than broadly labeling a system "latency-focused" or "throughput-focused," it can be more natural to think per operation.
Latency Degrades as You Approach the Throughput Limit
Latency and throughput are separate metrics, but they are not completely unrelated.
Suppose a server can handle at most about 1,000 req/s.
If only 500 req/s are coming in, there is plenty of headroom.
Incoming: 500 req/s
Capacity: 1,000 req/s
In this state, requests are processed immediately, so latency tends to stay stable.
But as load increases and approaches the limit, like:
900 req/s
950 req/s
990 req/s
requests start to wait.
Requests -> Queue -> Server
Even if the time the server actually spends processing is the same, if the time spent waiting in the queue grows, the latency seen by users gets worse.
Now suppose:
Incoming: 1,200 req/s
Capacity: 1,000 req/s
1,200 requests arrive each second, but only 1,000 can be processed each second.
The backlog then grows by 200 every second.
When that happens, three things occur at the same time:
- Throughput plateaus around 1,000 req/s
- The queue keeps growing
- Only latency keeps getting worse
That is why, in load testing, you should look not only at:
the maximum req/s achievable
but also at:
how latency changes as load increases
Looking Only at Average Latency Is Risky
When looking at latency, relying only on the average is also something to be careful about.
Average Latency: 100ms
Looking at just this, it seems quite fast.
But in reality, you might have:
p50: 50ms
p95: 300ms
p99: 2,000ms
In that case, most requests are fast, but some take more than 2 seconds.
With the average alone, these slow requests are hard to see.
So when looking at API latency, we often look at percentiles such as:
- p50
- p95
- p99
For example, if:
p95 = 300ms
it means:
95% of requests finished within 300ms
Compared to looking at the average alone, it also tells you how many slow requests exist.
So Which Should You Optimize?
Ultimately, it depends on the system and the operation.
As a tendency, you can group them like this:
| Operation | Metric usually emphasized |
|---|---|
| Web API | Latency |
| Batch processing | Throughput |
| Large-scale message processing | Throughput |
| IoT telemetry | Throughput |
| Real-time control | Latency |
But you cannot simply decide:
It's web, so latency
It's batch, so throughput
A Web API that receives heavy traffic also needs throughput, and a batch job may have strict completion-time requirements.
So rather than starting from:
should we improve latency or throughput?
it is better to start from:
what would be a problem if it were slow in this operation?
how much volume would be a problem if we could not process it?
It is also common to have requirements on both, such as:
p95 Latency < 300ms
Throughput > 5,000 req/s
Meeting the required latency while also securing the required throughput.
In practice, that is often how it ends up being approached.
"Good Performance" Alone Is Somewhat Vague
The statement "this system is fast" is quite vague when you think about it.
It might mean low latency.
It might mean high throughput.
For example, there are systems where:
each individual operation is quite fast, but it is weak under heavy traffic
And conversely, systems where:
it can process large amounts of data, but each individual operation takes time
So when talking about performance, you need to separate out:
what exactly is fast
The basis for that is latency and throughput.
If you want a rough way to remember them:
Latency = how long you wait for one operation
Throughput = how much can be handled in a given time
That is probably enough to start with.