The Feasible Region · Topic 09 of 33 · reading step 17 of 33 · Uncertainty

Queueing theory

Why do systems collapse before they are full?

  1. 01
  2. 02
  3. 16
  4. 03
  5. 17
  6. 04
  7. 05
  8. 19
  9. 21
  10. 10
  11. 06
  12. 18
  13. 07
  14. 08
  15. 20
  16. 22
  17. 09
  18. 23
  19. 11
  20. 12
  21. 13
  22. 14
  23. 15
  24. 25
  25. 26
  26. 27
  27. 28
  28. 29
  29. 30
  30. 31
  31. 32
  32. 33
  33. 24
About this section: Uncertainty

Topic 08 solved the shortest path under a quiet assumption: choosing an edge means taking it. Topic 20 removes that assumption on the same network, and topic 22 removes the assumption that the numbers are known before you commit; the two honest ways to cope (average over what might happen, or armour against the worst of it) give different answers on the same problem. Topic 09 is the mathematics of waiting when arrivals and service are random.

See it

100% 50% busy in system = 2 80% busy in system = 5 (2.5× worse) 95% busy in system = 20 (4× worse again) 20%40% 60%80% how busy the server is (utilisation) time in the system
The most expensive graph in operations management, and the least intuitive. Every point is the exact M/M/1 formula, not a sketch. The quantity plotted is time in the system (queueing plus being served, the number a patient or a driver actually experiences) which is W = 1/(μ−λ), or 1/(1−ρ) in units of one service time. Going from half-busy to 80% busy multiplies it by 2.5×. Going from 80% to 95% multiplies it by 4× again. Queueing time alone is ρ/(1−ρ) and rises even faster than this, 4× and then 4.75× across the same two steps, so the curve you are looking at is the conservative one. The hospital, the motorway and the CPU all fall over at the same shape, and they all fall over well before they are "full".

The intuition

Almost every manager's instinct is that a resource running at 95% is being used well and one running at 60% is being wasted. The mathematics says the opposite, and it is not close.

The reason is variability. If work arrived like clockwork and every job took exactly the same time, you could run at 100% with no queue at all. But arrivals clump and jobs differ, so idle time and backlog are not symmetric: an idle minute is gone forever, while a backlog persists and compounds. As utilisation rises there is less and less slack to absorb a clump, and the queue that forms has less and less opportunity to drain.

So the spare capacity that looks like waste is doing real work. It is what stops a random cluster of arrivals turning into a two-hour wait. This is why emergency departments target occupancy well below 100%, why motorways jam at flow rates below their theoretical maximum, and why a disk at 95% IO utilisation feels broken rather than efficient.

The practical rule that follows: if you cannot add capacity, reduce variability. Appointment systems, batching, and admission control all attack the variance rather than the mean, and they work.

The mathematics

The simplest useful model is M/M/1: Poisson arrivals at rate λ, exponential service at rate μ, one server. Write utilisation ρ = λ/μ. Then, in steady state:

L = ρ / (1 − ρ) average number in the system W = 1 / (μ − λ) average time in the system → ∞ as ρ → 1

The 1/(1 − ρ) factor is the whole story. At ρ = 0.5 it is 2. At ρ = 0.9 it is 10. At ρ = 0.99 it is 100. Nothing about the server changed; only how close it is to the wall.

Little's law is the deeper result and one of the few in this field that needs almost no assumptions: for any stable system whatsoever, regardless of distributions or scheduling discipline,

L = λ W

Average number in the system equals arrival rate times average time in it. It holds for hospitals, factories, and TCP connections alike, and it is the reason you can infer a quantity you cannot measure from two you can.

The honest caveat: M/M/1's exponential assumptions are a convenience. Real service times are rarely exponential, and the Pollaczek–Khinchine formula shows waiting scales with the variance of service time, not just its mean, which strengthens rather than weakens the argument above.

Where it actually runs

Emergency department capacity Bed occupancy is the single most-argued number in hospital management, and this curve is why. Pushing occupancy from 85% to 95% looks like a 12% efficiency gain on a spreadsheet and behaves like a crisis on the floor: ambulances queue outside because there is no slack to absorb a bad hour. The same curve sets call-centre staffing (via Erlang C, the multi-server sibling), airport security lanes, and how much headroom a cloud service keeps before autoscaling.