Why do systems collapse before they are full?
Topic 08 solved the shortest path under a quiet assumption: choosing an edge means taking it. Topic 20 removes that assumption on the same network, and topic 22 removes the assumption that the numbers are known before you commit; the two honest ways to cope (average over what might happen, or armour against the worst of it) give different answers on the same problem. Topic 09 is the mathematics of waiting when arrivals and service are random.
Almost every manager's instinct is that a resource running at 95% is being used well and one running at 60% is being wasted. The mathematics says the opposite, and it is not close.
The reason is variability. If work arrived like clockwork and every job took exactly the same time, you could run at 100% with no queue at all. But arrivals clump and jobs differ, so idle time and backlog are not symmetric: an idle minute is gone forever, while a backlog persists and compounds. As utilisation rises there is less and less slack to absorb a clump, and the queue that forms has less and less opportunity to drain.
So the spare capacity that looks like waste is doing real work. It is what stops a random cluster of arrivals turning into a two-hour wait. This is why emergency departments target occupancy well below 100%, why motorways jam at flow rates below their theoretical maximum, and why a disk at 95% IO utilisation feels broken rather than efficient.
The practical rule that follows: if you cannot add capacity, reduce variability. Appointment systems, batching, and admission control all attack the variance rather than the mean, and they work.
The simplest useful model is M/M/1: Poisson arrivals at rate λ, exponential service at rate μ, one server. Write utilisation ρ = λ/μ. Then, in steady state:
L = ρ / (1 − ρ) average number in the system
W = 1 / (μ − λ) average time in the system
→ ∞ as ρ → 1
The 1/(1 − ρ) factor is the whole story. At ρ = 0.5 it is 2. At ρ = 0.9 it is 10. At ρ = 0.99 it is 100. Nothing about the server changed; only how close it is to the wall.
Little's law is the deeper result and one of the few in this field that needs almost no assumptions: for any stable system whatsoever, regardless of distributions or scheduling discipline,
L = λ W
Average number in the system equals arrival rate times average time in it. It holds for hospitals, factories, and TCP connections alike, and it is the reason you can infer a quantity you cannot measure from two you can.
The honest caveat: M/M/1's exponential assumptions are a convenience. Real service times are rarely exponential, and the Pollaczek–Khinchine formula shows waiting scales with the variance of service time, not just its mean, which strengthens rather than weakens the argument above.
Emergency department capacity Bed occupancy is the single most-argued number in hospital management, and this curve is why. Pushing occupancy from 85% to 95% looks like a 12% efficiency gain on a spreadsheet and behaves like a crisis on the floor: ambulances queue outside because there is no slack to absorb a bad hour. The same curve sets call-centre staffing (via Erlang C, the multi-server sibling), airport security lanes, and how much headroom a cloud service keeps before autoscaling.
Check yourself
In the M/M/1 queue, what happens to the average time in the system when utilisation rises from 50% to 80%?
Time in the system is 1/(1 - rho) service times: 2 at 50% and 5 at 80%, so 2.5 times longer. Nothing about the server changed, only how close it is to the wall.