Can a perfectly certain whole have a genuinely uncertain half?
Every topic so far has talked about qubits, measurement, and algorithms without ever asking the most basic question a communication engineer would ask first: how much information is actually here, measured in bits? That question has a classical answer nearly 80 years old, and a quantum generalisation that both agrees with it in the classical limit and produces one result with no classical analogue at all: the fact that a perfectly certain whole can have a genuinely uncertain half. This closes two loose threads left open on purpose: topic 02 promised the object entropy is computed from, and topic 07 promised the number that finally says how much two systems are entangled, not just whether they are.
The von Neumann entropy generalises Shannon entropy from a classical probability distribution to a quantum density matrix, by applying the identical formula to ρ's eigenvalues, which are, after all, a genuine probability distribution (topic 02: they're exactly the weights in ρ's mixture of orthogonal pure states). S(ρ)=0 for any pure state. No uncertainty about which state you have, because you have a definite one. S(ρ)=log₂d for the maximally mixed state on a d-dimensional space: total ignorance, the quantum analogue of a uniform distribution.
Here's the fact that has no classical shadow at all. Take a Bell pair: a perfectly pure, perfectly certain joint state, S=0. Trace out one qubit (topic 02's partial trace) and look at the other alone: its entropy is a full bit, the maximum a single qubit can have. A whole with zero uncertainty has a half with maximum uncertainty. Classically this is impossible, if you know a joint outcome (x, y) with certainty, you know x and y individually with certainty too. Quantum mechanically, entanglement makes it not only possible but exact and provable.
This is also, finally, the honest answer to "how entangled is a state," promised back in topic 07. Schmidt decomposition showed every bipartite pure state reduces to r matched terms with coefficients λi; von Neumann entropy of either reduced state, S(ρA)=−Σλi²log₂λi², is the standard single number used to quantify how much: zero for an unentangled product state, and climbing to log₂(Schmidt rank) at maximal entanglement, exactly 1 bit for a Bell pair.
For a density matrix ρ with eigenvalues λi:
S(ρ) = −Tr(ρ log₂ρ) = − Σᵢ λᵢ log₂λᵢ
: literally the Shannon entropy of ρ's eigenvalue spectrum. Two structural facts, both consequences of Schmidt decomposition (topic 07): for a bipartite pure state, S(ρA)=S(ρB) always, since the reduced density matrices on both sides share the identical nonzero eigenvalues (the λi² from the Schmidt coefficients); and subadditivity holds for any bipartite state, pure or mixed: S(ρAB) ≤ S(ρA)+S(ρB), with equality exactly for uncorrelated product states.
The Holevo bound answers a different, sharper question: given an ensemble {pi, ρi} of quantum states (e.g. what Alice sends to encode a classical message), how much classical information can Bob actually extract by measuring? The Holevo quantity:
χ = S(ρ̄) − Σᵢ pᵢ S(ρᵢ) where ρ̄ = Σᵢ pᵢ ρᵢ
is an upper bound on the accessible information (the best mutual information any measurement, of any kind, can extract) and satisfies 0 ≤ χ ≤ log₂d for a d-dimensional system. (Holevo, Problems of Information Transmission 9, 177 (1973).) Verified numerically: S(ρA)=S(ρB) to within 2×10⁻¹⁵ over 350 random bipartite pure states; subadditivity held with zero violations over 500 random mixed bipartite states; χ stayed within [0, log₂d] with zero violations over 500 random ensembles.
A concrete check that ties directly back to topic 03's Helstrom bound: encode a bit as |0⟩ or |+⟩ (equal priors: the exact non-orthogonal pair topic 03 used). χ = 0.6009 bits. Running the actual optimal (Helstrom) measurement and computing the real mutual information obtained gives only I(X: Y) = 0.3991 bits: the bound holds (0.3991 < 0.6009) but isn't tight here, because the best single-shot binary discrimination just can't extract everything the ensemble's entropy would allow. The Helstrom success probability computed from the simulated measurement matched topic 03's closed-form trace-distance formula exactly (0.853553 both ways): two different bounds, on two different questions (best discrimination rate vs. best information rate), computed side by side and found consistent.
Closing the loop on superdense coding, and on entanglement itself Topic 06 asserted "never do better than 2 classical bits per qubit" without deriving it: the Holevo bound is exactly that limit, made precise: it caps accessible information at log₂d for a d-dimensional system, and in superdense coding Bob measures two qubits (his stored half plus the one he receives), so d=4 and the cap is log₂4=2 bits: precisely what the protocol achieves. Entanglement doesn't break the bound; it's spent in advance to let the received qubit's information reach that ceiling. And entanglement entropy (this topic's other half) is the number quoted whenever a physics paper states how much two systems are entangled, from the Bell pair's exactly 1 bit up to the much larger entropies that describe entanglement across a real many-body system or a black hole's event horizon, the frontier quantum mechanics vs. computing gestures at when it calls entanglement entropy a phase diagnostic in condensed matter.
Check yourself
A Bell pair is a perfectly pure joint state. What is the von Neumann entropy of one qubit taken alone?
The whole has zero uncertainty and each half has the maximum one bit. Classically a joint outcome known with certainty makes each part certain too; entanglement is what breaks that.