Pitfalls
Where EM quietly goes wrong
EM always makes progress, which is exactly why its failures are easy to miss: nothing crashes, the log-likelihood goes up, and the answer can still be mislabelled, unfinished, stuck or degenerate. Each one below shows up in, or right next to, the original notebook.
pitfall 1
Label switching
EM has no idea what its components mean. “Component 1” is just whichever bump the first random guess happened to steer towards. The notebook generated its data with group 1 as sci-fi lovers (mean 7.5), but after fitting, its component 1 sat on the low ratings: μ₁ = 4.02.
Its last cell then asked about a new user who rated a film 8.5, and read component 1 as “sci-fi lover”. The answer it printed, kept here exactly as it ran, is the wrong way round: 8.5 sits squarely in the high-mean component.
EXAMPLE: Classify a new user ======================================== New user rating: 8.5 P(Sci-fi lover | rating=8.5): 0.001 P(Romance lover | rating=8.5): 0.999 → This user is likely a ROMANCE LOVER! (confidence: 99.9%) This is how you use EM to make predictions on new data!
What each fitted component actually captured
| component | fitted μ | notebook calls it | matched by mean |
|---|---|---|---|
| 1 | 4.022 | Sci-fi lovers | Romance lovers |
| 2 | 7.361 | Romance lovers | Sci-fi lovers |
The fix is bookkeeping, not maths: after fitting, match each component to the group whose mean it is closest to (or simply sort components by mean). Every comparison with the true groups on this site does that first.
Is it bad luck? Rerun the notebook 100 times
Same data, same recipe (μ ~ U(3, 8), σ ~ U(0.5, 2), 15 iterations), 100 fresh random starts.
fitting in a background worker...
pitfall 2
Stopping early
The notebook calls fit(max_iterations=15, tolerance=1e-4). Its tolerance was never reached: at iteration 15 the log-likelihood was still rising by 0.029 per step, nearly 300 times the tolerance. The printed “final” parameters are a snapshot of a run in progress.
Continue from the same start and EM needs 90 iterations to meet its own stopping rule. It ends somewhere noticeably different, with a higher likelihood (−415.37) but a lower classification accuracy: 83.5% (Wilson 95% CI 77.7% to 88.0%) against 90.5% (85.6% to 93.8%).
Both fits classify the same 200 users, so the fair comparison is paired: a change of −7.0 percentage points (paired bootstrap 95% CI −11.0 to −3.0). 16 users went from right to wrong and 2 from wrong to right; the other 182 were classified the same way by both. Both accuracies are in-sample (each fit is scored on the ratings it was fitted to), so the intervals describe the 200 users drawn, not the uncertainty of either fit.
That is not a bug. EM maximises the likelihood of the data it is given, and these ratings were clipped at 10, piling seven of them onto one value and skewing the high group. The best-fitting pair of normal curves for that shape is not the pair that generated it.
Log-likelihood, same start, run to the notebook's tolerance
notebook, t = 15ℓ = -416.51
- low group
- π 0.34 · μ 4.02 · σ 1.25
- high group
- π 0.66 · μ 7.36 · σ 1.27
converged, t = 90ℓ = -415.37
- low group
- π 0.22 · μ 3.37 · σ 0.86
- high group
- π 0.78 · μ 7.02 · σ 1.45
true
- low group
- π 0.40 · μ 4.00 · σ 1.50
- high group
- π 0.60 · μ 7.50 · σ 1.20
Try it in the playground: keep the notebook's data and start, and raise the iteration cap.
pitfall 3
Local maxima
EM climbs uphill from wherever it starts, so it finds a peak of the likelihood, not necessarily the peak. The standard remedy is cheap: start from several guesses and keep the run with the highest log-likelihood. Below, each card is a complete EM run (up to 1,000 iterations, tolerance 10⁻⁶) from its own random start, best first. The fits run in a Web Worker so the page stays responsive.
Give EM three bumps but only two components and it has to choose which two bumps to merge; where it starts decides which it picks, and the two answers differ by about 14 in log-likelihood. A k-means++ start finds the better merge every time here.
On the notebook's own data nearly every random start reaches the answer the notebook's run was heading for (ℓ = −415.37). Now and then one finds a different peak with an even higher likelihood (ℓ ≈ −413.99): a narrow component sitting on a cluster of high ratings. Switch the data and draw new starts a few times to catch one, or see on the inference page how often it happens and what it means for “keep the best of many starts”.
data
start
restarts
Each thumbnail draws the lower-mean component in teal, whichever number EM gave it, so the same answer always looks the same. Which component ends up as “1” is label switching.
- Lower-mean component
- Higher-mean component
- Mixture
pitfall 4
Variance collapse
The mixture likelihood has no ceiling. Put one component's mean exactly on a data point and shrink its spread: that point's density grows without limit while every other point is still covered by the other component, so the log-likelihood heads to +∞. EM, doing its job, follows it there.
The notebook's data has a ready-made trap: clipping at 10 turned seven ratings into the identical value 10.0. A component that starts on them with a small spread shrinks every iteration until σ is zero (the next step divides zero by zero) or floating-point residue that no longer moves, which a naive loop would happily call “converged”. The notebook's random start (σ ≥ 0.5) never got close, so its run was safe.
With component 2 sitting on one point , the log-likelihood is at least
and the first term goes to as while the rest stays put.
The usual guards are a floor on σ (or a prior on it), dropping components whose weight collapses, and restarting.
Sit component 2 on…
Starting at μ₂ = 10.00 (the 7 ratings clipped to 10); component 1 starts at the overall mean and spread.
Collapsed at iteration 7: σ₂ reached exactly 0, so the density is 0/0 and the log-likelihood is NaN. Just before, it had reached ℓ = −394.64, already above the notebook's converged fit (−415.37). The likelihood has no maximum here: it grows without bound as σ₂ → 0.
σ₂ by iteration (log scale)
log-likelihood by iteration
Last fit before the collapse (iteration 6)
- Component 1
- Component 2
- Mixture