Skip to content
EM lab, home

Pitfalls

Where EM quietly goes wrong

EM always makes progress, which is exactly why its failures are easy to miss: nothing crashes, the log-likelihood goes up, and the answer can still be mislabelled, unfinished, stuck or degenerate. Each one below shows up in, or right next to, the original notebook.

pitfall 1

Label switching

EM has no idea what its components mean. “Component 1” is just whichever bump the first random guess happened to steer towards. The notebook generated its data with group 1 as sci-fi lovers (mean 7.5), but after fitting, its component 1 sat on the low ratings: μ₁ = 4.02.

Its last cell then asked about a new user who rated a film 8.5, and read component 1 as “sci-fi lover”. The answer it printed, kept here exactly as it ran, is the wrong way round: 8.5 sits squarely in the high-mean component.

as-run output · original notebook, final cell
EXAMPLE: Classify a new user
========================================
New user rating: 8.5
P(Sci-fi lover | rating=8.5): 0.001
P(Romance lover | rating=8.5): 0.999
→ This user is likely a ROMANCE LOVER! (confidence: 99.9%)

This is how you use EM to make predictions on new data!

What each fitted component actually captured

componentfitted μnotebook calls itmatched by mean
14.022Sci-fi loversRomance lovers
27.361Romance loversSci-fi lovers

The fix is bookkeeping, not maths: after fitting, match each component to the group whose mean it is closest to (or simply sort components by mean). Every comparison with the true groups on this site does that first.

Is it bad luck? Rerun the notebook 100 times

Same data, same recipe (μ ~ U(3, 8), σ ~ U(0.5, 2), 15 iterations), 100 fresh random starts.

fitting in a background worker...

pitfall 2

Stopping early

The notebook calls fit(max_iterations=15, tolerance=1e-4). Its tolerance was never reached: at iteration 15 the log-likelihood was still rising by 0.029 per step, nearly 300 times the tolerance. The printed “final” parameters are a snapshot of a run in progress.

Continue from the same start and EM needs 90 iterations to meet its own stopping rule. It ends somewhere noticeably different, with a higher likelihood (−415.37) but a lower classification accuracy: 83.5% (Wilson 95% CI 77.7% to 88.0%) against 90.5% (85.6% to 93.8%).

Both fits classify the same 200 users, so the fair comparison is paired: a change of −7.0 percentage points (paired bootstrap 95% CI −11.0 to −3.0). 16 users went from right to wrong and 2 from wrong to right; the other 182 were classified the same way by both. Both accuracies are in-sample (each fit is scored on the ratings it was fitted to), so the intervals describe the 200 users drawn, not the uncertainty of either fit.

That is not a bug. EM maximises the likelihood of the data it is given, and these ratings were clipped at 10, piling seven of them onto one value and skewing the high group. The best-fitting pair of normal curves for that shape is not the pair that generated it.

Log-likelihood, same start, run to the notebook's tolerance

iterationℓ(θ)
  • notebook, t = 15ℓ = -416.51

    low group
    π 0.34 · μ 4.02 · σ 1.25
    high group
    π 0.66 · μ 7.36 · σ 1.27
  • converged, t = 90ℓ = -415.37

    low group
    π 0.22 · μ 3.37 · σ 0.86
    high group
    π 0.78 · μ 7.02 · σ 1.45
  • true

    low group
    π 0.40 · μ 4.00 · σ 1.50
    high group
    π 0.60 · μ 7.50 · σ 1.20
after 15 (notebook)
after 90 (converged)

Try it in the playground: keep the notebook's data and start, and raise the iteration cap.

pitfall 3

Local maxima

EM climbs uphill from wherever it starts, so it finds a peak of the likelihood, not necessarily the peak. The standard remedy is cheap: start from several guesses and keep the run with the highest log-likelihood. Below, each card is a complete EM run (up to 1,000 iterations, tolerance 10⁻⁶) from its own random start, best first. The fits run in a Web Worker so the page stays responsive.

Give EM three bumps but only two components and it has to choose which two bumps to merge; where it starts decides which it picks, and the two answers differ by about 14 in log-likelihood. A k-means++ start finds the better merge every time here.

On the notebook's own data nearly every random start reaches the answer the notebook's run was heading for (ℓ = −415.37). Now and then one finds a different peak with an even higher likelihood (ℓ ≈ −413.99): a narrow component sitting on a cluster of high ratings. Switch the data and draw new starts a few times to catch one, or see on the inference page how often it happens and what it means for “keep the best of many starts”.

data

start

restarts

Running 24 fits in a background worker...

Each thumbnail draws the lower-mean component in teal, whichever number EM gave it, so the same answer always looks the same. Which component ends up as “1” is label switching.

    • Lower-mean component
    • Higher-mean component
    • Mixture

    pitfall 4

    Variance collapse

    The mixture likelihood has no ceiling. Put one component's mean exactly on a data point and shrink its spread: that point's density grows without limit while every other point is still covered by the other component, so the log-likelihood heads to +∞. EM, doing its job, follows it there.

    The notebook's data has a ready-made trap: clipping at 10 turned seven ratings into the identical value 10.0. A component that starts on them with a small spread shrinks every iteration until σ is zero (the next step divides zero by zero) or floating-point residue that no longer moves, which a naive loop would happily call “converged”. The notebook's random start (σ ≥ 0.5) never got close, so its run was safe.

    With component 2 sitting on one point xjx_j, the log-likelihood is at least

    ℓ(θ)  ≥  log⁡π2σ22π  +  ∑i≠jlog⁡(π1f(xi∣μ1,σ1))\ell(\theta) \;\ge\; \log\frac{\pi_2}{\sigma_2\sqrt{2\pi}} \;+\; \sum_{i \ne j} \log\big(\pi_1 f(x_i \mid \mu_1, \sigma_1)\big)
    ℓ(θ)  ≥  log⁡π2σ22π+∑i≠jlog⁡(π1f(xi∣μ1,σ1))\begin{aligned} \ell(\theta) \;\ge\;& \log\frac{\pi_2}{\sigma_2\sqrt{2\pi}} \\ &+ \sum_{i \ne j} \log\big(\pi_1 f(x_i \mid \mu_1, \sigma_1)\big) \end{aligned}

    and the first term goes to +∞+\infty as σ2→0\sigma_2 \to 0 while the rest stays put.

    The usual guards are a floor on σ (or a prior on it), dropping components whose weight collapses, and restarting.

    Sit component 2 on…

    Starting at μ₂ = 10.00 (the 7 ratings clipped to 10); component 1 starts at the overall mean and spread.

    not in the notebook

    Collapsed at iteration 7: σ₂ reached exactly 0, so the density is 0/0 and the log-likelihood is NaN. Just before, it had reached ℓ = −394.64, already above the notebook's converged fit (−415.37). The likelihood has no maximum here: it grows without bound as σ₂ → 0.

    σ₂ by iteration (log scale)

    iterationσ₂

    log-likelihood by iteration

    iterationℓ(θ)

    Last fit before the collapse (iteration 6)

    • Component 1
    • Component 2
    • Mixture
    00.10.20.30.4density