Skip to content
EM lab, home

Stepper

Four ratings, two groups, one step at a time

The explainer works EM through by hand on just four movie ratings, [2, 3, 7, 8]. Here you can replay it: each E-step works out how likely each rating is to belong to each group, and each M-step redraws the groups from those probabilities.

Starting guess

μ = 2.5 / 7.5, σ = 0.5, π = 0.5. Already at the answer, so EM stops at once.

  • Group 1
  • Group 2
  • Mixture
bars under each rating: γ₁ (teal) vs γ₂ (coral)
00.10.20.30.4densityμ₁ 2.50μ₂ 7.50

Start: the initial guess.

start

Initial guess (iteration 0)

We have four ratings, [2, 3, 7, 8], and suspect two kinds of user behind them. EM needs a starting guess for each group's average (μ), spread (σ) and share of users (π). It does not have to be good; the steps fix it.

guess
groupπμσ
10.50002.50000.5000
20.50007.50000.5000

Press play, or step with the arrow buttons.

Explain this iterationoptional · AI · your own key

Sends this iteration's numbers (parameters before and after, responsibilities, log-likelihood, stopping rule) and a fixed description of the page and data set, with no personal data, to Anthropic, from your browser. What is sent

Step to an iteration first: the starting guess has nothing to explain yet.

Log-likelihood

-5.6758
ℓ(θ)=∑ilog⁡(π1f(xi∣μ1,σ1)+π2f(xi∣μ2,σ2))\ell(\theta) = \sum_i \log\big(\pi_1 f(x_i \mid \mu_1, \sigma_1) + \pi_2 f(x_i \mid \mu_2, \sigma_2)\big)
ℓ(θ)=∑ilog⁡(π1f(xi∣μ1,σ1)+π2f(xi∣μ2,σ2))\begin{aligned} \ell(\theta) = \textstyle\sum_i \log\big(&\pi_1 f(x_i \mid \mu_1, \sigma_1) \\ &+ \pi_2 f(x_i \mid \mu_2, \sigma_2)\big) \end{aligned}
iterationℓ(θ)

What the hand-worked numbers got wrong (and right)

The explainer's worked example is the clearest part of it, so it is worth being exact about it. The density values it plugs in are off, but the story they tell is right: ratings 2 and 3 land in group 1 with near certainty, 7 and 8 in group 2, and the first M-step returns μ = 2.5 and 7.5 with σ = 0.5 and π = 0.5. In fact the starting guess was already a fixed point, so the second iteration changes nothing, exactly as the explainer says.

The original text is kept unchanged in original/em-explainer.md. The Maths page lists every correction in context.

Corrections to the worked example
quantityas writtenexact
normal_pdf(2; 2.5, 0.5)2 is exactly one standard deviation below 2.5, so the density is φ(1)/0.5.0.80.4839
normal_pdf(3; 2.5, 0.5)3 is also one standard deviation from 2.5, so it must equal the density at 2.0.60.4839
normal_pdf(3; 7.5, 0.5)3 is nine standard deviations from 7.5; the density is about 2 × 10⁻¹⁸.0.00012.06 × 10⁻¹⁸
normal_pdf(2; 7.5, 0.5)Eleven standard deviations away; the density is about 4 × 10⁻²⁷.0.00014.24 × 10⁻²⁷