E-step · expectation
Which group does each rating probably belong to?
Using the current guess of each group, Bayes' theorem gives every rating a probability of coming from each group: its responsibility.
Expectation–Maximisation, by doing
Imagine knowing only how people rate films, not what kind of viewer they are, and still recovering the groups behind the ratings. That is what the EM algorithm does. This site turns a long, maths-complete explainer and its Python notebook into something you can step through, poke and break.
Watch the guided tour· three short walkthroughs
A personal project by Sunchuangyu (Rin) Huang · written 2025, revived 2026
The idea
The explainer calls EM a smart detective. It cannot see who belongs to which group, so it alternates between two easier questions until the answers stop changing.
E-step · expectation
Using the current guess of each group, Bayes' theorem gives every rating a probability of coming from each group: its responsibility.
M-step · maximisation
Each group's share, average and spread become weighted averages, with ratings counting in proportion to how much they belong.
What the original notebook found
The notebook draws 200 “movie ratings”: 60% from sci-fi lovers around 7.5 and 40% from romance lovers around 4.0, clipped to 1 to 10. It then hides the labels and lets EM, written from scratch, find the groups from a random start. These are its printed results, reproduced by this site's TypeScript port to within 10⁻⁶.
| group | share π | mean μ | spread σ | fitted as |
|---|---|---|---|---|
| Sci-fi lovers fitted as component 2 | 0.656true 0.6 | 7.361true 7.5 | 1.273true 1.2 | component 2 |
| Romance lovers fitted as component 1 | 0.344true 0.4 | 4.022true 4.0 | 1.249true 1.5 | component 1 |
Note which component each group was fitted as: EM put the sci-fi lovers in component 2. That small detail is the first of the pitfalls.
Six ways in
Honest notes
Rebuilding the notebook number for number turned up three things worth knowing. The originals stay as they were; the site shows what happened and why.
Some hand-worked densities are off
The explainer uses and ; both are 0.4839. The conclusions hold. See both side by side.
The labels switched
Component 1 ended up as the low-mean group, so the notebook's final cell calls an 8.5 rater a romance lover (P = 0.001 for sci-fi). Why it happens.
Fifteen iterations was not convergence
The run stopped at its cap while still improving; finishing it changes the answer. Watch it finish.
About this project
The explainer and notebook are kept unchanged in original/. A script, export_parity.py, re-executes the notebook's cells with seed 42, checks every printed line against the saved output, and exports the ratings, the random start and the per-iteration trace. The TypeScript port of EMAnalyzer is tested against that trace to 10⁻⁶, and regenerates the notebook's console output character for character. Nothing here needs a server or an account; the one optional AI feature runs in your browser with your own key, and the methods page says what it does.