Skip to content
EM lab, home

Guided tour · start here

EM in three walkthroughs

Three short screen recordings of the main journeys, about 3 minutes in all: stepping through EM by hand on four ratings, fitting the notebook's mixture to 200 ratings (and watching a bad start get stuck), and asking how sure the fit is. Each step is listed beside its video as a transcript; press a step to jump to it. Below them, screenshots of every key feature.

Walkthrough 1 of 3 · 54 s

Step through EM by hand

The explainer's worked example on four ratings, one E-step and one M-step at a time, twice over, with the explainer's hand-worked numbers shown beside the exact ones.

No sound; the caption banner is part of the recording. Download the MP4 or the captions (WebVTT).

Steps and transcript

Data and settings
Ratings [2, 3, 7, 8]; the explainer's starting guess μ = 2.5 and 7.5, σ = 0.5, π = 0.5; tolerance 10⁻⁶. No randomness is involved.
Try it yourself
Overview Stepper

Walkthrough 2 of 3 · 1 min 21 s

Fit the mixture to 200 ratings

The notebook's own experiment, live: its 200 ratings and its random start, the log-likelihood climbing and the responsibilities settling, the run finished past the notebook's 15-iteration cap, then a bad start that lands on a lower local maximum.

No sound; the caption banner is part of the recording. Download the MP4 or the captions (WebVTT).

Steps and transcript

Data and settings
The notebook's 200 synthetic ratings (np.random.seed(42), exported from the notebook) and its own random start; tolerance 10⁻⁴. The bad start is a dragged guess: μ₁ moved to about 6.3 with σ₁ = 0.2, μ₂ = 8.5 and σ₂ = 1 left as they were.
Try it yourself
Playground Local maxima

Walkthrough 3 of 3 · 1 min 3 s

How sure are we? Intervals and the number of groups

Uncertainty around the notebook's fit: observed-information standard errors and parametric-bootstrap intervals, re-run in the browser with the published seed, then AIC and BIC for one to four groups and a bootstrap likelihood-ratio test of one group against two.

No sound; the caption banner is part of the recording. Download the MP4 or the captions (WebVTT).

Steps and transcript

Data and settings
Parametric bootstrap with B = 1,000 at seed 42 (re-run in the browser's Web Worker during the recording). Model choice: K = 1 to 4, best of 60 ordinary and 30 pile starts per K, variance floor σ ≥ 0.1. LRT: B = 500.
Try it yourself
Uncertainty Choosing K

Key features

Screenshots

Desktop shots at 1440 × 900 in light mode (the overview in dark mode too) and three phone shots. Select one to enlarge it, then use the arrow keys to move through them. The two AI shots show a mocked AI response for illustration: no key was entered and no provider was called.

  • Overview. The notebook's fit, the idea in two steps and the six ways in.
  • Overview, dark mode. The same page in dark mode.
  • Stepper: the E-step. Responsibilities for [2, 3, 7, 8], with the explainer's hand-worked numbers beside them.
  • Stepper: the M-step. Each update as a weighted average, the explainer's working underneath.
  • Playground. The notebook's 200 ratings, its random start and the log-likelihood by iteration.
  • A bad start, a local maximum. A dragged start converges to ℓ = −423.63, below the −415.37 of the notebook's start.
  • Pitfall: label switching. The notebook's last cell as it ran, and the components matched by mean.
  • Pitfall: local maxima. Complete EM runs from 24 random starts, best first, run in a Web Worker.
  • Standard errors and bootstrap intervals. Hessian SEs, Wald and parametric-bootstrap 95% intervals, with the misses marked.
  • Do 95% intervals cover 95%? Coverage across simulated data sets, each rate with its Wilson interval.
  • Choosing K. AIC and BIC for one to four components, with and without the clipped ratings.
  • Bootstrap likelihood-ratio test. The bootstrap null against the χ² reference that Wilks' theorem would suggest.
  • Maths. The explainer's derivations rendered with KaTeX, corrections labelled.
  • Methods. Data provenance, method, evaluation design, assumptions and limitations.
  • Decision record DR-002. Label switching: the decision first, then the options, what happened and changes.
  • Model card. Intended use, data provenance, evaluation with intervals and known failure modes.
  • Bring your own key. Optional AI settings: Anthropic by default, the key stays in this browser.
  • Explain this iteration (mocked). Mocked AI response for illustration: labelled AI-generated, grounding-checked, awaiting review.
  • AI audit log (mocked entry). Mocked AI response for illustration: the call, the reviewer's decision and JSON/CSV export.

On a phone (390 px wide)

  • Phone: overview. The overview at 390 px.
  • Phone: stepper. The E-step and the explainer's numbers on a phone.
  • Phone: inference. The bootstrap histograms, one per parameter.

Reproducible

How these were made

Every frame comes from a script, not a screen-capture session. A Playwright tour (web/e2e/showcase.spec.ts) drives the site in Google Chrome at 1280 × 800, adds the caption banner and the cursor highlight, and records the video. It also checks what it shows: the explainer's densities against the exact 0.4839, the notebook's log-likelihood after 15 iterations and at convergence, the lower maximum the bad start reaches, the bootstrap interval for π₁ and the browser re-run that reproduces it, BIC's choice of K and the bootstrap likelihood-ratio test. A broken feature fails the tour rather than producing a misleading video.

The ratings are the notebook's own, exported from it, and every simulation on the site is seeded, so a rerun records the same numbers. No API key is entered anywhere: the AI settings dialog is only opened, and the AI explanation in the screenshots is a mocked reply served inside the test browser, labelled as such on screen.