Skip to content
EM lab, home

Methods

How the numbers were made, and what they cannot tell you

The data provenance, the method and its 2026 additions, how each claim on the site is checked, the assumptions and limits, the decisions behind them, a model card for the fitted mixture, and what the optional AI feature does.

1 · provenance

Data provenance

Every rating is synthetic. The 2025 notebook draws 200 “movie ratings” with np.random.seed(42): each user is a sci-fi lover with probability 0.6 (ratings from N(7.5, 1.2²)) or a romance lover (N(4.0, 1.5²)), and the ratings are clipped to 1 to 10, which puts seven of them at exactly 10.0. No real people or personal information are involved.

export_parity.py re-runs the notebook's cells verbatim and refuses to export unless every printed line matches the saved output (DR-001). It recorded Python 3.13.9 and NumPy 2.5.3. Data generated in the browser (the playground's “make your own”, restart galleries, every simulation on /inference) use the site's own seeded generator, xoshiro128**, and are never presented as the notebook's.

Generated artefacts and the scripts that make them
artefactmade byholds
public/data/notebook-run.jsonuv run scripts/export_parity.pythe 200 ratings, true groups, random start, 15-iteration trace, printed output
src/lib/inference/__generated__/inference.jsonpnpm inferenceevery number on /inference, with its seed and settings (DR-004)
src/lib/stats/__fixtures__/reference*.jsonuv run scripts/stats_reference.py; Rscript scripts/stats_reference.RSciPy, statsmodels, numdifftools and R values for the statistics tests

2 · method

Method

The original. A two-component Gaussian mixture fitted by EM written from scratch: the E-step computes each rating's responsibilities by Bayes' theorem, the M-step updates each share, mean and spread as responsibility-weighted averages, and the run stops when the log-likelihood moves by less than the tolerance or at the iteration cap. The TypeScript port follows the notebook line for line. It adds two things, both off the notebook's path: an optional variance floor and a stop when a component collapses.

Uncertainty (2026). Standard errors from the observed information, the negative Hessian of the observed-data log-likelihood by central differences, inverted, at the local maximum the notebook's own start converges to (a higher maximum exists; see /inference). Parametric-bootstrap percentile intervals: B data sets drawn from the fitted mixture, each refitted from the fit (warm start), components ordered by mean (DR-002). A coverage study simulates data sets from the known truth and counts how often each nominal 95% interval contains it.

Choosing K. EM for K = 1 to 4 components, working in log space, from many starts (k-means++, Forgy and random) with a variance floor of σ ≥ 0.1 (DR-003); AIC and BIC; a parametric bootstrap likelihood-ratio test of one component against two, because the χ² reference fails at the boundary of the parameter space.

Diagnostics and comparisons. Log-likelihood traces checked for any decrease; iterations to each tolerance across random starts; the share of starts that reach the best maximum, with a Wilson interval, turned into the chance that the best of R starts gets there. Two methods on the same data sets are compared as a paired mean difference with a bootstrap interval over data sets. Coverage rates, selection rates, success rates and the headline accuracy carry Wilson intervals, and every seed behind /inference is listed in its seeds and sizes table.

3 · evaluation

Evaluation design

Each kind of claim on the site has a check, and each check runs in CI on every push.

Claims and how they are checked
claimevidencewhere
The TypeScript port is the notebook's algorithmEvery π, μ, σ and log-likelihood of all 15 iterations within 10⁻⁶ of the exported trace; the console output regenerated character for character; other seeds, a run to convergence and a 300-iteration run as extra tracesem.parity.test.ts
The statistics helpers are rightNormal and χ² functions, Wilson intervals, quantiles and the observed-information SEs checked against SciPy, statsmodels and numdifftools, and against R (prop.test, qchisq, numDeriv)stats.reference.test.ts
The published inference numbers are what the code producesFast analyses re-run in full, slow ones through small stored checks with the same seedsartefact.test.ts
The 95% intervals deserve the nameCoverage over 500 simulated data sets (Wald) and 200 (bootstrap, paired with Wald), with Wilson intervals on every coverage rate/inference#coverage
The number of componentsAIC and BIC for K = 1 to 4 (60 ordinary starts plus 30 with a narrow component on the ratings piled at 10.0), sensitivity checks for the variance floor (0.05, 0.1, 0.25) and without the clipped ratings, selection rates on 100 fresh samples, and a parametric bootstrap LRT (B = 500)/inference#choosing-k
EM behaves as the theory saysMonotone log-likelihood and iterations to tolerance over 200 random starts; random start against k-means++ paired over 200 simulated data sets/inference#convergence
The AI client is safe with a keyAdapters, errors, key storage, redaction, audit log and the grounding check, with the network mockedai.test.ts

4 · assumptions

Assumptions

  • Ratings are independent draws from a mixture of normal distributions. The clipping at 1 and 10 is not part of the model.
  • The reported fit is the maximum the notebook's own run was heading to. Its standard errors and bootstrap intervals describe that maximum, not the likelihood surface as a whole (a second, higher maximum exists).
  • Wald intervals assume the log-likelihood is close to quadratic near the maximum; the bootstrap assumes the fitted mixture is a fair stand-in for the truth.
  • Simulations use the site's generator, not NumPy's; their conclusions do not depend on which generator draws the numbers.

5 · limitations

Limitations

  • One synthetic data set of 200 ratings. Its accuracy, 90.5% as the notebook reported it, has a Wilson 95% interval of 85.6% to 93.8%; the converged fit's is lower (see the model card). Both are in-sample: the same ratings fitted the model, so the interval reflects which users happened to be drawn given the fitted rule, not the uncertainty of the fit itself, and it is optimistic about accuracy on new ratings.
  • At n = 200 the nominal 95% intervals cover the truth less often than 95%: Wald intervals between 84% and 92%, with the model exactly right or clipped like the notebook.
  • The bootstrap coverage study uses B = 200 bootstrap replicates per data set (200 data sets) to keep the compute manageable, where B = 1,000 is the usual minimum for percentile intervals. Each 2.5% and 97.5% endpoint then rests on about the 5th and 195th of 200 values, so the endpoints carry Monte Carlo error that this study does not quantify, and the bootstrap coverage rates and the paired bootstrap-against-Wald differences could move with a larger B.
  • The model ignores clipping, so the pile of ratings at 10.0 can be “explained” by an extra component: on this sample BIC chooses 3 components and AIC 4, and BIC's choice moves with the variance floor (DR-003).
  • The likelihood-ratio test's null distribution depends on the variance floor and on how many starts each fit gets (DR-003).
  • The AI feature has only been tested with the network mocked: this project has no API key.

6 · next time

What I'd change

  • Fit a censored mixture, so a rating clipped to 10 counts as “at least 10” (and one clipped to 1 as “at most 1”) rather than an exact value, and repeat the model choice.
  • Replace the hard variance floor with a weak prior on σ², and report the LRT for several floors.
  • Use profile-likelihood or bias-corrected bootstrap intervals, which should cover better than Wald intervals at this sample size.
  • Regenerate every artefact in a scheduled CI job and fail on any diff.
  • Build a small evaluation set for the AI explanations and score models on it, with intervals.

7 · decisions

Decision records

Each record states the decision first, then the options, the reasons, what actually happened (weak numbers included) and what I would change. Records are never edited after the fact; a new record supersedes an old one. The sources are in docs/decisions.

  1. DR-001Precomputed, seeded data for parity with the notebook
  2. DR-002Label switching, and how fitted components are named
  3. DR-003A variance floor for the inference fits, never for the parity path
  4. DR-004Precompute the slow simulations, and test the artefact
  5. DR-005"Explain this iteration": optional, bring your own key, grounded and audited
  6. DR-006Say exactly what the AI feature sends, which model answered, and how the site was built

DR-001 · Accepted · 2026-10-09

Precomputed, seeded data for parity with the notebook

Applies to: scripts/export_parity.py, web/public/data/notebook-run.json, /playground, every "original run" number on the site

Context

The notebook draws its 200 ratings, the hidden groups and EM's random starting guess from NumPy's global random stream after np.random.seed(42). That stream (the legacy Mersenne Twister, NumPy's legacy_gauss normal sampler and its uniform conversion) is not available in the browser. The revived site needed the exact same data and the exact same starting guess to claim that its TypeScript port of EMAnalyzer reproduces the notebook, and it needed fresh random data for the "make your own" playground and the restart galleries.

Decision

Run the notebook's code cells verbatim in scripts/export_parity.py (same order, same seed, same class), fail unless every printed line matches the output saved in the notebook, and export the ratings, true groups, random start and per-iteration trace to web/public/data/notebook-run.json. The TypeScript port is tested against that trace. Anything generated in the browser uses a separate seeded generator (xoshiro128** seeded through splitmix32) and the interface says it is not the notebook's data.

Options considered

  1. Port NumPy's generator to TypeScript (MT19937 plus the legacy normal sampler). Possible, but a large amount of code whose only purpose is to reproduce one draw, and any floating-point difference in the sampler would silently change the data.
  2. Copy the numbers out of the notebook by hand. The notebook never prints all 200 ratings or the full-precision starting guess, so this is not even possible without re-running it.
  3. Run Python in the browser (Pyodide). Exact, but a download of tens of megabytes for a page that only needs 24 KB of numbers.
  4. Re-execute and export with a script (chosen).

Why

Provenance is the point of the revival: every number on the site should trace back to the original code. A script that refuses to export unless the notebook's own printed output is reproduced is a stronger guarantee than any reimplementation, and it keeps the browser code simple.

What happened

  • The port matches the exported trace to within 10⁻⁶ in every π, μ, σ and log-likelihood of all 15 iterations, and the regenerated console output is identical, character for character, to what the notebook printed (em.parity.test.ts).
  • The export records the environment it ran in: Python 3.13.9 and NumPy 2.5.3.
  • The weak spot is the second random stream. The same seed gives different numbers in the playground than in NumPy, so "seed 2025" on the site and in Python describe different data sets. The site says so wherever generated data appear, but it is still a thing a reader has to keep in mind.
  • The same pattern now covers the inference results: a seeded script writes a committed artefact and a test re-runs it (DR-004).

What I'd change

  • Pin the export's dependencies exactly (uv lock --script) rather than with lower bounds, so a future NumPy cannot change the regenerated file.
  • Re-run export_parity.py in CI and fail on any diff, instead of relying on the parity tests against a committed file.

DR-002 · Accepted · 2026-10-09

Label switching, and how fitted components are named

Applies to: /pitfalls#label-switching, /playground, /inference, web/src/lib/em/labels.ts, web/src/lib/inference

Context

EM numbers its components arbitrarily: "component 1" is whichever bump the starting guess happened to steer towards. The notebook generated group 1 as the sci-fi lovers (mean 7.5), but its fitted component 1 ended up on the low ratings (μ₁ = 4.02). Its final cell then read component 1 as "sci-fi lover" and printed that a user who rates a film 8.5 is probably a romance lover. Its 90.5% accuracy is correct only because the accuracy line happens to encode the labels the switched way.

The upgrade made the problem bigger. A parametric bootstrap refits EM a thousand times and a coverage study refits it hundreds of times. Every refit has to name its components the same way, or the replicates of "μ₁" mix two different quantities.

Decision

  • The notebook's as-run output stays on the site, unchanged and labelled as such.
  • Wherever the truth is known (accuracy, the parameter tables, the playground), fitted components are matched to the true groups by mean before any comparison.
  • In every inference calculation (standard errors, bootstrap, coverage, choosing K), components are ordered by mean, μ₁ < μ₂. In one dimension, ordering by mean and matching to the truth by mean give the same pairing.

Options considered

  1. Fix the notebook's output. Rejected: the original stays as it ran (DR-001's principle).
  2. Order by mean (chosen): the usual identifiability constraint, easy to explain ("component 1 is the low group").
  3. Match each refit to a reference fit by its nearest parameters, or Stephens' relabelling algorithm. More robust when groups overlap, but harder to explain and unnecessary here.
  4. Order by weight. Unstable when the weights are similar.
  5. Report only label-free quantities (the mixture density, the sorted means). Honest but less useful for a reader who wants "the share of romance lovers".

Why

The notebook's groups are far apart (fitted means about 3.4 and 7.0, with standard errors near 0.2 to 0.3), so the ordering constraint never bites and its meaning is obvious to a reader.

What happened

  • In the label-switching census (100 notebook-style random starts, seed 7), 46 runs ended with component 1 as the low-mean group (46%, Wilson 95% CI 36.6% to 55.7%). The switch is close to a coin toss, not bad luck.
  • In the parametric bootstrap (B = 1,000, warm-started at the fit), the two means never crossed: 0 of 1,000 replicates needed relabelling.
  • Under seed 0 the notebook's own accuracy line reports 10.0% for a fit that is 90.0% right once labels are matched. The matched number is the one the site uses.
  • The weak spot: ordering by mean truncates the sampling distribution when the means are close. With overlapping groups (which a visitor can make in the playground) bootstrap intervals for the means would be biased. Nothing on /inference uses such data, but the rule would not be safe there.

What I'd change

  • Use a relabelling algorithm (or report label-free summaries) whenever the bootstrap means come within a few standard errors of each other, and flag it on screen.
  • Have the playground warn when the fitted means are too close for "component 1" to mean anything.

DR-003 · Accepted · 2026-10-09

A variance floor for the inference fits, never for the parity path

Applies to: /inference, /pitfalls#variance-collapse, web/src/lib/inference, web/src/lib/em/em.ts (varianceFloor)

Context

A Gaussian mixture's likelihood has no maximum: put one component on a single rating and shrink its σ, and the log-likelihood goes to +∞. The notebook's data make this easy, because clipping at 10 piled seven ratings onto exactly 10.0. The notebook's own run never got close (its σ stayed above 0.6), and the revival already had an opt-in floor for the collapse demo plus a guard that stops a run once σ falls below 10⁻⁸.

The upgrade fits hundreds of thousands of mixtures: K = 3 and 4 for model choice, and K = 2 on data simulated from a single normal for the likelihood-ratio test. Those are exactly the fits that find spikes. Without a bound some of them degenerate, and simply discarding them would bias the test's null distribution, because the replicates where a spike wins are the ones that matter.

Decision

All inference fits (standard errors, bootstrap, coverage, choosing K, the LRT, convergence studies of simulated data) use a floor of σ ≥ 0.1. The parity code path (fit with the default varianceFloor = 0) is unchanged. Every analysis counts the fits where the floor ended up binding and the page shows those counts.

Options considered

  1. No floor; drop degenerate runs. Biases the LRT null, as above.
  2. Absolute floor, σ ≥ 0.1 (chosen). Interpretable in rating points and far below any real group's spread (the true σ are 1.2 and 1.5).
  3. Relative floor (for example 5% of the sample SD) or scikit-learn's reg_covar = 1e-6 added to each variance. Scale-free, but 10⁻⁶ is far too small to stop a spike on seven identical values.
  4. A prior on σ² (penalised or MAP EM, as in Chen and Tan's penalised likelihood), which also has theory for testing the number of components. The principled option, but a different estimator from the notebook's.
  5. Model the clipping with a censored likelihood, so no component needs to absorb the pile at 10. Fixes the cause rather than the symptom.

Why

The floor is the smallest change that makes every fit well defined while leaving the fits that matter untouched, and it is easy to state and to check.

What happened

  • The floor never bound for the notebook's fit, its 1,000 bootstrap refits, or the coverage-study fits.
  • It decides how much the pile at 10.0 is worth when choosing K. The best three-component fit puts its third component on the clipped ratings with σ held at the floor (4.0% at μ = 9.96, ℓ = −403.23), and BIC picks K = 3 (BIC 848.85, against 851.79 for K = 4 and 854.47 for K = 2); AIC picks 4. The floor prevents degenerate spikes; it does nothing about misspecification.
  • That K = 3 fit was missed at first. Review found that none of the ordinary starts (k-means++, Forgy, random) could isolate seven identical values, so the first version reported a lower K = 3 maximum (ℓ = −409.17) and said both criteria chose K = 4. The fitter now adds starts with a narrow component on any value tied three or more times, and a test pins the K = 3 maximum and BIC's choice.
  • The choice moves with the floor. BIC picks K = 3 at σ ≥ 0.05 and 0.1, and K = 4 at σ ≥ 0.25, where K = 2, 3 and 4 are within 1.1 BIC points of each other. Without the seven clipped ratings the floor never binds and BIC picks K = 2.
  • It did bind in the LRT: in 40 of the 500 data sets simulated under the single-normal null, the best two-component fit had a component at σ = 0.1. The null distribution therefore depends on the floor. Leaving those 40 out moves the bootstrap's 95% point only from 11.49 to 11.10, still far above the χ² value of 7.81, so the conclusion does not hinge on it.

What I'd change

  • Report the LRT for several floors (0.05, 0.1, 0.25) as a sensitivity check, as the choice of K now is.
  • Replace the floor with a weak inverse-gamma prior on σ², which keeps the likelihood bounded without a hard edge.
  • Fit a censored mixture for ratings at 1 and 10, so the pile at 10 stops being a "component".

DR-004 · Accepted · 2026-10-09

Precompute the slow simulations, and test the artefact

Applies to: /inference, web/scripts/generate-inference.ts, web/src/lib/inference/__generated__/inference.json

Context

The inference page needs about 40,000 EM fits for the bootstrap coverage study alone (200 simulated data sets, each fitted once and refitted 200 times in its bootstrap), plus 500 null data sets for the likelihood-ratio test with nine starts each, and 200 data sets for the model-selection rates. On one laptop core that takes about two minutes. The site is static and deployed from web/.

Decision

Keep every seed and size in settings.ts. A script computes all results and writes __generated__/inference.json (about 90 KB). The page imports the artefact. artefact.test.ts recomputes the MLE, standard errors, bootstrap intervals, Wald coverage, model choice and convergence study in full, and checks the slow studies through small versions stored beside them with the same seeds. The bootstrap, the coverage study and the convergence study can be re-run in a Web Worker from the page.

Options considered

  1. Compute at build time in Server Components. One source of truth, but two minutes on every build and on every dev reload of the page.
  2. Compute in the browser on demand. Honest and interactive, but a visitor would wait minutes for the headline numbers, and the page would have nothing to show without JavaScript.
  3. Precompute with a script and test the artefact (chosen), with in-browser re-runs for the fast parts.

Why

It is the same trade as the notebook parity data (DR-001): a committed, regenerable artefact plus a test that fails when the code and the artefact disagree.

What happened

  • Generation takes about 130 seconds; the artefact test takes about 10 seconds.
  • Re-running the bootstrap in the browser with seed 42 reproduces the published intervals exactly, which is the simplest possible reproducibility check for a reader.
  • Same seed does not always mean the same last digit. Review found that Chrome 155 returns correctly rounded Math.exp and Math.log results where Node 26 sometimes differs in the last bit, and the numerical Hessian magnifies that to about 3 × 10⁻⁷ relative in standard errors and interval widths. The coverage re-run in Chrome therefore reported "different numbers" against an up-to-date artefact. The comparison now requires counts (coverage hits, failed fits, iterations) to match exactly and other numbers to agree to 10⁻⁵ relative, and the artefact test compares Hessian-derived values at 10⁻⁶ relative instead of 10⁻⁹, so a future Node with the newer maths library does not fail it for nothing.
  • The weak spot: the slow studies are verified only through their small versions. A change that altered only large-sample behaviour would slip through until someone regenerated, and a regeneration with an edited setting would not be flagged unless the settings test noticed.

What I'd change

  • Regenerate the artefact in a scheduled CI job and fail on any diff.
  • Spread the slow studies across worker threads so the generator finishes in seconds and could run at build time.
  • Compute the Hessian with automatic differentiation or analytic second derivatives, which would make the standard errors far less sensitive to the last bit of exp and log.

DR-005 · Accepted · 2026-10-09

"Explain this iteration": optional, bring your own key, grounded and audited

Applies to: /stepper, /playground, /ai-log, /methods#ai-use, web/src/lib/ai

Context

The lab is a teaching tool, and a plain-language reading of "what just happened in this E-step and M-step" is useful. It is also a small, realistic test of governing generative AI: the explanation is about numbers, which language models are prone to invent. There is no budget for an API key, no server, and nothing on the site may depend on AI.

Decision

  • Bring your own key, browser only. The visitor pastes an Anthropic (default, Claude Haiku 4.5; Claude Sonnet 5.5 optional) or OpenAI key. It stays in sessionStorage unless they choose "remember on this device" (localStorage), and "Forget key" removes it. Calls go straight from the browser to the provider. The key never reaches this site, a log or the repository.
  • Grounded input. The request contains only the iteration's numbers: parameters before and after, responsibilities (all four in the stepper; a summary of the 200 in the playground), the log-likelihood before and after, and the stopping rule.
  • Structured output. The reply must match a JSON schema (provider structured-output modes), validated again with zod against exactly the same contract. Display limits (long fields shortened, at most four caveats) are applied after validation, so a reply that obeys the schema is never thrown away for its length.
  • Refusals. Claude Sonnet 5.5 requests opt into Anthropic's server-side refusal fallback (fallbacks: "default", beta server-side-fallback-2026-07-01): if a safety classifier declines, Anthropic re-runs the request on the model it recommends for that category within the same call, and the audit log records the model that actually answered. Haiku 4.5 does not take the option. A refusal that still comes back is shown as a plain "declined" message.
  • Transparency. Every explanation is labelled "AI-generated", shows the model, latency and token usage, and runs a grounding check that lists numbers not found in what was sent. The check covers numbers with decimals, numbers above 10 and anything in scientific notation (1e-6, 1.5 × 10⁻⁵, 10^-6), matched at the precision they are written with; whole numbers from 0 to 10 are not checked, because they are usually counts, ratings or ordinals.
  • Human in the loop and audit. The visitor accepts, edits or rejects each explanation. Every call is written to an IndexedDB audit log without the key, viewable and exportable (JSON, CSV) at /ai-log. That includes failures: a refusal, a reply cut off at the token limit or one that fails validation keeps its raw reply and token usage, because the visitor paid for it.

Options considered

  1. No AI. Safe, but misses a useful explanation and a governance showcase.
  2. A server proxy with my key. Costs money, invites abuse and puts a secret on a server.
  3. An in-browser open model (WebLLM). No key needed, but a multi-gigabyte download for a paragraph of text.
  4. Bring your own key, from the browser (chosen).

On refusals: leaving the fallback out would keep the browser request a little simpler (no beta header), but a visitor on Sonnet 5.5 would then pay for a declined request and get nothing back. Anthropic's CORS preflight accepts the anthropic-beta header, so the fallback costs nothing in the browser.

Why

It keeps the site free, static and fully functional without AI, and it makes the governance controls concrete: what is sent, what comes back, who decided what, all visible to the visitor. The design is informed by the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency principles and the NIST AI Risk Management Framework; it is not a compliance claim.

What happened

  • The provider adapters, error handling, key storage, audit log and grounding check are unit-tested with the network mocked (ai.test.ts): the key only ever travels in a request header, and a key echoed in a provider error is redacted before it is stored or shown.
  • No live call was made during development or CI, because the project has no key. The request shapes follow the providers' documented APIs; a real call is the first thing to check after deployment.
  • The grounding check is lexical. It catches an invented number, but not a correct number used in a wrong sentence, so it is a prompt for the human reviewer, not a verdict. Review found two gaps before release, both fixed with tests: it read "3.1e-2" as accurate to ±0.05 because it ignored the exponent, and it skipped "3e-2" and "10⁻⁶" as small integers.
  • Review also found that the first version's zod schema was stricter than the JSON schema sent to the provider (length and item limits the providers do not enforce), so a valid five-caveat reply was discarded after the visitor had paid for it, and that failed calls were logged without their reply or usage. Both are fixed.

What I'd change

  • Build a small evaluation set (iterations with known facts: did the log-likelihood rise, did the run converge, which component owns each rating) and score explanations against it, with intervals, per model.
  • Offer an in-browser model as a no-key option once small models are good enough at this.
  • Check numbers semantically as well as lexically: parse each sentence's claim ("the change is below the tolerance") and test it against the snapshot.

DR-006 · Accepted (supersedes DR-005 in part) · 2026-10-10

Say exactly what the AI feature sends, which model answered, and how the site was built

Applies to: /methods#ai-use, /ai-log, README, web/src/lib/ai

Context

A review before release checked the AI use statement against the code and the repository, and found three places where it said less, or more, than was true:

  • DR-005 and /methods said the request contains "only the iteration's numbers". The snapshot also carries the page name, a one-line description of the data set (for example the seed of browser-generated ratings), its size, the run's stopping status and fixed notes on how to read the numbers. None of it is personal, but the statement was not literally accurate.
  • With Claude Sonnet 5.5, Anthropic's server-side fallback can answer a declined request with another model inside the same call. The audit log kept one model field, overwritten with the model that answered, so an auditor could not see what was requested or that a fallback happened.
  • /methods said every other text on the site "was written without" AI. The commit history shows the site's code and text were developed with an AI coding assistant. The runtime claim (only the optional feature calls a model) was true; the authorship claim was not.

Decision

  • What is sent. SNAPSHOT_FIELDS in web/src/lib/ai/explain-iteration.ts lists every field in plain words, /methods prints that list, and a test fails if a stepper or playground snapshot carries a field that is not on it, or the list names one that is never sent.
  • Which model answered. Each audit entry stores requested_model (what the visitor's settings asked for), model (what the provider reports) and fallback (true when a fallback content block or a fallback_message entry in usage.iterations shows the fallback ran). Both new fields are in the JSON and CSV exports and on /ai-log, and the explanation's label names the fallback when it happens. Entries written before this change have only model.
  • How the site was built. /methods and the README say that nothing on the site calls an AI model at runtime except the optional feature, that the code and text were developed with an AI coding assistant (Claude Code), and that every number comes from the code, the tests and the seeded scripts.

Options considered

  1. Trim the snapshot to numbers only, so the old wording became true. The page name and data-set line help the model explain the right thing (four hand-worked ratings, or 200 generated ones), and dropping them would make the explanations worse to save a sentence.
  2. Keep one model field and append "(fallback)" to it. Simple, but it mixes two facts into one string that the CSV export and any later analysis would have to parse.
  3. Separate fields and a field list tested against the code (chosen).

Why

A transparency statement that is wrong in small ways teaches readers to discount it in big ones. Each of these claims can be checked in a minute (by reading the request in the browser's network panel, or the repository's history), so each has to be exactly true. Tying the disclosure to the code with a test keeps it true when the snapshot changes. The design remains informed by, not compliant with, the frameworks named in DR-005.

What happened

  • The three issues were found in review before release, not by a visitor. No live call has been made (the project has no key), so the fallback detection is tested only against mocked responses shaped like Anthropic's documented ones.
  • The log view names the requested model only when the fallback ran. A provider that reports a dated model id (OpenAI often does) is therefore not mistaken for a fallback.

What I'd change

  • Show the exact JSON that will be sent before the first call, with a "send" button, rather than only describing it on /methods.
  • Record the fallback's switch points (which model declined, which continued) in the audit entry, not just a flag, if a fallback ever turns up in practice.

8 · model card

Model card

A model card for the statistical model at the centre of this project: the two-component, one-dimensional Gaussian mixture that the 2025 notebook fits to 200 synthetic movie ratings with the EM algorithm. Numbers here come from re-running the notebook (DR-001) and from the precomputed inference results (DR-004); every interval is a 95% interval.

Model details

  • Model: f(x) = π₁ N(x; μ₁, σ₁²) + π₂ N(x; μ₂, σ₂²), with π₁ + π₂ = 1: five free parameters.
  • Estimation: maximum likelihood by EM, written from scratch in the notebook (EMAnalyzer) and ported line for line to TypeScript (web/src/lib/em/em.ts), tested against the notebook's trace to 10⁻⁶.
  • Author and dates: Sunchuangyu (Rin) Huang. Notebook September 2025; inference and model card October 2026.
  • Licence: MIT.

Intended use

  • Teaching: showing how EM works, step by step, and how a fitted mixture should be reported (with uncertainty, diagnostics and an honest account of what can go wrong).
  • Not intended for segmenting, scoring or making decisions about real people, as a recommender system, or as evidence about real viewers' tastes.

Training data

  • 200 synthetic ratings generated by the notebook with np.random.seed(42): 60% "sci-fi lovers" from N(7.5, 1.2²), 40% "romance lovers" from N(4.0, 1.5²), clipped to [1, 10]. Clipping put seven ratings at exactly 10.0.
  • No real users and no personal information. The data are exported by scripts/export_parity.py with the Python and NumPy versions recorded.
  • The model ignores the clipping: it treats every rating as an unclipped normal draw.

Evaluation

WhatResult
Parity with the notebookEvery π, μ, σ and log-likelihood within 10⁻⁶ for all 15 iterations; console output identical
Accuracy as the notebook reported it (15 iterations; in-sample)90.5% (181 of 200), Wilson 95% CI 85.6% to 93.8%
Accuracy of the converged fit (same start, 90 iterations to the notebook's tolerance; in-sample)83.5% (167 of 200), Wilson 95% CI 77.7% to 88.0%
Converged parameters (low-mean group first), with observed-information SEsπ₁ 0.221 (0.054), μ₁ 3.36 (0.27), μ₂ 7.02 (0.19), σ₁ 0.86 (0.18), σ₂ 1.46 (0.14)
Parametric bootstrap 95% intervals (B = 1,000, seed 42)π₁ 0.110 to 0.359; μ₁ 2.86 to 4.11; μ₂ 6.60 to 7.44; σ₁ 0.48 to 1.22; σ₂ 1.16 to 1.76
Coverage of nominal 95% intervals (simulated data sets of 200)Wald intervals 84.4% to 91.2% depending on the parameter with the model exactly right, 85.2% to 92.0% clipped like the notebook (500 data sets each); percentile bootstrap 90.0% to 91.5% (200 data sets, B = 200 per data set). On the same 200 data sets the bootstrap covered more often for π₁, μ₁ and σ₁ (paired 95% intervals exclude 0); for μ₂ and σ₂ the difference is within simulation noise, and the bootstrap intervals are wider for 4 of the 5 parameters
Number of componentsOn this sample BIC prefers K = 3 and AIC K = 4. BIC's third component is the clipping pile (4.0% at 9.96, σ held at the 0.1 floor); at a floor of 0.25 BIC tips to K = 4 in a near-tie. Without the seven clipped ratings BIC prefers K = 2. On fresh samples from the recipe, clipped like the notebook, BIC picks K = 2 in 91 of 100
K = 1 vs K = 2 (parametric bootstrap LRT, B = 500)2Δℓ = 19.9, bootstrap p = 0.002 (none of 500 null statistics as large). The LRT's 18 starts on the observed data found the notebook's maximum, ℓ₂ = −415.37; the best K = 2 maximum, ℓ = −413.99, would give 22.6, so the statistic understates the evidence; p is unchanged

Both accuracies are in-sample: the 200 ratings that are scored are the ones the model was fitted to. A Wilson interval treats the 200 classifications as independent trials of a fixed rule, so it reflects which users happened to be drawn given the fitted rule, not the uncertainty of the fit itself, and it is optimistic about accuracy on new ratings.

All of these, with their seeds and settings, are on the site at /inference.

Known failure modes

  • Label switching. Component numbers are arbitrary; the notebook's own final cell misnames the groups (DR-002).
  • Stopping early. The notebook's 15 iterations were not convergence; finishing the run lowers the accuracy from 90.5% to 83.5%.
  • Local maxima. 5% of random starts (10 of 200, Wilson 2.7% to 9.0%) reach a different maximum with a higher likelihood (ℓ = −413.99 against −415.37): a narrow component around 8.07. The reported fit is the one the notebook's run was heading to.
  • Unbounded likelihood. A component can collapse onto one value; the clipped pile at 10.0 is a ready-made trap (DR-003).
  • Misspecification. The clipping is not modelled, so model-selection criteria add a component for the pile at 10, and how much that pile is worth depends on the variance floor.
  • Over-confident intervals at n = 200. Wald intervals cover only 84% to 92% of the time, not 95%; bootstrap intervals cover 90% to 92%, still short, and better than Wald by more than simulation noise for three of the five parameters only.
  • Monte Carlo error in the bootstrap coverage study. It uses B = 200 bootstrap replicates per data set for compute reasons, against the usual minimum of 1,000 for percentile intervals, so each endpoint rests on about the 5th and 195th of 200 values. That error is not quantified here, and the bootstrap coverage rates and the paired comparison with Wald could move with a larger B.

Ethical considerations

  • The group names ("sci-fi lovers", "romance lovers") are illustrative labels from the original explainer. Real audiences do not split into two tidy types, and naming clusters after people can harden stereotypes.
  • A responsibility (the probability that a rating came from a component) is a statement about a model, not about a person. It should not be used to label or treat individuals.
  • The optional AI explanation feature sends an iteration's numbers and a fixed description of the page and data set (no personal data) to a provider of the visitor's choosing, with the visitor's own key; see the AI use statement on /methods and DR-006.

Caveats and recommendations

Report a mixture fit with its uncertainty and its diagnostics: the number of starts and how many reached the reported maximum, whether the log-likelihood ever decreased, and how the number of components was chosen. Treat information criteria as evidence, not verdicts, especially when the data were censored, rounded or clipped.

9 · AI use

AI use statement

What the AI does

One optional feature, “Explain this iteration” on the stepper and the playground, writes a short plain-language reading of the EM iteration on screen. It is off until a visitor adds their own API key, and nothing else on the site calls an AI model while you use it.

Building the site is a different matter: its code and text were written with an AI coding assistant (Claude Code), as the commit history shows. Every number on the site comes from the code, the tests and the seeded scripts, not from a language model, and the 2025 explainer and notebook in original/ are kept as they were.

What it never does

  • It never computes or changes any number shown on the site.
  • It never sees anything except this iteration's numbers and the fixed description of the page and data set listed below. No personal data is sent.
  • It never runs without a click, and never with a key belonging to this site (there is none).
  • Its output is never shown without the “AI-generated” label.

What is sent, and where

A fixed system prompt and a JSON snapshot with these fields and no others (a test checks the snapshot against this list):

  • the page name (stepper or playground)
  • a one-line description of the data set (the notebook's synthetic ratings, or ratings generated in the browser with their seed) and its size
  • the iteration number and, on the stepper, which half of it is on screen
  • the shares, means and spreads before and after the iteration
  • the responsibilities: all four ratings' on the stepper; shares, counts and five sample ratings' on the playground
  • the log-likelihood before and after, and the change
  • the stopping rule: tolerance, iteration cap and the run's status
  • fixed notes on how to read these numbers

It goes straight from the browser to the provider the visitor chose: Anthropic (default Claude Haiku 4.5, or Claude Sonnet 5.5) or OpenAI (default gpt-5-mini, editable). The key is kept in the browser's session storage, or local storage if the visitor asks, and travels only in the request header.

Human in the loop, and the record

The reply must match a JSON schema and is validated again in the browser. A grounding check lists any number in it that was not in the snapshot; it checks numbers with decimals, numbers above 10 and anything in scientific notation; whole numbers from 0 to 10 are not checked. The visitor accepts, edits or rejects each explanation; one replaced by asking again, or still undecided when the visitor leaves the page, is closed in the log as superseded or abandoned rather than left pending. Every call is written without the key to an audit log in the browser's IndexedDB: time, feature, provider, the model asked and the model that answered, input, output, latency, token usage and the decision. Failed calls are recorded too, and a refusal, a cut-off reply or one that failed validation keeps whatever the provider sent back and its token usage. On Claude Sonnet 5.5 a declined request is retried by Anthropic's server-side fallback within the same call; the log records the model that was asked, the model that answered and whether the fallback ran. View or export it on the AI audit log.