Lab catalogue

8 labs are live and 34 more are mapped to the textbook. Clusters follow the fourteen DSPA3 chapters, so a lab can be dropped into a lecture or a homework set without rearranging anything else.

Foundations & Data Quality

Chapters 1, 2

Simulation, reproducibility, floating point, missingness mechanisms, robustness, and multiplicity.

Foundations, Data Science, Ethics, and the R Environment · Data Quality, Missingness, and Exploratory Visual Analytics

Ch 1 · 1.4
Simulator

Sampling Distributions & the Limits of the CLT

Separate the distribution of the data from the distribution of a statistic, and find where the central limit theorem stops applying.

15 min
Ch 2 · 2.3
Misconception

Missingness Mechanisms: MCAR, MAR, MNAR

Delete data three different ways and watch which repairs work, which fail, and why more data never fixes a biased mechanism.

18 min
Ch 1 · 1.4
Planned

Seeds, Streams, and Reproducibility

Why the same code gives different answers, and how a seed turns a random study into a repeatable one.

12 min
Ch 1 · 1.6
Planned

Floating Point Has Edges

Catastrophic cancellation, non-associative sums, and the variance formula that returns a negative number.

15 min
Ch 1 · 1.8
Planned

Fairness Criteria Cannot All Hold

Adjust a classifier and watch demographic parity, equal opportunity, and calibration compete.

20 min
Ch 2 · 2.4
Planned

Breakdown Points

Drag one outlier and see which summaries follow it and which refuse to.

12 min
Ch 2 · 2.7
Planned

The Same Data, Six Charts

Axis limits, binning, aspect ratio, and colour scale — each one changes the story.

15 min
Ch 2 · 2.9
Planned

Testing Many Things at Once

Family-wise error, false discovery rate, and the garden of forking paths.

18 min

Mathematical Core

Chapters 3, 4

Matrix computing, conditioning, projection geometry, and linear/nonlinear dimensionality reduction.

Linear Algebra, Matrix Computing, and Regression · Linear and Nonlinear Dimensionality Reduction

Ch 3 · 3.5
Diagnostic

Conditioning, Collinearity, and Unstable Coefficients

Watch the condition number of a design matrix explode and see exactly which quantities become unidentifiable — and which do not.

20 min
Ch 4 · 4.2
Explorer

PCA Geometry: Projection, Scale, and What Variance Means

Rotate a data cloud, change a measurement unit, and see principal components move — the projection is geometry, not magic.

18 min
Ch 3 · 3.2
Planned

Matrices as Motions

Watch a matrix act on the unit circle: rotation, scaling, shear, and collapse.

15 min
Ch 3 · 3.5
Planned

Regression Is a Projection

The hat matrix, the residual space, and why the residuals are orthogonal to the fit.

18 min
Ch 4 · 4.2
Planned

SVD as an Optimal Summary

Rank-k truncation on an image and a data matrix, with the error bound made visible.

20 min
Ch 4 · 4.6
Planned

Nonlinear Embeddings Invent Structure

Change perplexity and neighbours on data with no clusters, then watch clusters appear anyway.

22 min

Supervised Learning

Chapters 5, 6

Bayes error, kNN boundaries, metric families, leakage, kernels, ensembles, and backpropagation.

Supervised Classification · Black-Box Methods: Neural Networks, SVM, and Ensembles

Ch 5 · 5.6
Comparator

Accuracy Is a Trap: Metrics Under Class Imbalance

Move a threshold, move the prevalence, and watch which performance numbers stay honest and which quietly lie.

20 min
Ch 5 · 5.2
Planned

The Irreducible Floor

Overlapping class densities set a limit no classifier can pass. Try to beat it.

15 min
Ch 5 · 5.4
Planned

k Controls the Boundary

From jagged memorization at k = 1 to a nearly linear boundary at large k.

15 min
Ch 5 · 5.8
Planned

Where Leakage Hides

Impute before splitting, select features on all the data, or split a time series at random — and watch validation lie.

22 min
Ch 6 · 6.4
Planned

The Kernel Trick, Drawn

Lift two-dimensional data into a feature space where a plane separates it.

20 min
Ch 6 · 6.6
Planned

Bagging, Boosting, and Variance

Grow trees one at a time and watch which method reduces bias and which reduces variance.

25 min

Unsupervised Learning & Text

Chapters 7, 8

TF-IDF geometry, association rules, k-means, DBSCAN, spectral clustering, and mixtures.

Text Mining, NLP, and Association Rule Learning · Unsupervised Clustering

Ch 8 · 8.4
Comparator

Cluster Geometry: k-means, DBSCAN, and Shapes That Break Them

Compare centroid and density clustering on blobs, moons, and pure noise — and see why the elbow plot cannot tell you k.

22 min
Ch 7 · 7.3
Planned

Documents as Vectors

Term weighting, cosine similarity, and why raw counts mislead.

18 min
Ch 7 · 7.6
Planned

Support, Confidence, and Lift

A high-confidence rule can still be worthless. Lift explains why.

18 min
Ch 8 · 8.5
Planned

Clustering Through the Graph Laplacian

Non-convex shapes that defeat k-means fall apart cleanly in the spectral embedding.

22 min
Ch 8 · 8.7
Planned

EM, Step by Step

Soft assignments, monotone likelihood, and convergence to the wrong local optimum.

22 min

Validation & Feature Selection

Chapters 9, 11

Resampling, calibration versus discrimination, decision curves, LASSO paths, and knockoffs.

Model Performance Assessment and Validation · Variable Importance and Controlled Feature Selection

Ch 9 · 9.5
Misconception

Calibration Is Not Discrimination

Distort a model's probabilities without touching its ranking: AUC does not move a decimal place while every probability becomes wrong.

18 min
Ch 11 · 11.3
Explorer

LASSO Paths, Cross-Validation, and False Discoveries

Trace coefficients from saturated to empty, then discover that the cross-validated λ is tuned for prediction — not for finding the right variables.

25 min
Ch 9 · 9.2
Planned

Which Resampling Scheme?

Holdout, k-fold, repeated CV, LOOCV, and the bootstrap, scored on bias, variance, and cost.

22 min
Ch 9 · 9.7
Planned

Net Benefit and Threshold Choice

Accuracy cannot pick a threshold. A loss function can.

20 min
Ch 11 · 11.6
Planned

Controlled Selection with Knockoffs

Manufacture decoy variables, then use them to bound the false discovery rate.

25 min
Ch 11 · 11.2
Planned

Importance Depends on the Question

Impurity, permutation, and SHAP-style attributions disagree — especially under correlation.

22 min

Systems & Performance

Chapters 10

Columnar formats, chunking, Amdahl's law, and streaming computation in the browser.

Big Data, Formats, and Computational Performance

Ch 10 · 10.4
Planned

One Pass, Bounded Memory

Streaming means, variances, quantile sketches, and count-min counting.

20 min
Ch 10 · 10.7
Planned

Amdahl's Ceiling

Add workers to a job with a serial fraction and watch the speedup flatten.

15 min

Temporal & Longitudinal

Chapters 12

Autocorrelation, stationarity, filtering, survival curves, and forecast validation.

Time Series, Longitudinal, and Survival Analysis

Ch 12 · 12.2
Planned

Dependence Breaks the Standard Error

Serially correlated data with n = 500 can carry the information of n = 40.

20 min
Ch 12 · 12.5
Planned

Backtesting Without Peeking

Rolling origins, expanding windows, and the horizon at which skill disappears.

22 min
Ch 12 · 12.8
Planned

Censoring Is Information

Kaplan-Meier, hazard ratios, and why dropping censored cases biases everything.

22 min

Optimization

Chapters 13

Loss landscapes, gradient descent variants, constraints, and Bayesian optimization.

Optimization and Numerical Methods

Ch 13 · 13.3
Planned

Walking Downhill

Step size, momentum, and conditioning on a landscape you can rotate and stretch.

20 min
Ch 13 · 13.6
Planned

Constraints Bind

Lagrange multipliers and KKT conditions made visual on a feasible region you draw.

22 min
Ch 13 · 13.8
Planned

Where to Sample Next

A Gaussian process surrogate and the exploration-exploitation dial.

25 min

Deep Learning

Chapters 14

Backpropagation, convolution, sequence models, generative models, and generalization.

Deep Learning and Representation Learning

Ch 14 · 14.2
Planned

Backpropagation, One Edge at a Time

A tiny network with every forward value and every gradient shown on the graph.

25 min
Ch 14 · 14.4
Planned

What a Filter Sees

Design a kernel, slide it, and read the feature map it produces.

20 min
Ch 14 · 14.8
Planned

Overparameterized and Still Improving

Push past the interpolation threshold and watch test error fall a second time.

22 min