Practical Data Analysis

Sources and evidence

These references support the definitions, methods, and further reading in the book. All worked datasets and figures are original or drawn from the pinned scheduler simulation. Generated teaching examples are labeled synthetic. Source pages were checked on 2 October 2026; official documentation may evolve independently of the companion runtime.

S01. OpenTelemetry traces

Used in chapter 2. Trace, span, and attribute concepts for instrumentation.

S02. Prometheus histograms and summaries

Used in chapter 2. Aggregation limits of precomputed quantiles and histograms.

S03. Wickham, Tidy Data

Used in chapter 3. One variable per column, one observation per row, and one table per observational unit.

S04. pandas merge documentation

Used in chapter 3. Join cardinality validation and pandas null-key behavior.

S05. pandas scaling guidance

Used in chapter 3. In-memory scaling, reducing loaded data, and chunking.

S06. NIST exploratory data analysis

Used in chapter 4. Exploratory graphics and assumption checking.

S07. NIST intervals for paired differences

Used in chapter 5. Paired mean-difference confidence intervals.

S08. SciPy bootstrap

Used in chapter 5. Bootstrap methods and paired resampling.

S09. ASA statement on p values

Used in chapter 5. Interpretation limits of p values and threshold-only decisions.

S10. NIST multiple comparisons

Used in chapter 5. Bonferroni control for preselected multiple comparisons.

S11. Causal Inference What If

Used in chapter 7, 13. Counterfactual causal framework and further reading.

S12. NIST design selection

Used in chapter 7. Comparative, screening, and response-surface experiment objectives.

S13. Columbia notes on Little's law

Used in chapter 8. Little’s law, consistent system boundaries, and time averages.

S14. NIST process monitoring

Used in chapter 8. Control limits versus specification limits.

S15. Forecasting Principles and Practice on time-series cross-validation

Used in chapter 8. Time-ordered rolling-origin evaluation.

S16. Google SRE monitoring

Used in chapter 8. Latency, traffic, errors, saturation, and symptoms versus causes.

S17. Demystifying evals for AI agents

Used in chapter 9. Agent task, trial, grader, transcript, and outcome distinctions.

S18. Zheng and colleagues on LLM judges

Used in chapter 9. Documented judge position, verbosity, and self-enhancement biases; no universal accuracy claim.

S19. scikit-learn probability calibration

Used in chapter 9. Calibration versus discrimination and reliability diagrams.

S20. scikit-learn common pitfalls

Used in chapter 10. Training-only preprocessing and leakage prevention.

S21. scikit-learn cross-validation

Used in chapter 10. Grouped and time-aware cross-validation.

S22. Kapoor and Narayanan on leakage

Used in chapter 10. Leakage risks in machine-learning-based science.

S23. scikit-learn model-evaluation metrics

Used in chapter 10. Classification and regression metric definitions.

S24. Model Cards for Model Reporting

Used in chapter 10. Structured reporting of model use, performance, and limitations.

S25. W3C PROV primer

Used in chapter 11. Entities, activities, and agents in provenance.

S26. Rubin on inference with missing data

Used in chapter 11. Missing-data mechanisms and the conditions behind ignoring missingness.

S27. GeoPandas projections

Used in chapter 11. Coordinate reference systems and assigning versus transforming coordinates.

S28. Gebru and colleagues on datasheets

Used in chapter 11. Dataset motivation, composition, collection, use, and documentation.

S29. Python for Data Analysis

Used in chapter 13. Legally available author online book and practical reading path; no text reproduced.

S30. Fundamentals of Data Visualization

Used in chapter 13. Author manuscript for visualization reading; no figures or text reproduced.

S31. OpenIntro Statistics

Used in chapter 13. Foundational statistics reading recommendation and official availability.

S32. NIST engineering statistics

Used in chapter 13. Engineering-statistics reference reading recommendation.

S33. An Introduction to Statistical Learning

Used in chapter 13. Authors’ statistical-learning book and Python edition reading recommendation.

S34. Forecasting Principles and Practice

Used in chapter 13. Forecasting book reading recommendation.

M01. MIT projections and least squares

Used in chapter 14. Geometric connection between projection and least-squares residuals.

M02. NIST mean vector and covariance matrix

Used in chapter 14. Sample covariance definition and observation/feature orientation.

M03. NumPy covariance

Used in chapter 14. rowvar orientation and covariance API.

M04. MIT singular value decomposition

Used in chapter 14. Matrix factorization and singular directions; original example derived separately.

M05. NumPy singular value decomposition

Used in chapter 14. Returned U, singular values, Vh and reconstruction conventions.

M06. scikit-learn clustering guide

Used in chapter 14. Clustering objectives and contrasting method assumptions.

M07. scikit-image image data types

Used in chapter 15. Image dtype and range conventions.

M08. SciPy multidimensional convolution

Used in chapter 15. Convolution and boundary-mode concepts; companion implements its own explicit NumPy reflection convention.

M09. NumPy discrete Fourier transform conventions

Used in chapter 15. Frequency ordering, transform normalization, real-valued transform representation.

M10. scikit-image morphology

Used in chapter 15. Erosion, dilation, opening, closing and footprints.

M11. scikit-image region measurements

Used in chapter 15. Connected-region measurement, spacing, area and coordinates.

M12. Boyd and Vandenberghe Convex Optimization

Used in chapter 18. Author-hosted legal book availability and further study of convex problems.

M13. SciPy optimization guide

Used in chapter 18. Distinguishing solver families and problem structures.

M14. NIST law of propagation of uncertainty

Used in chapter 18. First-order sensitivities, variances, and covariance terms.

M15. NIST uncertainty budgets and sensitivity coefficients

Used in chapter 18. Relating input uncertainty to output uncertainty contributions.

M16. JCGM 101 Monte Carlo propagation

Used in chapter 18. Propagation of distributions, conditions and separation from validity of the input model.

G01. Projections

Used in chapter 16. CRS metadata, declaring versus transforming coordinates, and GeoPandas longitude/latitude order. Does not establish the accuracy of any particular local projection.

G02. Geodesic calculations

Used in chapter 16. Distinguishing ellipsoidal geodesics from the chapter’s explicitly spherical approximation. No claim that the synthetic sphere is survey grade.

G03. Merging data

Used in chapter 16. Attribute versus spatial joins, geometric predicates, one-to-many matches and nearest-join options. Rectangle half-open assignment is an original teaching convention.

G04. Geotransform tutorial

Used in chapter 16. Six affine coefficients, top-left corner coordinates, north-up negative pixel height and half-cell center offset. Zonal overlap example independently derived.

G05. GDAL Grid tutorial

Used in chapter 16. Inverse distance to a power interpolation and neighborhood parameters. Chapter fixture uses all points, p=2, zero smoothing; measurement-only variance is separately derived.

G06. How Kriging works

Used in chapter 16. Kriging relies on a modeled spatial relationship; prediction error is conditional on model assumptions. No kriging implementation or empirical calibration is claimed.

G07. Cross-validation strategies for data with temporal spatial hierarchical or phylogenetic structure

Used in chapter 16. Structured validation must reflect interpolation versus transfer to new space; dependence can make random validation misleading. Not a universal endorsement of every blocking design.

G08. Global Spatial Autocorrelation with Moran’s I

Used in chapter 16. Moran statistic, spatial lag and random-label permutation reference. Chapter explicitly fixes locations and weights and defines its own right tail; it does not assert generic point-process CSR or reproduce software default p-value semantics.

G09. Local Spatial Autocorrelation 1

Used in chapter 16. Local quadrant membership differs from significance; local tests raise multiplicity and permutation-resolution issues. Chapter implements descriptive local values only.

G10. Choropleth Map Design for Cancer Incidence Part 2

Used in chapter 16. Warnings about ecological inference, small denominators and geographic pattern interpretation. Chapter has no disease analysis and its four-person numerical counterexample is original.

F01. Present Value Relations Slides 1–36

Used in chapter 17. Present value, compounding and real/nominal consistency. All project cash flows, break-even numbers and IRR counterexample are original calculations; no recommended discount rate.

F02. Portfolio Theory Slides 1–46

Used in chapter 17. Weighted portfolio returns and the role of covariance in mean-variance analysis. Matrices and return observations are synthetic, not estimates for named assets.

F03. Foundations of Portfolio Theory

Used in chapter 17. Historical foundation and conditional role of mean-variance portfolio criteria. Does not supply the chapter’s numerical portfolio or current market claims.

F04. The Statistics of Sharpe Ratios

Used in chapter 17. Sharpe estimates have sampling error; simple annualization is not generally valid under serial correlation. The additive-sum variance identity is derived independently in the chapter.

F05. On the coherence of Expected Shortfall

Used in chapter 17. Expected shortfall definitions need care for discontinuous distributions; precise tail probability mass. Chapter uses upper-tail loss confidence alpha, which must not be confused with an author’s lower-tail probability convention.

F06. How Fees and Expenses Affect Your Investment Portfolio

Used in chapter 17. Fees reduce returns and the capital left to compound. Does not justify the toy 0.001 transaction cost as realistic or current.

F07. The Probability of Backtest Overfitting

Used in chapter 17. Trying many alternatives creates selection risk even without a direct look-ahead bug; retain experiment history. Chapter does not implement CSCV or estimate a PBO.

F08. The Delisting Bias in CRSP Data

Used in chapter 17. Historical research example of missing delisting outcomes. Supports auditing terminal observations, not a claim that current CRSP data still has the reported defect.