Validation ledgers#

Validation is a release requirement, not a retrospective exercise. These records connect each public implementation to its hypotheses, primary-paper formula, calibration, literal or trusted oracle, numerical edge cases, invariants, and known differences from SHT 0.1.9.

Mean tests

Classical formulas, high-dimensional trace and diagonal statistics, Behrens–Fisher approximations, randomization, Bayes factors, multi-group procedures, and the validation-blocked sparse maximum-test audit.

Classical mean tests: formula and validation ledger
Variance tests

Alternative-tail coverage, robust scale handling, degenerate samples, and independent SciPy comparisons.

Classical variance tests: formula and validation ledger
Covariance tests

Null-covariance whitening, literal U-statistics, two-sided projection tests, multi-group formulas, and conditional-regression Bayes factors.

Covariance tests: formula and validation ledger
Mean and variance

Joint normal-population likelihood ratios, corrected rejection tails, component combinations, and stable exact quadrature.

Joint mean and variance: formula and validation ledger
Mean and covariance

Nonidentity-null whitening, high-dimensional centering, classical LRT checks, and unequal-sample trace estimators.

Joint mean and covariance: formula and validation ledger
Distributional equality

Exact enumeration, corrected Monte Carlo inference, floating-point ties, scale stability, and a literal distance-statistic oracle.

Biswas–Ghosh (2014) two-sample test: validation record
Normality

Shapiro approximations, moment formulas, finite-sample null-size audits, and Monte Carlo defaults.

Normality tests: formula and validation ledger
Rectangular uniformity

Interpoint moments, quantile transforms, support boundaries, and asymptotic versus Monte Carlo calibration.

Rectangular-uniformity tests: formula and validation ledger
Simplex uniformity

Dirichlet likelihoods, symmetric and general optimization, strict boundaries, and Wilks-regime checks.

Simplex uniformity: formula and validation ledger

Multivariate-mean method ledgers#

What the evidence labels mean#

Evidence

Question answered

Formula ledger

Does the code evaluate the intended finite-sample quantity?

Hand fixture

Can a reader reproduce at least one result independently?

Trusted oracle

Does a separate established implementation agree where contracts overlap?

Literal reference

Does an intentionally simple implementation reproduce the optimized path?

Invariance check

Does the result respect transformations implied by the mathematics?

Numerical stress test

Does finite precision preserve inferential ordering and physical units?

Null simulation

Is an approximation calibrated in the regime where it is advertised?

No single row is sufficient by itself. The appropriate combination depends on the method and its calibration. A public asymptotic option remains labeled as an approximation when finite-sample simulation does not support a stronger claim.

The release-simulation runner records the maintained scenario and random-stream contracts for the null and targeted alternative audits, and explains why historical Fisher rows are excluded from the 0.1.0 release evidence.

Legacy corrections#

The legacy audit records confirmed defects and how pySHT addresses them. Important corrections include null-covariance whitening, literal scale-equivariant covariance U-statistics, genuine two-sided Wu–Li tests, corrected likelihood-ratio tails, fixed CPH group indexing, defined AJB/RJB defaults with Monte Carlo uncertainty, fixed auxiliary randomness during permutation, log-domain exact and Bayesian calculations, and removal of the invalid distribution-equality asymptotic branch.