hv-falsification-os
A NumPy-only toolkit for validating measurement claims. Answers eight methodological questions per call, each with its own null. Packages the measurement discipline from the XuanJi SKILLS archive into callable functions.
Size: ~22 KB source, no model artifact (stateless). Dependencies: NumPy only.
The eight checks
| check | question | output |
|---|---|---|
measure_cv |
What is the noise of my metric? | mean, std, CV, 95% CI |
same_code_test |
Is this delta real? | delta, ratio-to-CV, verdict |
amortization_sweep |
Where is the optimum? | peak param, peak value, full curve |
noise_floor |
What is the smallest detectable signal? | floor at 2Ο |
precondition_test |
Does the technique apply here? | pass rate, holds |
rejection_rate |
Does the test actually reject anything? | rate, informative flag |
reversal_detection |
Is the winner stable across runs? | reversal flag, interpretation |
four_question_probe |
Is rank divergence present? | effective rank, cross-cov rank, probe RΒ² |
Plus one static helper:
five_percent_rule(rel_delta, cv)β classify a delta against a known CV.
Source discipline
The three rules the toolkit packages, from the SKILLS archive:
Ch 2 β Sweep before polish. Any "new method" must be compared against
a working baseline before it is polished. The amortization_sweep check
finds the peak of a parameter-versus-ratio curve and reports whether the
peak beats the baseline. In the benchmark, bit-slicing peaks at 0.589 β
below 1.0, i.e. loses to scalar. That is the correct finding, and polish
would never have surfaced it.
Ch 3 β Measure the noise floor before believing the delta. Every
"result" inside 2Γ CV is noise. measure_cv returns the CV of any
callable; same_code_test returns the delta-over-CV ratio and a verdict
in {noise, marginal, real}. noise_floor gives the minimum detectable
signal for a metric at a chosen confidence level.
Ch 5 β State and test the precondition before implementing. Every
technique has a hidden precondition. precondition_test runs a boolean
function on N independently-seeded instances and reports the pass rate.
The benchmark demonstrates the precondition check on clustered vs uniform
data: the same test that passes on clustered data fails on uniform.
Headline numbers
| benchmark | result |
|---|---|
| CV calibration (5 magnitudes, N=200) | all within 2.5% of designed |
| Same-code discrimination (9 cases) | 9/9 verdicts match design |
| Amortization peak recovery | 0.589 at param 500 |
| Precondition test | 100% pass on clustered, 0% on uniform |
| Rejection rate on informative tests | 31% on avalanche, 0% on bias |
| Reversal detection | 3 cases, all correctly classified |
| Four-question probe | effective_rank = 62, probe RΒ² = 0.91 |
| 21/21 consistency checks | β |
The five-percent rule
The empirical rule that came out of the SKILLS Chapter 3 work:
| ratio (|Ξ| / CV) | verdict | |---|---| | < 2.0 | noise β do not report | | 2.0β3.0 | marginal β increase samples | | > 3.0 | real β the delta is measurable |
Implemented in same_code_test and exposed as the static method
five_percent_rule.
How to use
from hv_falsification_os import FalsificationOS
os_ = FalsificationOS()
# 1. Measure the noise of a metric
cv = os_.measure_cv(my_callable, n_trials=500)
# cv = {'mean': ..., 'std': ..., 'cv': 0.048, ...}
# 2. Test whether a delta is real
r = os_.same_code_test(fn_baseline, fn_candidate, n_trials=500)
# r['verdict'] = 'real' | 'marginal' | 'noise'
# r['ratio_to_cv'] = 3.4
# 3. Sweep a parameter and find the peak
s = os_.amortization_sweep(fn, [1, 10, 100, 1000, 10000], n_trials=30)
# s['peak_param'] = 500, s['peak_value'] = 0.589
# 4. Compute the noise floor
nf = os_.noise_floor(metric_fn, n_trials=500, confidence=2.0)
# nf['floor'] = 0.098 β below this, nothing is detectable
# 5. Test a precondition before implementation
p = os_.precondition_test(my_boolean_test, n_trials=30)
# p['holds'] = True | False
# 6. Check if a test actually rejects anything
rr = os_.rejection_rate(my_test_fn, n_designs=200)
# rr['is_informative'] = True if 2% < rate < 98%
# 7. Detect reversals between runs
rv = os_.reversal_detection((1.0, 1.10), (1.02, 1.11))
# rv['reversal'] = False
# 8. Four-question probe on a feature matrix
probe = os_.four_question_probe(H, Y)
# probe['effective_rank_H'] = 62.0
# probe['cross_cov_rank'] = 1.0
# probe['fresh_probe_r2'] = 0.91
- Downloads last month
- 16