hv-falsification-os

A NumPy-only toolkit for validating measurement claims. Answers eight methodological questions per call, each with its own null. Packages the measurement discipline from the XuanJi SKILLS archive into callable functions.

Size: ~22 KB source, no model artifact (stateless). Dependencies: NumPy only.

The eight checks

check question output
measure_cv What is the noise of my metric? mean, std, CV, 95% CI
same_code_test Is this delta real? delta, ratio-to-CV, verdict
amortization_sweep Where is the optimum? peak param, peak value, full curve
noise_floor What is the smallest detectable signal? floor at 2Οƒ
precondition_test Does the technique apply here? pass rate, holds
rejection_rate Does the test actually reject anything? rate, informative flag
reversal_detection Is the winner stable across runs? reversal flag, interpretation
four_question_probe Is rank divergence present? effective rank, cross-cov rank, probe RΒ²

Plus one static helper:

  • five_percent_rule(rel_delta, cv) β€” classify a delta against a known CV.

Source discipline

The three rules the toolkit packages, from the SKILLS archive:

Ch 2 β€” Sweep before polish. Any "new method" must be compared against a working baseline before it is polished. The amortization_sweep check finds the peak of a parameter-versus-ratio curve and reports whether the peak beats the baseline. In the benchmark, bit-slicing peaks at 0.589 β€” below 1.0, i.e. loses to scalar. That is the correct finding, and polish would never have surfaced it.

Ch 3 β€” Measure the noise floor before believing the delta. Every "result" inside 2Γ— CV is noise. measure_cv returns the CV of any callable; same_code_test returns the delta-over-CV ratio and a verdict in {noise, marginal, real}. noise_floor gives the minimum detectable signal for a metric at a chosen confidence level.

Ch 5 β€” State and test the precondition before implementing. Every technique has a hidden precondition. precondition_test runs a boolean function on N independently-seeded instances and reports the pass rate. The benchmark demonstrates the precondition check on clustered vs uniform data: the same test that passes on clustered data fails on uniform.

Headline numbers

benchmark result
CV calibration (5 magnitudes, N=200) all within 2.5% of designed
Same-code discrimination (9 cases) 9/9 verdicts match design
Amortization peak recovery 0.589 at param 500
Precondition test 100% pass on clustered, 0% on uniform
Rejection rate on informative tests 31% on avalanche, 0% on bias
Reversal detection 3 cases, all correctly classified
Four-question probe effective_rank = 62, probe RΒ² = 0.91
21/21 consistency checks βœ“

The five-percent rule

The empirical rule that came out of the SKILLS Chapter 3 work:

| ratio (|Ξ”| / CV) | verdict | |---|---| | < 2.0 | noise β€” do not report | | 2.0–3.0 | marginal β€” increase samples | | > 3.0 | real β€” the delta is measurable |

Implemented in same_code_test and exposed as the static method five_percent_rule.

How to use

from hv_falsification_os import FalsificationOS

os_ = FalsificationOS()

# 1. Measure the noise of a metric
cv = os_.measure_cv(my_callable, n_trials=500)
# cv = {'mean': ..., 'std': ..., 'cv': 0.048, ...}

# 2. Test whether a delta is real
r = os_.same_code_test(fn_baseline, fn_candidate, n_trials=500)
# r['verdict'] = 'real' | 'marginal' | 'noise'
# r['ratio_to_cv'] = 3.4

# 3. Sweep a parameter and find the peak
s = os_.amortization_sweep(fn, [1, 10, 100, 1000, 10000], n_trials=30)
# s['peak_param'] = 500, s['peak_value'] = 0.589

# 4. Compute the noise floor
nf = os_.noise_floor(metric_fn, n_trials=500, confidence=2.0)
# nf['floor'] = 0.098 β€” below this, nothing is detectable

# 5. Test a precondition before implementation
p = os_.precondition_test(my_boolean_test, n_trials=30)
# p['holds'] = True | False

# 6. Check if a test actually rejects anything
rr = os_.rejection_rate(my_test_fn, n_designs=200)
# rr['is_informative'] = True if 2% < rate < 98%

# 7. Detect reversals between runs
rv = os_.reversal_detection((1.0, 1.10), (1.02, 1.11))
# rv['reversal'] = False

# 8. Four-question probe on a feature matrix
probe = os_.four_question_probe(H, Y)
# probe['effective_rank_H'] = 62.0
# probe['cross_cov_rank'] = 1.0
# probe['fresh_probe_r2'] = 0.91
Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support