python3 -m pytest tests/ -q ........................................................................ [ 11%] ........................................................................ [ 23%] ........................................................................ [ 35%] ........................................................................ [ 46%] ........................................................................ [ 58%] ........................................................................ [ 70%] ........................................................................ [ 81%] ........................................................................ [ 93%] ....................................... [100%] ---------------------------- pytest-regtest report ----------------------------- total number of failed regression tests: 0 total number of failed snapshot tests : 0 615 passed in 466.90s (0:07:46) python3 bench/audit.py ========================================================================== AUDIT — checking our claims against our artifacts, not against ourselves ========================================================================== [1] statistics vs scipy PASS t_ppf matches scipy to 1.92e-13 over 54 points PASS t_sf matches scipy to 2.11e-15 over 35 points PASS paired_t p-value matches ttest_1samp to 1.28e-07 PASS MDE at n=3 realises power 0.8000 (want 0.80), method=exact_noncentral_t PASS MDE at n=5 realises power 0.8000 (want 0.80), method=exact_noncentral_t PASS MDE at n=10 realises power 0.8000 (want 0.80), method=exact_noncentral_t [2] vendored files and provenance PASS vendor/eprocess.py byte-identical to statefabric/stats/eprocess.py PASS vendor/provenance_estate.py byte-identical to remote/_provenance.py PASS vendor/seal.py byte-identical to verifycore/seal.py PASS vendor/bench_harness_estate.py byte-identical to tools/honestbench/bench_harness.py PASS vendor/mooncake_loader.py byte-identical to statefabric/experiments/mooncake_loader.py NOTE estate HEAD has moved to 515540069e82; PROVENANCE.md records an older one. NOT a failure -- the content hashes above are the binding check. Refresh the note when convenient. [3] document figures vs measured artifacts PASS E1 A/A null: every quoted figure matches the artifact PASS E4a starved: every quoted figure matches the artifact PASS E4b eager: every quoted figure matches the artifact (87 individual figures checked) [4] our own certificates carry provenance PASS active_expert_ceiling.json: stamped PASS agentic_public_traces.json: stamped PASS agentic_residency_cc.json: stamped PASS agentic_residency_tracelab.json: stamped PASS cache_ledger_cc.json: stamped PASS cache_ledger_tracelab.json: stamped PASS d1_engine_hit_rate.attempt1_no_counters.json: stamped PASS d1_engine_hit_rate.attempt2_timestamp_as_counter.json: stamped PASS d1_engine_hit_rate.attempt3_broken_dataclass.json: stamped PASS d1_engine_hit_rate.json: stamped PASS d1_engine_hit_rate.prereg.json: sealed PASS d1_engine_hit_rate.prereg.json: seal re-derives from its own recorded threshold PASS d2_positions_per_pass.json: stamped PASS d3_determinism_2x2.json: stamped PASS d4_step_law_profile.attempt1_no_profiler.json: stamped PASS d4_step_law_profile.attempt3_zero_events.json: stamped PASS d4_step_law_profile.json: stamped PASS d5_knee_binding_by_shards.json: stamped PASS d6_isolation_tax.json: stamped PASS d7_admission_above_knee.json: stamped PASS d8_expert_rank.json: stamped PASS estate_acceptance_triples.json: stamped PASS estate_axiom_coverage.json: stamped PASS estate_imports.json: stamped PASS estate_negatives_and_pins.json: stamped PASS estate_orphan_theorems.json: stamped PASS estate_publish_force_risk.json: stamped PASS estate_published_datasets.json: stamped PASS estate_release_readiness.json: stamped PASS estate_verify_all.json: stamped PASS fleet_receipt_provenance.json: stamped PASS fleet_report_091494c4.json: stamped PASS fleet_report_92041851.json: stamped PASS fleet_report_a465f767.json: stamped PASS fleet_report_c8e16960.json: stamped PASS fleet_report_e636fdfb.json: stamped PASS fleet_report_validation.json: stamped PASS kv_footprint_ladder.json: stamped PASS l2_adaptive_ttl.json: stamped PASS lossless_probe.json: stamped PASS lossless_throughput.json: stamped PASS mooncake_reconciliation.json: stamped PASS multimodel_demand.json: stamped PASS offload_ceiling_generalization.json: stamped PASS order_effect.json: stamped PASS plan_status.json: stamped PASS r1_residency_policy.json: stamped PASS r2_joint_parking.json: stamped PASS r4_gateway_signals.json: stamped PASS repo_split_status.json: stamped PASS s01_aa_null_sharded.json: stamped PASS s01_aa_null_tiny.json: stamped PASS s01_positive_control.json: stamped PASS s01_positive_control_eager.json: stamped PASS s01_variance_levels.json: stamped PASS s01_variance_levels_L1.json: stamped PASS s01_variance_levels_L2.json: stamped PASS s01_variance_levels_L3.json: stamped PASS s02_paired_knee.json: stamped PASS s02_paired_knee.prereg.json: sealed PASS s02_paired_knee.prereg.json: seal re-derives from its own recorded threshold PASS s02_sweep_sharded.json: stamped PASS s02_sweep_tiny.json: stamped PASS s02_sweep_tiny_cpu32.json: stamped PASS s06_bailian_fleet_audit.json: stamped PASS s06_bailian_fleet_audit.prereg.json: sealed PASS s06_bailian_fleet_audit.prereg.json: seal re-derives from its own recorded threshold PASS s06_c1_fixed_bridge_n8.json: stamped PASS s06_verifier_parity.json: stamped PASS s06_verifier_parity.prereg.json: sealed PASS s06_verifier_parity.prereg.json: seal re-derives from its own recorded threshold PASS s06_weka_prefix_hit_rate.json: stamped PASS s06_weka_prefix_hit_rate.prereg.json: sealed PASS s06_weka_prefix_hit_rate.prereg.json: seal re-derives from its own recorded threshold PASS s06_weka_residency.json: stamped PASS s06_weka_residency.prereg.json: sealed PASS s06_weka_residency.prereg.json: seal re-derives from its own recorded threshold PASS s08_o14_engine_rows.json: stamped PASS s08_parity_bar.json: stamped PASS s08_parity_bar.prereg.json: sealed PASS s08_parity_bar.prereg.json: seal re-derives from its own recorded threshold PASS s08_verifier_contract.json: stamped PASS s28_crossgpu_A10G.json: stamped PASS s28_crossgpu_L40S.json: stamped PASS s46_apc_real_trace.json: stamped PASS s46_apc_real_trace.prereg.json: sealed PASS s46_apc_real_trace.prereg.json: seal re-derives from its own recorded threshold PASS s46b_apc_fixed_bridge.json: stamped PASS s46b_apc_fixed_bridge.prereg.json: sealed PASS s46b_apc_fixed_bridge.prereg.json: seal re-derives from its own recorded threshold PASS s46b_engine_hit_rate.json: stamped PASS s46c_apc_fixed_bridge_n8.json: stamped FAIL s46c_apc_fixed_bridge_n8.json: does_not_prove is present but EMPTY PASS s46c_apc_fixed_bridge_n8.prereg.json: sealed PASS s46c_apc_fixed_bridge_n8.prereg.json: seal re-derives from its own recorded threshold PASS s47_aa_null_real_trace.json: stamped PASS s47_aa_null_real_trace.prereg.json: sealed PASS s47_aa_null_real_trace.prereg.json: seal re-derives from its own recorded threshold PASS s47_aa_null_real_trace.superseded.prereg.json: sealed PASS s47_aa_null_real_trace.superseded.prereg.json: seal re-derives from its own recorded threshold PASS s47b_aa_fixed_bridge_n8.json: stamped FAIL s47b_aa_fixed_bridge_n8.json: does_not_prove is present but EMPTY PASS s47b_aa_fixed_bridge_n8.prereg.json: sealed PASS s47b_aa_fixed_bridge_n8.prereg.json: seal re-derives from its own recorded threshold PASS s48_connector_smoke.json: stamped PASS s53_demo.json: stamped PASS slo_sensitivity.json: stamped PASS spend_ledger.json: stamped PASS sprint10_shadow.json: stamped PASS sprint11_refactor.json: stamped PASS sprint11_type_agnosticism.json: stamped PASS sprint12_unified_allocation.json: stamped PASS sprint13_consensus_weighted.json: stamped PASS sprint13_delta_census.json: stamped PASS sprint14_codec_frontier.json: stamped PASS sprint18_expert_census.json: stamped PASS sprint19_20_gpucount.json: stamped PASS sprint19_expert_oracle.json: stamped PASS sprint19_expert_oracle_topk.json: stamped PASS sprint20_partitioning.json: stamped PASS sprint21_expert_prefetch.json: stamped PASS sprint21b_confidence_gate.json: stamped PASS sprint22_roofline.json: stamped PASS sprint22_speculation.json: stamped PASS sprint23_joint.json: stamped PASS sprint24_customer_proof.json: stamped PASS sprint24_nvlink_recomposed.json: stamped PASS sprint27_forecasting.json: stamped PASS sprint28_censoring.json: stamped PASS sprint28_heterogeneous.json: stamped PASS sprint29_energy.json: stamped PASS sprint29_pareto.json: stamped PASS sprint32_above_the_wall.json: stamped PASS sprint32_lead_sweep.json: stamped PASS sprint32_reconstructed.json: stamped PASS sprint3_endogeneity.json: stamped PASS sprint3_simulator_calibration.json: stamped PASS sprint44_knee_binding.json: stamped PASS sprint44_knee_claim_path.json: stamped PASS sprint44_variance_attribution.json: stamped PASS sprint45_trace_bridge.json: stamped PASS sprint46_report.json: stamped PASS sprint47_report.json: stamped PASS sprint4_baseline_costs.json: stamped PASS sprint50_offload_oracle.json: stamped PASS sprint5_recompute.json: stamped PASS sprint6_oracle.json: stamped PASS sprint7_affinity_baseline.json: stamped PASS sprint7_cross_family.json: stamped PASS sprint7_cross_family_reconstructed.json: stamped PASS sprint7_liveness.json: stamped PASS sprint7_liveness_reconstructed.json: stamped PASS sprint7_real_waits.json: stamped PASS sprint8_connector_api_probe.json: stamped PASS sprint9_signal_value.json: stamped PASS sprint9_signal_value_reconstructed.json: stamped PASS x2_sibling_prefix_economics.json: stamped PASS 94 artifact(s) written after the first commit carry a REAL git_rev; none carries "unknown" [5] simulation claims re-derived from code PASS Sprint 0.0 full-trace wall 0.28906 == 0.28906 PASS Sprint 0.0 test-split wall 0.19645 == 0.19645 PASS Sprint 0.0 reconciled cells: 8/8 reproduce PASS Sprint 32 condition: Belady ATTAINS the wall at cap >=1693 PASS H005 prefetch oracle at lead 8: +69.67 pp above the wall PASS H005 oracle bandwidth 1.1325x -- not the constraint PASS Sprint 32 NEGATIVE holds: predictor 0.19621 <= wall 0.19645 PASS Sprint 32 headline 0.19621 == 0.19621 as published PASS Sprint 4 families bracket the real capture: chat 0.0000 < 0.2891 < agentic 0.9083 PASS Sprint 3 loads all 24 sweep points from the artifact (not 10 by hand) PASS Sprint 3 NEGATIVE RESULT holds: low load alone puts the knee at 49.6 qps against a measured 40-45 PASS Sprint 3 engine saturates at batch 249.3, below the configured 256 PASS Sprint 3 leakage finding stands: the retired constants imply 46.71 ms at the cap, inside the held-out plateau 46.10-49.00 ms PASS Sprint 3 contaminated constants are not back in engine.py PASS Sprint 3 sensitivity is the corrected one: d ln s/d ln a = 1, d ln s/d ln b = 1.17 (the retracted law claimed one gain for both) PASS Sprint 5 verdict holds on the real capture: R = 22.2%-26.3%, refuting ~2% and short of 30-40% PASS Sprint 5 inherits CacheCeiling.lean: distinct 38788 == compulsory misses 38788 PASS Sprint 5: at 10% capacity LRU recomputes 22.2% and Belady 0.5% -- RECOMPUTATION is an eviction failure (Sprint 6 corrected the inference: it is only 22% of the prefill bill) PASS Sprint 5 warning holds: retention alone spans the gate, 4.7% to 88.8% PASS Sprint 6: eviction never reaches gold -- best Belady ratio 0.782 > 0.55 PASS Sprint 6: prefetching IS gold at realistic lead -- 0.264 to 0.339 across 100 GbE to 10 GbE PASS Sprint 6 correction stands: recomputation is only 22.2% of the prefill bill, so eliminating all of it is a 0.782 ratio -- Sprint 5 read a component ratio as a whole-system one PASS Sprint 6: LRU is a fair baseline -- within 1.5% of the best demand policy PASS Sprint 10 refusal threshold is DERIVED: rho=0.7000 is exactly where the band reaches 20% PASS Sprint 10 interval is inverted not mirrored: honest headline 22.22% vs the mirrored 23.00% that shipped first PASS Sprint 10 never over-claims: 0 of 1328 scored reports exceeded the truth PASS Sprint 10 refusal renders as a refusal, not as a small savings figure PASS Sprint 10 materiality floor 1.0% blocks a 0.004% headline PASS Sprint 12 equivalence: the gap to unified allocation shrinks monotonically as the split grid refines, to 0.102933% PASS Sprint 12 verdict stands: 0.1029% is below the 5% floor PASS Sprint 11: 0 real leaks -- the abstraction holds PASS Sprint 12 allocator repair holds: unified/best-split 1.0094 at 5% pressure (plain greedy gave 0.498 and would have inverted the finding) PASS Sprint 11: refetching a 14.68 MB KV block beats recomputing it at every tier -- worst 11.74 ms against 85.3 ms PASS Sprint 13 consensus is reported against the random-subspace null k/d PASS Sprint 13 parameter-weighted consensus re-derives: 1.7397x chance PASS Sprint 13 drops zero-delta contributions per (model, tensor): 17 dropped -- a zero matrix has no singular subspace and would enter as a RANDOM one PASS Sprint 13 verdict stands: only 6.2% of delta parameters sit in tensors above 2x chance -- deflating for Sprint 15 PASS Sprint 13 exclusions are by measurement: 5 of 12 excluded, including 1 exact re-upload(s) PASS Sprint 15 stands dead: per-model SVD beats the shared basis at equal error on 4/4 tensors PASS Sprint 14: no structural codec gets below 0.324 error -- the deltas are near full-rank and dense PASS Sprint 14 benchmarks the prior art: BitDelta reaches 16.0x PASS Sprint 14 keeps its strawman labelled: the naive single-scale quantizer scores 1.334 at 2 bits -- worse than storing zeros PASS Sprint 18: routing is near-uniform -- top-8 share 0.266 against a uniform 0.125, only 2.1x PASS Sprint 18: routing is NOT memoryless -- H(next|prev) 3.37 against H(next) 5.36, 1.99 bits of layer-to-layer information PASS Sprint 18 duplicate-domain defect stands recorded: prose and long_context are byte-identical streams, so there are THREE domains, not four PASS Sprint 19: oracle upside 82.3% over the BEST STATIC placement clears the 20% funding threshold PASS Sprint 19: the prefetch oracle needs 70x fewer fetches than static at 50% HBM PASS Sprint 21 safety guarantee is structural: the prefetcher exposes ['fit', 'predict', 'predict_prior'] and no routing method PASS Sprint 21: the predictor works -- 31.3pp hit-rate uplift over the popularity control PASS Sprint 21 verdict stands: 0 of 54 configurations move fewer bytes than they save PASS Sprint 21 volume gap: 95 experts moved per token to cover 8, against the oracle's ~0.4 PASS Sprint 19 corrected onto the full 8-of-64 selection: ceiling holds at 78.7%, and the artifact records what it supersedes PASS Sprint 20: 31.7% less cross-device traffic than EPLB clears the 20% gate PASS Sprint 20 reports the COST beside the gain: worst per-layer imbalance 1.86 -> 2.33, a trade not a free win PASS Sprint 20 memoryless control holds: 5.5% on random routing (it scored 60.5% before per-layer balance was enforced) PASS Sprint 20 tolerance floors at the baseline's own imbalance -- a tighter one produced a flat 0.00% that was a result about the constraint PASS Sprint 22: acceptance 0.76844 is 99.92% of the machine-checked ceiling 0.76903 -- it cannot be the variable, which is why no GPU sweep ran PASS Sprint 22: speculation is a slowdown at every depth at the measured operating batch of 249 (0.854 to 0.409) PASS Sprint 22: the reversal survives a FREE drafter -- depth 8 crosses at batch 51, so it is physics and not a draft-cost artefact PASS Sprint 23: the do-nothing baseline is already infeasible on PCIe -- Sprint 19 measured EPLB placement at 25.3 GB/s against a 25 GB/s link PASS Sprint 23: joint beats best-single by 33.2% PASS Sprint 23: joint EQUALS naive stacking in compute -- the controller is not earned, ship all the wedges PASS Sprint 23: speculation is never selected in any feasible configuration PASS Sprint 23: 3 pairs are synergistic -- worthless alone, valuable together, because the constraint binds PASS Sprint 27 KILL stands: predictive +1.50% against a 15% gate PASS Sprint 27 reachability gap: oracle reaches 23.2% where the forecaster cannot PASS Sprint 28 best heterogeneous gain 3.59% against a 20% gate PASS Sprint 28 reports UNDERPOWERED, not KILL -- reps=1 cannot decide a 20% gate on a rig whose knee CV is 15.73% PASS Sprint 28 stamps pricing_verified: false -- an unverified rate cannot silently become a claim PASS Sprint 28: prefill $/perf is flat across three GPU classes -- the arbitrage is already priced out PASS Sprint 29: energy vs GPU-seconds r=0.99948 -- one objective PASS Sprint 29: J/token swings 10.1x with load -- utilisation is the lever, not policy PASS Sprint 24 capstone stands: composed ACHIEVABLE GPU reduction 0.00% -- the gate's <10% branch PASS Sprint 24 declines the GPU run rather than staging one against a configuration assembled from ceilings PASS Sprint 24: all 10 wedges name their sprint and gap kind PASS Sprint 24 points at the VOLUME gap -- the one kind that is an engineering problem rather than a dead end PASS Sprint 21b: the confidence gate cuts prefetch volume 12.0x (94.6 -> 7.9 per token) PASS Sprint 21b: gated net traffic is 1.011x of doing nothing -- essentially free, against 1.70x ungated PASS Sprint 21b reports the COST: uplift +31.3pp -> +5.5pp, 82% of it spent on the volume cut PASS Sprint 21b records that closing the volume gap is necessary and NOT sufficient -- static placement alone already exceeds PCIe PASS SPRINTS.md carries the execution-coverage warning above its state of play PASS execution audit: 5 of 30 MEASURED sections ran their specified experiment (21 partial, 4 substituted or not run) PASS execution audit still grades Sprint 8 as C -- the specified experiment -- 'Real GPU. Demonstrate an HBM -> host -> HBM KV life PASS execution audit still grades Sprint 14 as C -- the gate metric was substituted. The artifact says so in its own field: gate_not PASS execution audit still grades Sprint 22 as C -- the specified experiment -- 'sweep batch x workload {code,chat,JSON,agentic} x d PASS execution audit still grades Sprint 24 as C -- the specified experiment -- 'Real rented GPUs. Baseline = latest tuned vLLM/SGLa PASS execution audit carries the CORRECTED grades 5A/21B/4C, not the original 12A/14B/4C (got 5A/21B/4C) PASS 9 grade revisions are recorded with their evidence, including Sprint 11 C->A PASS Sprint 11's refactor IS done -- predictors/ imports InferenceObject (predictors/liveness_objects.py) PASS Sprint 11's port is byte-identical on every cell it ran (16 cells, 0 mismatches) PASS Sprint 11 is graded B because its gate was SUBSTITUTED: it ran capacities [157, 314, 629, 1258] against a fresh in-process run, while the gate names results/sprint7_liveness.json -- no overlap, nothing cross-checkable PASS Sprint 11 is graded B, the same standard applied to Sprint 14 -- a substituted gate is a substituted gate regardless of how good the substitute is [6] certificate interrogation PASS guards.fail_closed.require_artifact gates this section's evidence -- missing OR EMPTY now raises instead of skipping PASS every certificate is interrogated: 129/129 named by a checker PASS s01_aa_null_sharded.json: analysis re-derives from its own raw replicates PASS s01_aa_null_tiny.json: analysis re-derives from its own raw replicates PASS s01_positive_control.json: analysis re-derives from its own raw replicates PASS s01_positive_control_eager.json: analysis re-derives from its own raw replicates PASS s01_variance_levels.json: analysis re-derives from its own raw replicates PASS s01_variance_levels_L1.json: analysis re-derives from its own raw replicates PASS s01_variance_levels_L2.json: analysis re-derives from its own raw replicates PASS s01_variance_levels_L3.json: analysis re-derives from its own raw replicates PASS the known arm-order mismatch is still exactly that and nothing more: cert ['graphs', 'eager'] vs raw ['eager', 'graphs'], ratios exact reciprocals (0.477232 and 2.095417) -- a ratio is not comparable without its arm order PASS sprint11_type_agnosticism spans 5 object types -- the abstraction claims five and the original artifact demonstrated one PASS sprint11_type_agnosticism actually PERMUTED the type field (720 labels moved) -- a permutation test that permutes nothing proves nothing PASS sprint11_type_agnosticism: admission_identical under a type permutation PASS sprint11_type_agnosticism: eviction_trace_identical under a type permutation PASS sprint11_type_agnosticism: final_resident_identical under a type permutation PASS sprint11_type_agnosticism spans 1700.9x in object size -- a narrow spread is where a type leak would hide PASS sprint7_affinity_baseline sweeps 4 capacities, matching the original artifact's PASS sprint7_affinity_baseline measures BOTH baselines at every capacity PASS sprint7_affinity_baseline records that the predictive EVICTION arm loses to plain LRU at every capacity -- the original artifact reports only pct_of_belady_gap and pct_of_prefetch_oracle_gap, in which this is invisible PASS sprint7_affinity_baseline does not claim a clean sweep against the affinity baseline PASS sprint7_affinity_baseline records that the winning arm is ORACLE-ASSISTED, so the comparison flatters us PASS s53_demo runs 6 identical DIRECT calls as a control -- two can only say 'differ or not' and cannot separate a cold-start boundary from run-to-run nondeterminism PASS s53_demo records WHY the control exists -- its first version compared one gateway call to one direct call and concluded THE PROXY CHANGED THE OUTPUT PASS s53_demo: the gateway is self-consistent across two calls PASS s53_demo: the gateway's output matches a WARM direct call -- the proxy alters nothing, which is Track B's correctness gate PASS s53_demo: the engine IS deterministic once warm, so the cold-start difference is a boundary and not chaos PASS s53_demo: first-differs-rest-agree, the signature of a cold-start boundary PASS s53_demo: the engine is NOT byte-identical across its own cold start -- the finding that makes the target claim's 'byte-identical outputs' conditional PASS sprint47_report quotes s47_aa_null_real_trace's ratio unchanged PASS sprint47_report quotes the source's pooled knee CV unchanged PASS s47_aa_null_real_trace's arms are a1/a2, matching the declared ratio_orientation its seal now requires PASS sprint47_report: the A/A is NOT significant (p=0.646293) -- an A/A that finds an effect means the rig manufactures one PASS sprint47_report: the A/A confidence interval straddles 1.0 ([0.933486, 1.049156]) PASS sprint47_report: Sprint 2's gate PASSES on the A/A (2.42% vs 10.0%) -- 19.6's condition (b) may be satisfiable as written PASS sprint47_report records that the POOLED knee CV puts the between-arm effect inside a reproducibility number PASS sprint47_report: S46's arms each reproduce INSIDE the gate while its pooled CV fails -- the demonstration that the pooled statistic inverts the gate PASS sprint47_report: every arm across both experiments passes within-arm PASS sprint47_report states whether 20.9's PROPOSED R1 passes -- reporting only that the old gate passed would be selecting the flattering criterion PASS sprint47_report records that 1 of 10 arm-knees was CLIENT_LAG-bound, so by the harness's own criterion this run's ratio is not quotable PASS s48_connector_smoke: the connector loaded PASS s48_connector_smoke: the evidence is conclusive (log present AND module importable) -- an empty table means 'not measured', not 'never called' PASS s48_connector_smoke recorded a non-empty call table PASS s48_connector_smoke: get_num_new_matched_tokens fired 42 times for 40 requests -- at least once per request is what 'driven per request' means PASS s48_connector_smoke: worker-side methods fire too, so the seam is exercised on both the scheduler and worker paths rather than only where we looked PASS s48_connector_smoke: no KV was moved -- the connector is the identity, so this is a statement about the INTERFACE and not about any policy PASS sprint4_baseline_costs anchors on the apc_off arm -- the BASELINE. Anchoring on apc_on would price the trace at the CACHED capacity and understate every baseline cost in the table PASS sprint4_baseline_costs has at least one MEASURED knee to anchor to -- a cost with no measured capacity anywhere is arithmetic on an assumption PASS sprint4_baseline_costs reports how many families are SCALED rather than measured -- collapsing the two would present four assumptions as four results PASS sprint4_baseline_costs records that its own scaling assumption is refuted by S46 on the one trace where it can be checked PASS estate_verify_all measured 26 Lean corpora -- a parse that finds none would report GREEN on an empty set, which it did on its first run PASS estate_verify_all: 62 targets over 26 corpora -- targets must outnumber corpora, several mechanized-* rules share a directory PASS estate_verify_all: pass + not-pass reconciles against the corpus count PASS estate_verify_all records that every corpus was CACHED -- a cached build verifies the cache is current, and verify-all's own header claims to be a clean-checkout gate, which is a different and stronger thing PASS estate_verify_all records that it covers only the mechanized-* Lean corpora -- a green Lean spine does NOT mean `make verify-all` exits 0 PASS sprint46_report records both arm names PASS sprint46_report's orientation matches the source artifact's arm_names (['apc_off', 'apc_on'] vs ['apc_off', 'apc_on']) PASS sprint46_report quotes the source's ratio unchanged PASS sprint46_report: reciprocal arithmetic is consistent (0.766419 * 1.304769 = 1.000000) -- an orientation error lives exactly here PASS sprint46_report: reciprocating the CI also REVERSES it -- 1/[lo,hi] is [1/hi,1/lo], and getting that backwards would widen or invert the interval PASS sprint46_report: the uplift is positive (30.477%) -- APC on a trace with reuse can only help, and a null would mean a broken bridge rather than a working cache PASS sprint46_report: 0 arm-knees harness-bound -- S44p found 6 of 9 historically, and a harness-bound knee makes the ratio a comparison of load-generator events PASS sprint46_report: the ratio is quotable by the S44p criterion PASS sprint46_report RECORDS that the sealed threshold's numeric form contradicts its own words -- quietly quoting the reciprocal and leaving 0.766 in the artifact for a reader to find later is the failure this check prevents PASS sprint46_report: SLO budget used at the knee is 0.55-0.81, above the 0.25 that SPRINTS.md 20.9 R3 would require to call it an SLO knee rather than a stationarity limit PASS order_effect EXCLUDES the real A/B runs -- on an A/B, order confounds with treatment and pooling them would manufacture or mask an effect PASS order_effect: both arms actually appear first ({'a1': 20, 'a2': 10}) -- a fixed order makes the test impossible, so this is a precondition not a statistic PASS order_effect pools 30 A/A replicates PASS order_effect: 0 metrics show a significant first-vs-second difference -- on an A/A that could only be state leaking between arms, and it would contaminate every paired ratio in this repo PASS order_effect states the effect size it EXCLUDES (2.745% on ttft_p95) rather than reporting a bare null -- a null without a bound is not a result PASS sprint29_pareto uses all 24 of Sprint 29's telemetry rows PASS sprint29_pareto sweeps the objective weight over many orders of magnitude -- a single weight cannot show invariance PASS sprint29_pareto: the argmin of cost + w*joules does NOT move with w, which is the ARGMIN test Sprint 29's Spearman 0.978 could not perform PASS sprint29_pareto records that the cost-optimal load sits ABOVE the measured knee -- a frontier over cost and energy alone selects an operating point the system cannot hold, and reporting it without that would be a recommendation to run past saturation PASS sprint29_pareto reports the best SLO-FEASIBLE point beside the unconstrained one, so the reachable answer is on record PASS sprint6_oracle carries its capacity sweeps PASS sprint6_oracle: Mooncake working set pinned at 38,788 blocks (got 38788) PASS sprint6_oracle: 54,559 block references on the real trace (got 54559) PASS sprint6_oracle sweeps 7 capacity fractions -- a single capacity cannot establish a ceiling's shape PASS sprint32_above_the_wall: the demand wall is pinned two-sided at 0.19645 (got 0.19645) -- this is the ORPHANED original, kept as evidence the gap existed, and it must not drift PASS sprint32_above_the_wall: held-out split pinned at 21,069 refs / 16,930 distinct PASS sprint32_above_the_wall carries all 18 published cells (18) PASS sprint12_unified_allocation's single-type CONTROL passes -- without it the 0.0004% result cannot be distinguished from a broken allocator PASS sprint12_unified_allocation sweeps 6 split-grid resolutions -- the residual is grid quantisation, and one resolution could not show that PASS sprint12_unified_allocation: unified beats the best static split by <0.05% at EVERY grid resolution -- the thesis-killing result, pinned PASS sprint12_unified_allocation: 0 real type leaks in the Sprint 11 ledger PASS sprint22_speculation: measured acceptance pinned at 0.76844 (got 0.76844) PASS sprint22_speculation: measured acceptance 0.76844 does not exceed its proven alpha=1-TV ceiling 0.76903 -- a measurement above a proven bound is an instrument error, not a result PASS sprint10_shadow refuses above rho=0.70 -- the refusal threshold is the finding (real fleets run above it), so it is pinned PASS sprint10_shadow: all gates pass PASS sprint10_shadow runs 4000 trials per rho PASS sprint44_variance_attribution has all three variance levels present (3 present, 0 missing) -- an attribution over a partial ladder attributes to whatever happened to run PASS sprint44_variance_attribution carries its per-level rows PASS sprint45_trace_bridge's gate actually ran PASS sprint45_trace_bridge: 20,000 prefix pairs checked, 0 violations (got 20000/0) -- the bridge result S46 was authorised on PASS sprint45_trace_bridge's inventory contains at least one REAL trace with non-zero reuse -- the whole point of the sprint PASS sprint5_recompute sweeps 9 capacities PASS sprint5_recompute uses the trace's declared block_tokens=512, not a downscaled default PASS lossless_throughput sweeps 18 codec configurations PASS lossless_throughput pins the ratio 19.1 needs at 2.55 PASS lossless_throughput: NO configuration meets both the 2.55x ratio and PCIe4's 25 GB/s -- 19.1's decision rests on this being empty PASS kv_footprint_ladder priced every available model (14/14), refusing 0 -- a ladder with silent gaps is not a ladder PASS mooncake_reconciliation uses the trace's declared block_tokens PASS mooncake_reconciliation holds out a test split PASS sprint7_liveness_reconstructed: wall pinned two-sided at 0.90834 (got 0.90834) -- these are the C1 reconstructions, and a reconstruction that drifts silently is worse than the orphan it replaced PASS sprint7_liveness_reconstructed: n_refs pinned two-sided at 68666 (got 68666) -- these are the C1 reconstructions, and a reconstruction that drifts silently is worse than the orphan it replaced PASS sprint9_signal_value_reconstructed: capacity pinned two-sided at 629 (got 629) -- these are the C1 reconstructions, and a reconstruction that drifts silently is worse than the orphan it replaced PASS sprint9_signal_value_reconstructed: n_refs pinned two-sided at 22045 (got 22045) -- these are the C1 reconstructions, and a reconstruction that drifts silently is worse than the orphan it replaced PASS sprint7_cross_family_reconstructed carries 4 families PASS sprint7_cross_family_reconstructed records the RECOVERED lead=8 -- the published artifacts record no lead at all, and every pct_of_prefetch_oracle_gap in both Sprint 7 halves was computed against it PASS spend_ledger prices 24 GPU artifacts from recorded wall_s and the GPU each run actually GOT -- a spend total over an empty set would be zero, and zero would look like thrift PASS spend_ledger declares itself a LOWER BOUND: wall_s excludes container boot and model load, and runs that crashed before writing an artifact cost real money and appear nowhere PASS spend_ledger stamps pricing_verified from the rate table rather than asserting it -- pricing/2026-08.yaml carries inherited placeholders whose own header records five inconsistent estate copies differing by up to 50% PASS spend_ledger lists the 3 artifact(s) that show a GPU and could not be priced rather than dropping them -- a silent drop pushes a spend total down and looks like frugality PASS spend_ledger records that the ~$85.0 total reported throughout this session has no derivation, against a derived floor of $11.4106 (7.45x) -- and that only Modal's billing settles it PASS spend_ledger's total is the sum of its derived and declared parts, so the headline cannot drift from the table beside it PASS offload_ceiling_generalization compares 3 block streams including a granularity-matched CONTROL arm -- this program has twice produced a dramatic negative from a comparison with no control, both false PASS offload_ceiling_generalization: block granularity explains 0.5% of the between-trace gap -- finer blocks raise a hit rate mechanically, so without this the difference could be an artefact rather than a trace effect PASS offload_ceiling_generalization: the offload ceiling IS trace-dependent (headroom ratio 0.0494) -- Phase 3's premise that one trace is one data point is now measured, not asserted PASS offload_ceiling_generalization: S50's 12pp gate is cleared on one trace and not the other ({'mooncake_timed_600': True, 'wildchat_coarsened_to_64': False}) -- with an UNBOUNDED, FREE host tier. A sprint whose result is decided by trace choice is not yet a result about connectors PASS offload_ceiling_generalization reports each trace's COLD-MISS FLOOR beside its headroom -- no connector can move a first-touch, so a headroom quoted without its floor is uninterpretable across traces PASS sprint22_roofline: the step law's FIXED cost is weight streaming -- 10.2667 ms predicted from 3.08 GB over the device datasheet vs 12.41031 ms fitted (83%), with no free parameters PASS sprint22_roofline: the MARGINAL per-sequence cost is NOT memory traffic -- KV reads explain only 22% of it (4.536x gap). Sprint 22 scales b by tokens verified; that assumes the non-bandwidth majority scales the same way, and nothing establishes it PASS sprint22_roofline records the batch at which KV traffic equals weight traffic (B=331.5) -- below it decode is dominated by re-reading weights, which is the regime every single-stream latency claim in this program was measured in PASS sprint22_roofline offers the match to Sprint 3's underived a=12.078 (at 0.85 of peak bandwidth) as an OBSERVATION with its implied fraction stated -- the assumed fraction was chosen before the comparison, and a reader must be able to judge the coincidence PASS active_expert_ceiling: the 1.648x active-expert bandwidth ceiling is pinned two-sided (got 1.6484) -- 19.4 makes it a PRE-REGISTERED upper bound on Sprint 36, and a pre-registered bound that can drift is not one PASS active_expert_ceiling decomposes at least one real model config rather than quoting the number from prose -- 19.3's table had no script until C26 PASS active_expert_ceiling states that it bounds BANDWIDTH and not storage -- 19.3's other half is that storage upside is large, and conflating them is how a 1.648x cap gets read as a cap on the whole wedge PASS plan_status classifies 98 plan items PASS plan_status: ungated done + ungated undone = ungated total, so the headline cannot drift from the item list beside it PASS multimodel_demand establishes C29's true half by INSPECTION -- Mooncake's requests carry ['hash_ids', 'input_length', 'output_length', 'timestamp'] and no model field, so it cannot be partitioned by model even after the fact PASS multimodel_demand records that a labelled multi-model trace DOES exist -- C29's conclusion that none was available did not follow from Mooncake's limitation, and was false PASS multimodel_demand carries the trace's OWN classification (is_real_trace=False, derived_from_real_text=True) -- promoting it to REAL or demoting it to SELF-AUTHORED would both be wrong in ways that matter PASS bench/trace_io classifies the labelled trace in its own DERIVED_FROM_REAL class rather than folding it into REAL_TRACES PASS multimodel_demand: real multi-model demand is MORE concentrated than the self-authored trace (HHI 0.538019 vs 0.093309, 5.766x) -- the Zipf-over-60 premise behind a $2,000 sprint was invented PASS multimodel_demand states that two hosted API models are not Sprint 17's scenario -- the datum widens the premise, it does not cancel the sprint PASS bench/trace_schema sees 7 trace files -- a declaration check over an empty set is the vacuity this repo has caught before PASS every trace declares whether it was RECORDED or GENERATED (7 files) -- 20.17 found the estate's fleet anchor carrying `synthetic: true` unread while three GPU runs were scheduled against it as real PASS every trace marked as a SLICE carries its source digest (none missing) -- D4 exists because a 600-request slice was being cited as a 12,031-request trace PASS RESEARCH_CONSTITUTION.md exists -- 17 names it, and every sprint is graded against the clause it holds PASS 2 is still locatable in SPRINTS.md (30 lines) -- a parser failure reported as agreement would make the drift check vacuous PASS RESEARCH_CONSTITUTION.md still matches SPRINTS.md 2 verbatim (0 line(s) drifted) -- a stale constitution means sprints are graded against text the plan no longer says PASS estate_imports records what was imported; an empty import list reported as done is the vacuity this repo has caught before PASS estate_imports hash-pins all 6 imported assets with a full sha256 -- an import without a digest is a copy, not provenance PASS estate_imports records the estate HEAD at import (515540069e82) -- the estate is a moving submodule, so an import without a sha is pinned to nothing PASS every asset estate_imports claims is present actually exists on disk PASS estate_imports names the rows it deliberately did NOT import, with reasons -- a partial import reported as complete is the same defect as a green check over an unstated subset PASS estate_imports states that holding a certificate is not confirming it -- agentic_fleet_gpu's 1.142x stays unconfirmed until 18 Run A runs PASS fleet_receipt_provenance re-derives every quoted number of agentic_fleet_gpu.json from its ARMS rather than its stored gains block ({'goodput_reqs_s_ratio': 1.142378, 'goodput_frac_ratio': 1.165468, 'ttft_p99_ratio_theirs_over_ours': 0.839293, 'vs_round_robin_ratio': 1.950924}) PASS fleet_receipt_provenance independently picks the strongest cache-aware arm (cache_aware@p1) and it matches the declared one (cache_aware@p1) -- so 1.1424x is against the best competitor, not a strawman PASS fleet_receipt_provenance: NO gate substitution -- the claim declares goodput@SLO and the headline IS that measure (declared=True, headline_matches=True) PASS the fleet receipt's workload is SYNTHETIC (calibrated to mooncake_fast25_conversation_trace) and this repo's prose says so -- 18 schedules three GPU runs against it, and S45 refused a trace for less PASS repo_split_status enumerates all 11 rows of 17's import table, not only the ones that are done PASS repo_split_status: the done count (8) is derived from the per-import check, so it cannot drift from the table beside it PASS repo_split_status: the cited-but-absent count (3) is the length of the named list PASS repo_split_status: cited_but_absent is exactly (cited AND not held) for every row -- the flag is computed, not asserted PASS repo_split_status reports the Lean imports as held, and .lean files really are in the tree -- a citation this repo cannot open is a claim a reader cannot check PASS the imported Lean files carry provenance: source path, estate HEAD sha, and a plain statement that this repo does NOT build them -- copying a proof is not checking it PASS repo_split_status separates 'not imported' from 'no longer importable' ([]) -- an import whose source is gone is a different problem from one nobody has done yet PASS estate_negatives_and_pins re-reads the register's own counts (52 preserved, 39 not_probed, 8 with no detector) from the view rather than quoting the plan PASS estate_negatives_and_pins records that the negatives register CANNOT be regenerated -- generator present=False, registry/ present=False PASS estate_negatives_and_pins reconciles the register's own MANIFEST.sha256 against the tree -- an unquantified 'the inputs are missing' is an impression, not a measurement PASS estate_negatives_and_pins: the missing-input count (24) is the length of the named list, not a separately typed number PASS estate_negatives_and_pins: every manifest input still present also MATCHES its digest (9 of 33) -- the defect is absence, not corruption, and those are different findings PASS estate_negatives_and_pins enumerates every tool under simulators/external rather than only the ones the plan names PASS estate_negatives_and_pins: the pinned count (1) is derived from the per-tool classification, so it cannot drift from the table beside it PASS estate_negatives_and_pins names the tools present on disk with NO vendoring recipe (['homa_src', 'hpcc_src', 'alloy.jar']) -- unpinned and unfetchable are different failures and it separates them PASS estate_negatives_and_pins states that a recorded checkout SHA is EVIDENCE of what was used and not a pin -- recording it must not be mistaken for fixing it PASS estate_publish_force_risk reached 20 live remotes -- a risk report built from the local tree alone would be a guess about what is on GitHub PASS estate_publish_force_risk pushed nothing: GET requests only, comparing content by git blob SHA-1 rather than transferring files PASS estate_publish_force_risk measured 19 packages at risk against the plan's 3 named ones -- the recalled list was an undercount, and the artifact says by how much PASS estate_publish_force_risk reports the recall error in BOTH directions -- named-but-safe and at-risk-but-unnamed -- rather than only the direction that flatters the plan PASS estate_publish_force_risk names the specific paths a force-push would delete, not merely a count PASS estate_publish_force_risk flags deletions under .github/workflows/ separately (['certhead', 'isolation-tax']) -- removing a workflow turns CI OFF on a published repo rather than failing it, which is the one deletion that hides its own consequence PASS estate_published_datasets measured 4 published datasets against their own manifests -- an empty measurement reported as clean is the vacuity this repo has caught before PASS estate_published_datasets: every published dataset matches its declared sha256 and row count -- so the plan's F8 mechanism ("sha256 mismatch, 280 rows") does NOT reproduce, and the artifact says so PASS estate_published_datasets records that the plan's stated F8 mechanism failed to reproduce, rather than quietly re-describing a different defect under the same label PASS estate_published_datasets counts columns shipped but NOT declared in the manifest (9), derived from the rows rather than from the exporter's hand-written table PASS estate_published_datasets names the published datasets the estate's own --check does not cover (['abstain-corpus']) -- a green check over an unstated subset is not a green check PASS estate_published_datasets reconciles coverage against what is LIVE: 3 covered + 1 uncovered = 4 published PASS estate_published_datasets modified nothing published -- rewriting a manifest or a card changes public content under an irrevocable grant, which is the owner's action and not a measurement's PASS estate_orphan_theorems reports BOTH the orphan count (480) and the count that would mean UNVERIFIED (0) -- quoting the first alone is what put '480 unaudited theorems' in our own plan PASS estate_orphan_theorems carries the estate tool's own limit: reachability is TEXTUAL, so the orphan count is a lower bound -- a number quoted without its direction of error is not a measurement PASS estate_orphan_theorems: the reassuring reading is the CONJUNCTION of full corpus coverage and a GREEN exhaustive audit -- scope alone does not license it PASS estate_orphan_theorems records that audit_axioms_exhaustive.py --all --strict was actually RUN, with its exit status -- not that it exists PASS estate_orphan_theorems: the exhaustive audit obtained an axiom footprint for every theorem it parsed (1715 of 1715) -- parsed-but-unresolved is the gap that makes a green audit vacuous PASS estate_orphan_theorems counts theorems in corpora that are FLAGGED in scope for the exhaustive audit and skipped at run time for want of a lakefile PASS estate_orphan_theorems: 14 theorem(s) are in scope and unaudited, and the verdict says so rather than reporting the corpus coverage as clean PASS sprint50_offload_oracle: cold-miss floor 0.831282 + max achievable hit rate 0.168718 = 1 -- the ceiling is a DECOMPOSITION of the reference stream, not a fitted quantity PASS sprint50_offload_oracle sweeps 4 capacities -- a single capacity could not show whether the headroom is a property of the trace or of one configuration PASS sprint50_offload_oracle: the unbounded host tier is at least as good as GPU-only at EVERY capacity -- an oracle that loses to the thing it bounds would be an implementation error, not a ceiling PASS sprint50_offload_oracle is computed at the capacity S46 ACTUALLY RAN (num_gpu_blocks_override=4096, and the S46 certificate records 4096) -- a ceiling at a different capacity bounds a different system PASS sprint50_offload_oracle records the tension between its MODELLED 3.9% hit rate and S46's MEASURED 30.5% knee uplift, rather than reporting the headroom as if the two were consistent PASS sprint50_offload_oracle: s50_authorised_to_spend is the CONJUNCTION of clearing the 12pp gate and having measured vLLM's own prefix-cache counters -- not the gate alone PASS sprint3_endogeneity sweeps 6 noise levels -- a single sigma could not distinguish a structural bias from a noise artefact PASS sprint3_endogeneity: the reconstructed regressor reports a SMALLER held-out error at every noise level -- the bias flatters, which is the finding; classical errors-in-variables would go the other way PASS sprint3_endogeneity: the optimism factor is flat ([1.0787, 1.0927]) across a 10x noise range, so it is structural and more replicates will not remove it PASS sprint3_endogeneity: r_squared inflation GROWS with noise even though the held-out optimism does not -- that separation is why r_squared is the sensitive tell PASS sprint3_endogeneity: the effect is MODEST at the repo's measured noise -- recorded as sharpening Sprint 3's own self-assessment, not overturning it PASS slo_sensitivity sweeps 8 SLO multipliers -- one threshold is what Sprint 2's E3 was owed for PASS slo_sensitivity covers 21 arm-sweeps across 6 raw artifacts -- both on-disk shapes must be read or half the evidence is silently skipped PASS slo_sensitivity reproduces a knee at 1.0x nominal -- the control PASS slo_sensitivity: the knee is UNCHANGED at 2x and 5x the nominal SLO -- if it moved, the knee would be SLO-bound and S44p's finding would not generalise PASS slo_sensitivity: at the nominal SLO, most arms are bound by STATIONARITY, not the SLO ({'STATIONARITY': 17, 'BOTH': 3, 'NONE_FEASIBLE': 1}) PASS sprint19_20_gpucount reports a phi threshold per load PASS sprint19_20_gpucount: at least one reference load saves NO whole GPU even at phi=1.0 -- integer quantisation is the finding, and an analysis where every load pays would have lost it PASS sprint19_20_gpucount: saving one GPU needs phi >= 0.6301 -- a bar Sprint 49 must clear, not merely a non-zero derivative PASS sprint19_20_gpucount names Sprint 49 as the blocking measurement rather than substituting an assumed phi PASS sprint28_censoring: the A10G sweep DID saturate -- without a saturated arm there is no contrast and the diagnostic proves nothing PASS sprint28_censoring: the L40S sweep did NOT saturate -- its quoted throughput is the offered load, not a capacity PASS sprint28_censoring: the unsaturated GPU's delivery ratio holds flatter than the saturated one's -- that ordering IS the diagnostic PASS sprint28_censoring: L40S needs only 1.4444x its ladder top to reverse the ordering -- a headline this close to reversal is not established PASS estate_acceptance_triples found 22 distinct acceptance triples -- the plan hand-counted 3, and a sweep that agreed with the hand-count would mean the sweep was not looking PASS estate_acceptance_triples separates 63 correctly-labelled superseded quotations from live contradictions -- without that split it would report README.md as self-contradicting, which it is not PASS estate_acceptance_triples: stale-current is a subset of asserted-current PASS estate_axiom_coverage deduplicates by lean_dir -- three CORPORA entries name theory/lean and summing per-entry counts triple-counts its 454 declarations PASS estate_axiom_coverage: 606 declarations named of 1642 found -- the banner says ALL, and the coverage is the measurement of that quantifier PASS estate_axiom_coverage 36.9% corroborates the plan's independent D-4 figure of 32.4% by a different parse PASS estate_release_readiness counted the publish gate's boxes from its own file -- four different counts circulate in prose (22, 18, 17, 15) PASS estate_release_readiness: publish-gate boxes reconcile (ticked + unticked == total) PASS estate_release_readiness: 17 stamped + 127 unstamped reconcile against 144 certs PASS estate_release_readiness measured 17 stamped GPU certificates against the plan's hand-count of 6 -- this one runs the estate's way and must not be quietly kept at the worse number PASS sprint32_lead_sweep sweeps 7 leads -- the amendment calls lead time the binding variable and DEPTHS=(1,2,4) is not a sweep of it PASS sprint32_lead_sweep: the predictor hit rate is CONSTANT across every lead ([0.19621]) -- chaining a first-order table forward buys no hits PASS sprint32_lead_sweep: prefetch volume RISES monotonically with lead (2580 -> 2947, +14%) while hits do not -- the cost of lead is real and its benefit is zero PASS sprint32_lead_sweep FAILS the amended gate at every lead; if this ever passes it REOPENS the KV wedge and must be reported, not absorbed PASS sprint32_lead_sweep best share of oracle excess is ~0 (-0.0003) PASS s01_variance_levels.json declares reps_usable (5) -- a run that reports no count cannot be distinguished from a run that never happened PASS s01_variance_levels.json reports a usable run and carries a paired s_y -- the quantity the whole variance decomposition exists to produce PASS s01_variance_levels_L1.json declares reps_usable (5) -- a run that reports no count cannot be distinguished from a run that never happened PASS s01_variance_levels_L1.json reports a usable run and carries a paired s_y -- the quantity the whole variance decomposition exists to produce PASS s01_variance_levels_L2.json declares reps_usable (5) -- a run that reports no count cannot be distinguished from a run that never happened PASS s01_variance_levels_L2.json reports a usable run and carries a paired s_y -- the quantity the whole variance decomposition exists to produce PASS s01_variance_levels_L3.json declares reps_usable (5) -- a run that reports no count cannot be distinguished from a run that never happened PASS s01_variance_levels_L3.json reports a usable run and carries a paired s_y -- the quantity the whole variance decomposition exists to produce PASS s02_paired_knee.json: carries its own per-rep evidence PASS s28_crossgpu_A10G.json: carries its own per-rep evidence PASS s28_crossgpu_L40S.json: carries its own per-rep evidence PASS s01_aa_null_tiny.json: carries its own per-rep evidence PASS s01_aa_null_sharded.json: carries its own per-rep evidence PASS s01_positive_control.json: carries its own per-rep evidence PASS re-derived the knee summary of 3 sweep certificates from per-rep data PASS lossless_probe: headline 1.4973 RE-DERIVES as 1.4973 from its own per-tensor codec table PASS lossless_probe: 1.4973 against the 2.548 PCIe needs -- this is the measurement 19.1 turns on, and it says the constitution conflict is REAL PASS sprint8 probe validated vLLM 0.28.0 -- it must match the version the GPU runs used (0.28.0), which an earlier run of this probe did not PASS sprint8: all 9 KVConnectorBase_V1 methods the estate overrides are present in vLLM 0.28 -- no fork required PASS sprint9 carries 7 signal-quality levels including a no-signal control -- the arm that makes its convexity claim readable PASS sprint3 still records its gate as NOT CLOSED -- the calibration never closed and that must not quietly become a pass PASS sprint7_cross_family carries its per-family rows PASS re-derived 32 stored ratios across 8 paired-arm GPU certificates from per-replicate raw data (only s01_* carry paired arms; the sweeps are covered above) PASS all 16 headline numbers are pinned TWO-SIDED to their published values -- a wrong number now fails, not just a deleted one PASS fleet_report_091494c4 reproduces the offload oracle's cold-miss floor 0.831282 on the same stream (got 0.831282) PASS fleet_report_091494c4 reproduces the offload oracle's headroom 13.0039 pp (got 13.0039) PASS fleet_report_091494c4 REFUSES the block-hit -> GPU-hour conversion (C11 is unmeasured; a number here would be invented) PASS fleet_report_091494c4: host-backed hit rate equals 1 - cold-miss floor -- the ceiling is a decomposition, not a fit PASS fleet_report_validation: the floor-only predictor is KILLED as pre-registered (worst LOO rel err 100.6595 on wildchat_labelled@16); the report ships the curve PASS fleet_report_validation spans 9 traces with reuse -- a generalization claim needs more than the two it started from PASS d6_isolation_tax: the unsalted control LEAKED (1372 cross-tenant hits), so the salted arm's zero is readable PASS d6_isolation_tax: the salted arm leaked nothing PASS d6_isolation_tax: tax 0.1436 pp < 2.0 and the verdict says KILLED -- the number and the word agree PASS d7_admission_above_knee: the hard cap never loses to open loop at >= 1.25x the knee PASS d7_admission_above_knee: every past-knee hard-cap row carries a nonzero shed count -- the ledger is present beside the latency it bought PASS d7_admission_above_knee: the controller AS SHIPPED sheds > 90% at every load (min 0.965333) -- the defect that turned the Little's bound off by default is reproduced, not remembered PASS d5_knee_binding_by_shards: client-lag-bound fraction 1.0 at 1 shard -> 0.0455 at 8 -- sharding removes the harness-bound knees (2605.24217's result on our data, not ours) PASS d5_knee_binding_by_shards: at 8 shards 20 of 22 knees are stationarity-bound -- what remains after the harness artefact is removed is still not the SLO PASS d5 re-reads every synthetic arm-knee sprint44_knee_binding has (10 of 10) -- no row was dropped to make the table cleaner PASS d1 attempt 2: hit rate reads exactly 1.0 on all 8 APC-OFF points -- the timestamp-as-counter defect is reproduced from the artifact, not remembered PASS d1 attempt 2 recorded the raw `_created` metric lines that explain the defect PASS d1 attempt 2's certificate carries its knees (apc_off 7.0617, apc_on 9.976062) -- the absolute knee moved 15% from attempt 1 in a fresh container, the paired ratio less, as Sprint 2 measured PASS d1: the APC-off control reports zero hits -- the counters read what they claim PASS d1: engine hit rate 0.386463 EXCEEDS the trace's theoretical maximum 0.168718 -- the number cannot be a property of the trace PASS d1: the legacy bridge mapped 13712 block ids to 33 texts -- the collapse is re-derived, not asserted PASS d1: engine rate 0.386463 matches LRU on the collapsed stream 0.410488 within 25% -- the engine measured the bridge PASS d1: LRU on the real id stream 0.038678 reproduces the offload oracle's 0.038678 -- the model was right about the trace PASS d1: the fixed mapping keeps 13712 of 13712 block ids distinct PASS d1: verdict is BRIDGE COLLISION, S46 is withdrawn as real-trace evidence, and S50 stays unauthorised -- the flags agree with the numbers PASS d3: no divergence in either APC-off cell (0, 0) -- the effect is cache-state dependent PASS d3: APC on / invariant 0 diverges on 6 boot-prompts -- S53's cold-start finding reproduces under serialised requests PASS d3: APC on / invariant 1 diverges on 0 -- the shipped flag closes it PASS d3: the verdict names the existing flag as the fix, not a contribution PASS d3: the warm response is byte-identical across every boot in every cell -- the divergence is a deterministic cold-start effect, not run-to-run noise PASS d4 attempt 1 refused itself (profiler endpoint absent), invented no split, and recorded the engine's 'Unknown vLLM environment variable' line as the cause PASS d4 attempt 3: all 4 cells recorded zero kernel events from a profiler_out_*.txt table parsed as JSON -- its split of zeros is unquotable PASS d4: every cell recorded thousands of kernel events ([26596, 59640, 49852, 44402]) -- not attempt 3's zeros PASS d4: per-step wall agrees with the sweep's fitted a + b*running within 4.2% at every B -- the step count is not inflated PASS d4: the attention slope 0.05342 re-derives from the cells (0.05342) PASS d4: attention is 49.5% of the per-sequence slope -- the largest class, under the pre-registered 60% line, and the verdict says so PASS d4: the CPU gap is 2.2% of the per-sequence slope -- no host-side residual survives MRV2 PASS d4: the measured per-sequence slope 0.10789 is more than twice the KV-bytes roofline 0.03643 -- the roofline's omission is confirmed, and named PASS agentic_real_cc provenance: content hashed and discarded, class DERIVED_FROM_REAL PASS agentic_real_cc: 130942 API calls over 394 programs -- a sample, not an anecdote PASS agentic_real_cc: median cache-read share 0.997205 -- agent prompts are almost entirely re-read prefix, as the public numbers say PASS agentic_real_cc provenance carries no path, branch or prompt fields PASS agentic_residency_cc: hold shares by wait class sum to 1.000 PASS agentic_residency_cc: the hold shares re-derive from the unbounded 1 h run PASS agentic_residency_cc: 98% of hold cost is in human-paced or > 6 s waits -- the gateway-visible classes PASS agentic_residency_cc: oracle park recomputes nothing and restores from host -- the park/wake claim's first number is a bound, and says so PASS agentic_residency_cc: LRU at 1 h hold keeps hit rate >= 0.98 down the capacity ladder -- one operator's sessions rarely overlap; the fleet curve is TraceLab's PASS x2: fork-aware saving is 0.088432 of peak live KV, under the 0.1 kill line, and the verdict says KILLED PASS x2 measured 4613 subagent programs -- a fan-out corpus, not a case PASS cache_ledger_cc: re-created tokens by cause sum to the total (977614116) PASS cache_ledger_cc: avoidable share of cache spend 0.094598 is under the 10% gate and the verdict says so -- a report, not a product, on this corpus PASS cache_ledger_cc: the rate card is marked UNVERIFIED -- dollars are conditional, tokens are the ledger PASS cache_ledger_tracelab: re-created tokens by cause sum to the total (951988358) PASS cache_ledger_tracelab: avoidable share of cache spend 0.083019 is under the 10% gate and the verdict says so -- a report, not a product, on this corpus PASS cache_ledger_tracelab: the rate card is marked UNVERIFIED -- dollars are conditional, tokens are the ledger PASS agentic_residency_tracelab: 99% of hold cost in gateway-visible waits on 43 developers -- the Claude Code shape reproduces PASS agentic_residency_tracelab spans 357149 calls PASS agentic_public_traces: two traces loaded with licences recorded; Exgentic is marked NOT loaded rather than counted PASS r4: the request reaches the upstream byte-identical while finish_reason and the session are read from what the proxy already relays PASS r4: the free signal reaches precision 0.8515 / recall 0.981 against oracle park-after-6s on the Claude Code trace PASS r4: and 0.7374 / 0.8984 on TraceLab -- the shape holds across operators PASS r1: the sealed gate is recorded as NOT passed on the default card PASS r1: keep-5-minutes-then-re-prefill is the cheapest policy on every (trace, model) cell -- parking loses to recompute on PCIe fleets at these context sizes PASS r1: 2 of 36 sensitivity cells pass, all with free link traffic and a >= 200 GB/s host link -- the hardware condition is in the artifact PASS r2: the joint gate is recorded as failed on all four cells PASS r2: checkpointing the sandbox ALONE saves >= 30% vs holding both on every trace -- the surviving product is the sandbox half, not the joint controller PASS r2: on PCIe the joint policy collapses to sandbox-only to the cent -- the joint claim adds nothing there PASS l2: the $30/dev-month gate is recorded as NOT passed on every trace PASS l2: the oracle saves and the tool-median adaptive policy loses on every trace -- the gap is prediction, not policy PASS l2: always-1h beats the 5-minute default by $36.02/dev-month on TraceLab -- the configuration residue is in the artifact PASS t4: the refit TOOL_MIX has 8 measured tools summing to one, with the human-paced class present PASS t4: the recorded eviction-loses-everywhere flags re-derive from the per-capacity rows PASS t4: eviction loses to LRU at every capacity on BOTH mixes and the verdict says REPRODUCES -- the policy, not the generator PASS t4: the refit moved the generator's reuse (0.9083 -> 0.8724) -- the calibration was not a no-op PASS d8: the headline medians (top-25% share 0.40915, normalized rank 0.8914500000000001) re-derive from the 16 per-layer rows PASS d8: top-25% share 0.40915 is between the sealed gates (0.40 / 0.60) and the verdict says so PASS d8: normalized effective rank 0.8914500000000001 -- the bank spans most of its nominal directions; 2607.28308's finding on our model PASS d8 records the observed GPU name -- unlike sprint18_expert_census, something establishes a GPU was involved PASS s46b: engine hit rate 0.037596 is under the sealed ceiling 0.169 -- the fixed bridge no longer manufactures sharing PASS s46b: engine hit rate 0.037596 agrees with the hash-only LRU model 0.038678 within 10% -- the $0 model and the engine's own counters now say the same thing PASS s46b: the APC-off control reports zero hits PASS s46b: knee ratio apc_off/apc_on 0.926664 at n=1 -- APC still helps, by a fraction of the withdrawn 0.766 PASS s46b: the run is not void PASS s46b's sweep certificate carries its knees (apc_off 8.2904, apc_on 8.9465) -- the analysis quotes the sweep, not a recollection PASS d2: prose control at 1.0396 positions per pass -- the oracle finds ~nothing to commit on free text, as it must PASS d2: the BUILTIN JSON grammar forces 0.0 of structured tokens -- schema-less JSON mode has nothing to jump; every grammar gain comes from a declared schema PASS d2: declared-schema ceiling 1.542 on structured output is between the sealed gates (1.5 dead / 3.0 sprint) and the verdict says so PASS d2: code patches reach 4.831 from prompt-copy alone -- SpecDecode-Bench's code-editing result, reproduced and not claimed PASS d2: every scored tool call parses as JSON after wire-format normalisation (195 of 200 did not before it) PASS d1 attempt 3: all 2 arms record the 'engine_timeline' constructor failure -- the broken-dataclass attempt is on file as what it was PASS d1 attempt 3's certificate records reps_usable=0 -- it claims nothing PASS d1 attempt 1: engine_prefix_cache is null at all 16 load points -- the sharded merge dropped the counters, and the artifact records that rather than a number PASS d1 attempt 1 reproduces S46's arm knees (apc_off 8.2904, apc_on 10.516637 qps) -- the instrument was consistent even while its counters were lost PASS all 53 hypothesis documents name their evidence PASS every certificate a hypothesis document names exists PASS all 21 hypothesis documents from 032 onward carry a `## Prior art` section PASS every recorded prior-art section is dated and cites at least one URL PASS README's test count matches reality: 615 [7] S06 certificates: sealed clauses re-read from the artifacts PASS s06_c1_fixed_bridge_n8: the primary bar FAILS and the A/A bar passes, as the certificate records PASS s06_c1_fixed_bridge_n8: S46c paired median 0.949884, p 0.62503 (not significant at n=8) PASS s06_c1_fixed_bridge_n8: engine hit rate 0.0385 with APC on against the model's 0.0387, 0 hits with it off PASS s06_c1_fixed_bridge_n8: S47b A/A median 1.024752, not significant PASS s46c_apc_fixed_bridge_n8.json: paired_knee_analysis.median_of_per_rep_ratios = 0.949884 PASS s46c_apc_fixed_bridge_n8.json: does_not_prove is EMPTY and is recorded as such (STALE_CLAIMS.md; left red on purpose, never stamped) PASS s47b_aa_fixed_bridge_n8.json: paired_knee_analysis.median_of_per_rep_ratios = 1.024752 PASS s47b_aa_fixed_bridge_n8.json: does_not_prove is EMPTY and is recorded as such (STALE_CLAIMS.md; left red on purpose, never stamped) PASS s06_bailian_fleet_audit: headroom 44.5672 / 30.9651 / 21.1422 / 51.609 pp on the four Bailian files PASS s06_bailian_fleet_audit: the Mooncake positive control reproduced and the fresh-id control was refused by the runner while the instrument said OK (the finding is preserved) PASS fleet_report_92041851.json: bailian_qwen_traceA_blksz_16 cold-miss floor 0.402025 PASS fleet_report_c8e16960.json: bailian_qwen_traceB_blksz_16 cold-miss floor 0.384727 PASS fleet_report_e636fdfb.json: bailian_qwen_thinking_blksz_16 cold-miss floor 0.538158 PASS fleet_report_a465f767.json: bailian_qwen_coder_blksz_16 cold-miss floor 0.336277 PASS s06_weka_prefix_hit_rate: 0.965737 over 129,409,824 blocks reproduces the published 96.57%, controls clean PASS s06_weka_residency: 176,406,170,212 HBM token-seconds at 1 h hold on 739 public agent traces PASS s06_verifier_parity: 12 of 12 seals and 68 of 96 MDE rows recompute and 0 refusals fired; the bar failed and says so [8] S08 certificates: the verifier contract, the parity bar, the O14 rows PASS s08_verifier_contract: all 5 inputs carry their verdict in the exit code PASS s08_verifier_contract: 4 distinct exit codes over 5 verdict classes (0 verified, 2 failed, 3 no-rule, 4 unreadable) PASS s08_verifier_contract: tampering config.reps is NOT caught and is recorded as declared scope, not as a pass PASS s08_verifier_contract: the 6-file list verifier/index.html fetches imports on its own PASS s08_parity_bar: accept-everything and reject-everything are graded first and both fail the bar PASS s08_parity_bar: the two degenerate strategies fail in opposite directions (accept-all agrees with everything, reject-all with nothing) PASS s08_parity_bar: every bundled module is byte-identical to its source module PASS s08_parity_bar: C2, the clause byte-identity cannot see -- the entry point's own file list imports alone PASS s08_parity_bar: all 4 rules are tripped by a mutant of a real certificate PASS s08_parity_bar: 3 of the 4 rules are never exercised in the refusing direction by any committed certificate, and the bar says so PASS s08_parity_bar: the 28 pre-audit MDE disagreements are reported and still red, not gated away PASS s08_parity_bar: a mutation of a field no rule reads trips nothing PASS s08_o14_engine_rows: the planted and the scrubbed control both fire, so the derivation reads the completions PASS s08_o14_engine_rows: engine determinism recomputes to 6 divergent boot-prompts with caching on and 0 with it off, 0 under the batch-invariant flag PASS s08_o14_engine_rows: the engine-counter analysis re-runs with every headline value identical to the committed certificate PASS s08_o14_engine_rows: the engine's own counters read 0.0385 with caching on and 0 hits with it off ========================================================================== PASS README's audit-check count matches this run: 646 2 FAILURES - s46c_apc_fixed_bridge_n8.json: does_not_prove is present but EMPTY - s47b_aa_fixed_bridge_n8.json: does_not_prove is present but EMPTY make: *** [audit] Error 2 exit=2