Every core claim is measured on real, public, non-toy workloads and released with a one-command reproduction harness — flywheel run all-local — that reproduces the results with no database or cluster required.
EAV-DT (arXiv:2606.29280, Jun 2026) and CORTEX hallucination prevention (arXiv:2602.17691, Jan 2026) — peer-visible, publicly reproducible claims.
HELIX serving performance on the official vLLM-project GuideLLM harness (AMD EPYC 9254), plus the confidence-score evidence behind the Inference Engine.
Claim-by-claim source links with methodological status and legal evidence labels — every headline metric maps to a documented artefact.
The layered Hybrid LLM architecture, with Kevin (sovereign Australian knowledge) and OmniTX (biomedical) as evidence instances — architecture narrative, not live products.
The Living Decision Flywheel harness reproduces the core results with no database or cluster required: flywheel run all-local.
On a real 251k-case procurement log, the continually-retrained flywheel beats a frozen policy (signed-rank p≈0.009). A label-permutation placebo collapses the gain 389× — the value is freshness, not volume. M
411× cheaper than full re-materialisation; 13–814× vs a strong incremental baseline on a real 251k-cell workload. M
Zero argmax mismatches between the C++/ONNX production engine and the Python reference, with ~20× tighter tail latency. M
Sub-layer latent activation steering recovers 98.5% of native FP16 reasoning at 4-bit (73.7% vs 74.8% FP16). M
A randomised trial + heterogeneity sweep establish a falsifiable rule: a learned policy beats a fixed rule only in the regime of pricing, discounting, and material re-sourcing. M
On OULAD, the policy shows no significant disparity by disability or deprivation — auditable via lineage, not asserted. M
HELIX v1.7 Build #58 performance is documented separately on the benchmarks page (official vLLM GuideLLM, AMD EPYC 9254).
Book a call for the full evidence dossier and reproduction access.