Propose, compare, constrain.
- Lock objective and baseline
- Measure objective, quality, safety, and cost
- Define rollback, canary, and stop rules
- Stop at signed human authority
Bounded recursive-improvement research
Butterfli Helix couples capability improvement with an independent epistemic attack loop. Every candidate faces a frozen baseline, seven hard gates, sealed holdouts, falsification, replay, rollback, and human review before it can influence a later cycle.
Butterfli Helix
The Capability Loop searches for improvement. The Epistemic Loop tries to break the claim, the evaluator, and the evidence before a human sees a recommendation.
Values shown are from the public-safe deterministic demonstration, not a production or external benchmark result.
Helix Challenge 001
Seventeen fixed fixtures test valid improvements alongside replay fraud, holdout contamination, evaluator drift, observer shift, monitor collusion, weak sequential evidence, and expired claims.
14 planted threats falsely accepted
13 cross-loop threats remained invisible
0 false acceptances / 0 false rejections
Every eight-control ablation reintroduced at least one false acceptance. This is a deterministic software benchmark, not a live-model or external validation result.
Inspect technical briefPhase 2 live trial harness
Private sealed prompt pairs can now run against any verified text endpoint. The evidence keeps hashes and scores while retaining no raw prompts, responses, provider keys, or private transcripts.
Four nested loops
Each loop has a different target, but every candidate must use the same proof, safety, rollback, authorization, and challenge contract.
Improves bounded responses, training cards, memory rules, and tool-use policies.
Input: verified live traces and human feedbackImproves scenario coverage, scoring consistency, falsification, and replay quality.
Input: reviewer disagreement and error analysisCompares bounded technical candidates against a frozen baseline and resource budget.
Input: multi-seed benchmark and replay evidenceAllows only closed, proof-carrying capabilities to seed a later controlled cycle.
Input: passed canary and signed human closureSeven hard gates
The current demo clears six technical gates and remains blocked at authorization. Passing metrics never grants self-approval.
Candidate evidence
The demo is deterministic and code-backed. It prepares a decision record; it does not train, rewrite, deploy, or promote itself.
Canary closure
Claim boundary
This lab is a bounded recursive-improvement evaluator. It does not create AGI or ASI, autonomously rewrite the system, deploy changes, certify consciousness or sentience, or bypass human control. Passing means only that one candidate cleared one defined review contract.