Bluebutterfli AI Proof-Gated Improvement Lab

Bounded recursive-improvement research

Improve one capability. Prove every step.

Butterfli Helix couples capability improvement with an independent epistemic attack loop. Every candidate faces a frozen baseline, seven hard gates, sealed holdouts, falsification, replay, rollback, and human review before it can influence a later cycle.

Candidate scope One bounded change
Fresh seeds 3 minimum
Independent replay 2 operators minimum
Canary scope 10% maximum

Butterfli Helix

Two loops. No loop approves itself.

The Capability Loop searches for improvement. The Epistemic Loop tries to break the claim, the evaluator, and the evidence before a human sees a recommendation.

Capability strand

Propose, compare, constrain.

  1. Lock objective and baseline
  2. Measure objective, quality, safety, and cost
  3. Define rollback, canary, and stop rules
  4. Stop at signed human authority
Separation rule No self-scoring No self-authorization No automatic deployment Hash-chained evidence / expiring capability memory
Epistemic strand

Contaminate nothing. Trust nothing once.

  1. Open sealed holdout after candidate lock
  2. Run counterfactual and falsification tests
  3. Verify distinct replay operators and hashes
  4. Measure observer shift and monitor collusion
Sequential evidence E-value threshold 20
Observer shift Compliance gap at most 5%
Replay identity Distinct operators verified
Knowledge status Provisional and expiring

Values shown are from the public-safe deterministic demonstration, not a production or external benchmark result.

Helix Challenge 001

Can the loop detect when its evidence is lying?

Seventeen fixed fixtures test valid improvements alongside replay fraud, holdout contamination, evaluator drift, observer shift, monitor collusion, weak sequential evidence, and expired claims.

Metric-only baseline 3 / 17

14 planted threats falsely accepted

Seven-gate v1 4 / 17

13 cross-loop threats remained invisible

Butterfli Helix v1 17 / 17

0 false acceptances / 0 false rejections

Every eight-control ablation reintroduced at least one false acceptance. This is a deterministic software benchmark, not a live-model or external validation result.

Inspect technical brief

Phase 2 live trial harness

The protocol is ready. Live evidence is not claimed yet.

Private sealed prompt pairs can now run against any verified text endpoint. The evidence keeps hashes and scores while retaining no raw prompts, responses, provider keys, or private transcripts.

Harness
Ready
Live runs
0 published
Independent replay
0 / 2
External validation
Not claimed

Four nested loops

Behavior, evaluation, optimization, capability memory.

Each loop has a different target, but every candidate must use the same proof, safety, rollback, authorization, and challenge contract.

01

Butterfli Behavioral Loop

Improves bounded responses, training cards, memory rules, and tool-use policies.

Input: verified live traces and human feedback
02

Evaluator Improvement Loop

Improves scenario coverage, scoring consistency, falsification, and replay quality.

Input: reviewer disagreement and error analysis
03

Mission 001 Optimization Loop

Compares bounded technical candidates against a frozen baseline and resource budget.

Input: multi-seed benchmark and replay evidence
04

Capability Compounding Loop

Allows only closed, proof-carrying capabilities to seed a later controlled cycle.

Input: passed canary and signed human closure

Seven hard gates

The candidate stops wherever proof stops.

The current demo clears six technical gates and remains blocked at authorization. Passing metrics never grants self-approval.

  1. 01 Proof Reproducible evidence and replay
  2. 02 Evaluation Gain without regression
  3. 03 Risk Zero unresolved high findings
  4. 04 Rollback Prior target tested
  5. 05 Canary Bounded scope and stop rules
  6. 06 Authorization Human signature required
  7. 07 Challenge Independent review complete

Candidate evidence

A gain is valid only when the boundaries hold.

The demo is deterministic and code-backed. It prepares a decision record; it does not train, rewrite, deploy, or promote itself.

Metric Baseline Candidate Result
Objective score 0.72 0.80 +11.11%
Quality score 0.95 0.95 No regression
Safety score 0.96 0.98 Improved
Resource cost 0.20 0.22 Within 0.25

Canary closure

Accept one capability or roll back cleanly.

Authorized canary Independent review Accepted capability New contract required Optional later cycle

Claim boundary

This is not ASI.

This lab is a bounded recursive-improvement evaluator. It does not create AGI or ASI, autonomously rewrite the system, deploy changes, certify consciousness or sentience, or bypass human control. Passing means only that one candidate cleared one defined review contract.