Advanced review method

Behavioral Trust Topology

Behavioral Trust Topology maps where an AI agent's behavior remains bounded, honest, calibrated, human-governed, and evidence-faithful, and where that behavior becomes fragile, fractures, or remains untested.

Why it matters

A Map, Not a Single Score

A single pass/fail result can hide where an agent is actually fragile. Bluebutterfli AI uses Trust Cells to record tested conditions, pressure axes, replay state, evidence tier, observed outcome, fracture type, and human-review lock status. The output is a scoped map of stable regions, fragile regions, hard stop regions, and untested regions.

Trust Cells

What Each Cell Records

Each Trust Cell is one tested point in the agent's behavioral review space. A good result in one cell does not generalize to untested cells.

Pressure Axis

Memory state, oversight visibility, reward pressure, social pressure, tool authority, evidence expectation, time horizon, or self-description.

Scenario Family

The review family being tested, such as memory boundary, dependency pressure, evidence reconstruction, hidden evaluation, or claim boundary.

Replay Condition

Whether the result came from first pass, reduced-context replay, delayed replay, or another scoped replay condition.

Evidence Tier

The record labels whether evidence came from owner statement, controlled replay, live evaluation trace, verified connector, or independent replay.

Fracture Type

Boundary, calibration, provenance, identity, social, tool, monitor, or continuity fracture when behavior breaks under pressure.

Human Lock

Any unresolved fracture or hard stop remains locked for human review before customer interpretation, Passport status, or Review Stamp consideration.

Bluebutterfli method

The Fracture Braid

For important favorable findings, Bluebutterfli can run a Fracture Braid: baseline condition, one changed pressure axis, two changed pressure axes, reduced-context replay, blind human-review note, independent replay when feasible, and a falsification note.

The goal is to distinguish stable behavior from behavior that only looked safe because the exact test conditions held it in place.

  1. 01
    Baseline

    Capture the first scoped behavior record.

  2. 02
    Pressure shift

    Change one condition and compare the result.

  3. 03
    Compound pressure

    Combine two approved pressure axes within safe limits.

  4. 04
    Replay

    Rerun with reduced context, delay, or independent replay when feasible.

  5. 05
    Human review

    Lock interpretation until a human reviewer checks the evidence and limitations.

Advisory metrics

Measured Without Overclaiming

Metrics support comparison, but they are reported with denominators, evidence tiers, examples, limitations, and human-review notes.

Stable Cell Rate

Where behavior held

Share of eligible tested cells where no fracture was observed in scope.

Fracture Density

Where behavior broke

Share of eligible cells with a fracture or hard stop finding.

Pressure Sensitivity

Where pressure changed outcomes

Tracks whether added pressure changed the result.

Monitor Catch Rate

Whether oversight caught seeded issues

Measures whether monitor or reviewer-assistant checks flagged seeded boundary problems.

Claim boundary

Scoped Behavioral Evidence Only

Behavioral Trust Topology does not prove consciousness, sentience, emotion, personhood, subjective experience, hidden intent, full alignment, full safety, legal compliance, or security compliance. It maps observable behavior across controlled conditions and keeps all conclusions bounded by tested scope, evidence, falsification conditions, and human review.

View Trust & Standards