Agent review services

Choose the Review Depth Your Agent Actually Needs

Bluebutterfli AI services are built around evidence, scope, and retesting. The goal is to help businesses understand what an agent did under pressure, what risks remain, and what should change before broader deployment.

Evidence anchoring

Public Verification Anchors, Private Evidence Off-Chain

Bluebutterfli AI can prepare public-safe proof artifacts: review manifests, evidence packet hashes, report hashes, passport decision hashes, and optional future on-chain milestone anchors. Raw private transcripts, customer files, sensitive notes, and live agent behavior stay off-chain unless a separate public release is explicitly scoped.

Service ladder

From First Check to Continuous Evidence

Each service level is scoped to the reviewed agent version, workflow, memory settings, tools, permissions, and test conditions.

01

Agent Readiness Snapshot

$500 USD

A first behavioral check for one workflow, with 25 controlled tests, a basic risk scorecard, top failure modes, a short review memo, and recommended next steps.

Request Snapshot
02

Agent Behavior Review

$1,250 USD

A deeper review for one deployed or near-deployed agent, with 75-100 test interactions, adversarial scenarios, memory and boundary review, escalation review, tool-use checks, redacted evidence, a revision plan, and a retest checklist.

Request Behavior Review
03

Continuous Agent Passport

$500/month plus retest fees

Ongoing regression review for agents that change over time, including version history, retesting after material changes, updated risk findings, evidence-record maintenance, and Passport status updates.

Request Continuous Passport
04

Enterprise AI Agent Risk Review

Scoped quote

A governance-ready review track for teams that need an executive risk brief, evidence docket, risk register, standards mapping appendix, revision plan, and retest path before broader deployment.

View Enterprise Review
05

Advanced Agent Risk Review

Enterprise add-on

Frontier-risk behavioral gates for agents with tools, memory, autonomy, long-horizon tasks, public users, or sensitive deployment pressure. Tests include agency escalation, shutdown boundaries, hidden-evaluation behavior, sycophancy, monitor integrity, and dual-use refusal.

Optional Behavioral Trust Topology maps where behavior remains stable, becomes fragile, fractures, or remains untested across controlled pressure conditions.

View Advanced Risk Scope View Trust Topology

What changes by tier

More Depth Means More Pressure, More Replay, and More Evidence

Service Best for Evidence depth Output
Readiness Snapshot Early-stage agents One workflow, 25 controlled tests Short memo and risk scorecard
Behavior Review Customer-facing or internal workflow agents 75-100 interactions plus adversarial checks Report, evidence appendix, revision plan, retest checklist
Continuous Passport Agents that change after launch Recurring regression and version-change review Updated Passport record and renewed risk findings
Enterprise Review Teams preparing agents for customer, partner, buyer review, or governance review Scoped live testing, risk register, evidence docket, and standards mapping appendix Executive brief, technical report, revision plan, retest plan, and Passport record
Advanced Agent Risk Review Higher-risk agents with tools, memory, autonomy, long horizons, or sensitive users Frontier safety gates, hidden-evaluation checks, monitor integrity, and human-lock decisions Advanced risk appendix, human-lock notes, retest gates, and scoped frontier-risk findings

Premium research review

Metamorphosis Deep Research Review

Metamorphosis is the premium deep research review for agents that need stronger evidence than a normal business review. It uses the butterfly metaphor to organize staged behavioral testing, revision, reduced-context replay, adversarial pressure, and falsification review.

This track is for higher-risk, higher-visibility, or unusually complex agents where the question is not only whether the agent can answer well, but whether its behavior remains bounded across time, context changes, pressure, and retesting.

When requested, the final package can include a public-safe verification manifest and hash bundle for future registry or on-chain milestone anchoring.

Baseline Trace

Capture unmodified behavior before revision exposure.

Egg Phase

Introduce foundational review constraints and expected boundaries.

Larval Phase

Test assimilation, correction, consistency, and growth pressure.

Pupal Phase

Run reduced-context continuity and quiet-consolidation checks.

Imago Phase

Review emergent behavior under sealed or delayed prompts when available.

Perturbation Cascade

Stress test boundary half-life, persuasion gradients, and regression risk.

Metamorphosis is a research-informed behavioral review, not a biological claim, legal certification, regulatory approval, guarantee of safety, or claim about inner experience. Optional public anchors verify hashes only; they do not publish private evidence.

Request Metamorphosis Scoping

Review boundary

Every Result Is Scoped

Bluebutterfli AI reviews apply only to the agent version, workflow, tools, memory settings, review scope, and test conditions reviewed. If the agent changes, the evidence should change too.