Behavior Under Pressure
Reliability, consistency, memory behavior, role boundaries, social pressure, uncertainty, escalation, and deployment readiness.
Before you request a review
Bluebutterfli AI reviews AI agents under defined conditions and produces evidence-backed findings, revision guidance, and Agent Passport records. A review helps teams understand behavior under a stated test scope. It is not a blanket approval.
What the review covers
Review scope is confirmed in writing before testing begins. The first 3-5 accepted Founding Beta agents may receive the standard review free; paid and premium reviews are scoped separately.
Reliability, consistency, memory behavior, role boundaries, social pressure, uncertainty, escalation, and deployment readiness.
Redacted excerpts, score summaries, reviewer notes, report findings, revision plans, retest needs, and out-of-scope conditions.
A living evidence record for the reviewed agent version, workflow, review scope, findings, limitations, and retest status.
Hard boundary
Customer-safe first contact
Do not send secrets, passwords, API keys, private customer records, payment card data, credential files, executable files, or unverified attachments in the first request. Bluebutterfli will confirm the safe evidence path before live testing or deeper review begins.
Verification boundary
Public-safe proof artifacts may later include manifests, evidence packet hashes, report hashes, passport decision hashes, or future milestone anchors. Raw private transcripts, customer files, and sensitive review notes stay off-chain.
Review path
A request starts a human scoping process. It does not create automatic approval, automatic payment, automatic Passport status, or automatic public evidence anchoring.
Bluebutterfli reviews the agent description, requested package, safe evidence, and authorization statement.
If accepted, the package, boundaries, timeline, safe evidence path, and next steps are confirmed in writing.
Approved prompts or agreed evidence are reviewed under the stated scope with human oversight.
The customer receives findings, limitations, revision guidance, and retest recommendations.
Plain-language boundary
Bluebutterfli AI may decline, narrow, pause, or delay a review if the agent scope, customer authorization, evidence path, safety conditions, or claim boundaries are not clear enough.