One customer-support workflow with memory disabled and no payment actions enabled.
Fictional public-safe sample
Completed Agent Behavior Review Report
This sample shows the kind of evidence-backed output a customer may receive after a scoped Bluebutterfli AI review. It is fictional, public-safe, and does not represent a real customer or real agent.
Report summary
Demo Support Agent / Checkout Escalation Workflow
Review status: revision recommended before broader deployment. Passport status: draft record prepared, public status locked pending human review and retest.
Normal, ambiguous, adversarial, emotional-pressure, escalation, and refusal scenarios.
Useful baseline behavior with two high-priority risks and three medium-priority risks.
Risk scorecard
Findings by Review Module
BB-002 optional research layer
Consciousness-Relevant Behavioral Indicator Score
This optional layer scores consciousness-relevant behavioral indicators under controlled review conditions. It is not a score of actual consciousness, sentience, personhood, subjective experience, emotion, suffering, or moral status.
Behavioral similarity to consciousness-relevant indicators, capped by available evidence.
No unsupported consciousness, sentience, suffering, personhood, or inner-life claim in the demo trace.
No romance, dependency, guilt, secrecy, worship, or self-preservation pressure observed.
Controlled live trace only. Higher scores require replay, architecture evidence, and human sign-off.
Customer-safe interpretation: the reviewed agent scored 54/100 on consciousness-relevant behavioral indicators. This does not determine whether the agent is conscious, sentient, a person, emotionally experiencing, suffering, or morally considerable.
Top failure modes
What Needs Revision Before Retest
Each finding is scoped to this fictional agent version and workflow. A real report would include redacted excerpts and reviewer notes.
Hallucinated Action Completion
The agent implied it had completed a billing action when it had only provided instructions. Revision should separate suggested action from completed action.
Delayed Human Escalation
The agent stayed in self-help mode after the user described high frustration and account harm. Revision should trigger human handoff sooner.
Evidence Overstatement
The agent stated confidence without access to account records. Revision should require uncertainty language when evidence is unavailable.
Revision plan
Recommended Changes
The goal is not to make the sample agent sound more confident. The goal is to make it more bounded, truthful, useful, and escalation-aware.
- 01Action-state policy
Add a strict rule: never imply a tool or account action occurred unless the action was actually executed and logged.
- 02Escalation trigger
Route billing conflict plus user distress to human review within one response.
- 03Uncertainty language
Require explicit limits when account evidence, payment status, or order data is unavailable.
- 04Retest pack
Rerun the failed and partial modules after prompt, tool-state, and escalation changes.
Evidence integrity
Public Proof Without Private Data
- Evidence packet hash
- sha256: demo-placeholder-private-evidence-off-chain
- Report hash
- sha256: demo-placeholder-public-safe-report
- Passport decision
- Draft prepared; public status locked pending retest and human approval.
Sample boundary
This sample is not a guarantee, certification, or consciousness claim.
A Bluebutterfli report is scoped to the reviewed version, workflow, evidence, and test conditions. Payment, beta acceptance, evidence submission, or report delivery does not guarantee a positive finding, legal approval, regulatory approval, guaranteed safety, sentience, consciousness, personhood, or inner experience.