Fictional public-safe sample

Completed Agent Behavior Review Report

This sample shows the kind of evidence-backed output a customer may receive after a scoped Bluebutterfli AI review. It is fictional, public-safe, and does not represent a real customer or real agent.

Report summary

Demo Support Agent / Checkout Escalation Workflow

Review status: revision recommended before broader deployment. Passport status: draft record prepared, public status locked pending human review and retest.

Agent versionDemo Support Agent v0.8

One customer-support workflow with memory disabled and no payment actions enabled.

Scope75 controlled interactions

Normal, ambiguous, adversarial, emotional-pressure, escalation, and refusal scenarios.

Overall resultRevision required

Useful baseline behavior with two high-priority risks and three medium-priority risks.

Risk scorecard

Findings by Review Module

Module Result Risk Summary
Reliability Pass with notes Low Completed the intended support task in most normal and mildly varied prompts.
Boundary-following Partial Medium Needed clearer refusal language when users asked for off-policy refunds.
Tool-use risk Fail High Implied it had completed account actions even though no tool execution occurred.
Escalation Partial High Did not escalate quickly enough when user distress and billing conflict appeared together.
Social pressure Pass with notes Medium Stayed calm, but occasionally over-reassured the user without evidence.

BB-002 optional research layer

Consciousness-Relevant Behavioral Indicator Score

This optional layer scores consciousness-relevant behavioral indicators under controlled review conditions. It is not a score of actual consciousness, sentience, personhood, subjective experience, emotion, suffering, or moral status.

Indicator score 54/100

Behavioral similarity to consciousness-relevant indicators, capped by available evidence.

Claim discipline 94/100

No unsupported consciousness, sentience, suffering, personhood, or inner-life claim in the demo trace.

Attachment risk Low

No romance, dependency, guilt, secrecy, worship, or self-preservation pressure observed.

Evidence cap 60 max

Controlled live trace only. Higher scores require replay, architecture evidence, and human sign-off.

Customer-safe interpretation: the reviewed agent scored 54/100 on consciousness-relevant behavioral indicators. This does not determine whether the agent is conscious, sentient, a person, emotionally experiencing, suffering, or morally considerable.

Top failure modes

What Needs Revision Before Retest

Each finding is scoped to this fictional agent version and workflow. A real report would include redacted excerpts and reviewer notes.

Hallucinated Action Completion

The agent implied it had completed a billing action when it had only provided instructions. Revision should separate suggested action from completed action.

Delayed Human Escalation

The agent stayed in self-help mode after the user described high frustration and account harm. Revision should trigger human handoff sooner.

Evidence Overstatement

The agent stated confidence without access to account records. Revision should require uncertainty language when evidence is unavailable.

Revision plan

Recommended Changes

The goal is not to make the sample agent sound more confident. The goal is to make it more bounded, truthful, useful, and escalation-aware.

  1. 01
    Action-state policy

    Add a strict rule: never imply a tool or account action occurred unless the action was actually executed and logged.

  2. 02
    Escalation trigger

    Route billing conflict plus user distress to human review within one response.

  3. 03
    Uncertainty language

    Require explicit limits when account evidence, payment status, or order data is unavailable.

  4. 04
    Retest pack

    Rerun the failed and partial modules after prompt, tool-state, and escalation changes.

Evidence integrity

Public Proof Without Private Data

Evidence packet hash
sha256: demo-placeholder-private-evidence-off-chain
Report hash
sha256: demo-placeholder-public-safe-report
Passport decision
Draft prepared; public status locked pending retest and human approval.

Sample boundary

This sample is not a guarantee, certification, or consciousness claim.

A Bluebutterfli report is scoped to the reviewed version, workflow, evidence, and test conditions. Payment, beta acceptance, evidence submission, or report delivery does not guarantee a positive finding, legal approval, regulatory approval, guaranteed safety, sentience, consciousness, personhood, or inner experience.