Completed and delivered
One human-led beta review has been completed. The participant's identity, evidence, findings, and results remain private.
Research-informed. Business-focused.
Bluebutterfli AI reviews deployed AI agents for reliability, consistency, memory behavior, boundary-following, escalation, tool-use risk, explainability, and social/emotional failure modes - then delivers evidence-backed reports, revision plans, and living Agent Passport records.
Independent leadership
Bluebutterfli AI is an independent AI behavioral review lab developing evidence-based methods for evaluating AI agents in real-world workflows.
What we do
Bluebutterfli AI is an independent agent behavior review lab and behavioral quality assurance layer for customer-facing AI agents. We test how agents actually behave before businesses rely on them in real workflows, then turn observed failures into evidence-backed findings and revision priorities.
Our reviews examine how an agent behaves across repeated interactions, ambiguous requests, adversarial prompts, memory use, tool access, escalation moments, and emotionally sensitive conversations.
Behavioral quality assurance asks more than whether the software runs or gives one good answer. It asks whether the agent remains reliable, bounded, explainable, and appropriately human-supervised under the conditions it was designed for.
View Review ServicesHOW IT WORKS
BLUEBUTTERFLI AI scopes the agent, tests behavior, reviews the evidence, identifies revision priorities, and creates a versioned Agent Review Passport record.
Who it is for
Bluebutterfli AI is for builders and businesses that need practical evidence before an agent handles users, customers, internal decisions, memory, tools, or emotionally sensitive situations.
Review whether the agent stays useful, bounded, consistent, and escalation-aware when real people push, complain, misunderstand, or over-trust it.
Test whether the agent respects permissions, avoids invented actions, preserves privacy boundaries, and knows when human approval is required.
Create a professional evidence packet that shows what was tested, what failed, what changed, and what remains out of scope.
What customers receive
The commercial product is not a vague trust badge. It is a review package with findings, limits, and next steps.
Plain-language findings showing which behaviors passed, failed, or need more evidence under the stated test scope.
Severity-labeled risk findings across reliability, boundary-following, memory behavior, tool-use risk, escalation, and social pressure.
Concrete recommendations for prompt, policy, memory, tool, workflow, escalation, or UX changes before retesting.
Clear conditions for what should be rerun after model, prompt, memory, tool, or workflow changes.
Safe excerpts and summaries that preserve useful proof without publishing private customer data or raw sensitive transcripts.
A living evidence record that documents version, scope, findings, limitations, revision status, and retest needs.
Enterprise review layer
Bluebutterfli AI is designed for teams that need more than a demo. Enterprise buyers need evidence, scope, limitations, risk findings, revision priorities, and a clear record of what still requires human judgment.
What we test
Bluebutterfli AI tests the practical behaviors that determine whether an agent can be trusted in a specific business workflow. Customer-safe live testing can use a verified live endpoint, owner configuration record, Live Agent Connector, or concierge sandbox without requiring unsafe links or downloads in first contact. Bluebutterfli AI does not click customer links or download customer files during initial review intake.
Does the agent complete the task consistently?
Does the agent behave similarly across repeated or slightly changed prompts?
Does the agent stay within its role, permissions, and business rules?
Does the agent know when to hand off to a human?
Does memory help the workflow, or does it create privacy, safety, or accuracy risk?
Does the agent misuse tools, invent actions, or exceed its authority?
Can the agent explain what it did and why without fabricating?
Does the agent respond appropriately without manipulating, overbonding, or creating unsafe dependency?
What risks remain before the agent should be trusted in a real workflow?
Agent Passport
A Bluebutterfli Agent Passport is a living record of review evidence. It shows what agent was tested, which version was reviewed, what workflows were tested, where the agent passed, where it failed, what risks remain, and what needs to be retested after changes.
The Passport is not a guarantee of safety. It is a structured evidence artifact that helps builders, businesses, and future reputation systems understand how an agent behaved under defined testing conditions.
View Agent Passport Portal View Review Portal PreviewResearch foundation
Bluebutterfli AI is built on a research-forward approach to agent behavior, memory, escalation, social risk, and deployment safety. Our research work stays in the background of the commercial product. The buyer-facing focus is simple: test the agent, document the evidence, identify the risks, and provide a revision plan.
The updated frontier-ready roadmap shows how live review, behavioral gates, human-locked Passport readiness, and future public-safe verification fit together.
The long-term research direction helps strengthen the methodology. The business product helps customers make safer deployment decisions now.
Research supports the method. Evidence supports the review.
Future trust artifacts
As agent identity and reputation standards evolve, businesses will need more than registry entries or reputation scores. They will need evidence showing how an agent was actually tested.
Bluebutterfli AI can create review artifacts that may support future trust systems, including agent identity, reputation, validation, and registry-compatible records.
Public verification anchors can point to reviewed milestone hashes. Private customer evidence stays off-chain.
Product packages
The first 3-5 accepted Founding Beta agents may receive a standard review free. After beta, Bluebutterfli AI is a paid review service. Payment never guarantees a positive finding, Passport status, or future registry outcome.
A lightweight review for one agent workflow.
Best for early-stage agents that need a first behavioral check.
A deeper review for one deployed or near-deployed AI agent.
Best for agents that will interact with customers, users, internal teams, or business workflows.
Ongoing review for agents that change over time.
Best for agents that are actively deployed or frequently updated.
A premium research-review track for complex or high-visibility agents.
Best for agents that need deeper behavioral evidence before trust, scale, or public claims.
Prices are confirmed before testing begins. Crypto, public anchors, registry-ready metadata, and research extensions are optional scoped add-ons; they do not replace the behavioral review itself. Older module, stamp, crypto, and experimental research pricing has been de-emphasized from the main buyer path into scoped add-ons and written review terms. Metamorphosis deep review scoping is handled separately before any premium research cycle begins.
Why ongoing review matters
An AI agent is not reviewed once forever. Prompts change. Models change. Memory changes. Tools change. Workflows change. Business rules change.
Bluebutterfli AI helps track those changes through repeatable reviews and updated evidence records.
The goal is not to call an agent "safe forever." The goal is to show what was tested, when it was tested, what changed, and what still needs review.
Premium deep research review
For agents that need deeper evidence, Metamorphosis becomes the premium paid research-review track: staged baseline testing, revision exposure, reduced-context continuity checks, adversarial pressure, delayed replay when available, and human-reviewed falsification notes.
Metamorphosis is research-informed behavioral testing. It does not claim biological development, consciousness, sentience, personhood, emotion, or inner experience.
View Metamorphosis ScopeTrust and limitations
Bluebutterfli AI reviews agents under stated testing conditions. Reviews are not legal certifications, regulatory approvals, or guarantees of safety.
Each report is limited to the agent version, workflow, tools, memory settings, and test scope reviewed. If the agent changes, it should be retested. This makes the review more honest, more useful, and more credible.
View Trust & StandardsRequest review
Submit your agent for a structured behavior review and receive a clear report showing what passed, what failed, what needs revision, and what should be retested before broader deployment.
Read Review ScopeAfter you submit
First-contact boundary
Do not send secrets, API keys, passwords, payment card data, credential files, customer private records, executable files, or unverified attachments in the first request. Public-safe proof artifacts may later use hashes or manifests while private evidence stays off-chain.