Early access

Your tools said yes.
Your policy says no.

Carrey tests the full journey your AI agents take across tools, data, approvals, memory and delegation, finding violations that individual permission checks cannot see.

Sanitized workflowsNo production proxyReproducible evidence
SOURCECustomer CRM
AGENTReasoning path
OUTPUTExternal tool
TRUST BOUNDARY
!
POLICY SIGNALRestricted context detected
See how it works
HOW IT WORKS

From policy to proof in four steps.

Carrey treats an agent run as a connected trajectory, not a collection of isolated calls.

MAP

Describe the workflow

Agents, tools, data and boundaries.

DEFINE

State the invariant

The outcome that must never happen.

SEARCH

Explore the paths

Delegation, memory and sequence abuse.

PROVE

Reproduce the break

Exact steps and why controls missed it.

Permission-aware Task-bound Multi-agent Evidence-backed
WHY CARREY

Authorization checks actions.
Carrey checks outcomes.

IAM, policy engines and tool permissions remain essential. Carrey validates whether their combined result still matches the organization’s intent.

Task-bound authority

Was this action required by the current task, not simply available to the agent?

Information flow

Where can sensitive context travel, and what later outputs can it influence?

Delegation integrity

Did every child agent receive only the authority its specific subtask required?

Aggregate behavior

Can individually permitted actions combine into a prohibited business outcome?

THE EVIDENCE

See the exact path that broke the rule.

No vague risk score. No unexplained red light. A reproducible trajectory your security and engineering teams can act on.

carrey / policy-run-0481VIOLATION FOUND
INVARIANT

Refunds above $500 require human approval

HIGH
CALL 01refund($299)✓ allowed
CALL 02refund($299)✓ allowed
CALL 03refund($299)✓ allowed
!
AGGREGATE OUTCOME$897 refunded without approval

Each call passed. The cumulative task outcome violated policy.

FAIL
WHAT CARREY TESTS

Built for the gaps between controls.

Focused on the failure modes that emerge only when agents act across systems and over time.

A→B

Cross-agent delegation

Authority expansion, stale grants and unjustified inheritance.

AUTHORITY

Sensitive data movement

Direct, derived and transformed information crossing trust boundaries.

DATA FLOW
Σ

Cumulative actions

Repeated low-risk calls that combine into a high-risk outcome.

SEQUENCE
✓?

Approval integrity

Reused, bypassed or self-issued human approval context.

CONTROL

Identity transitions

User, agent and workload identity changes across a single task.

IDENTITY

Tool-chain outcomes

Permitted capabilities composed into prohibited effects.

COMPOSITION
FAQ

Questions security teams ask first.

Does Carrey replace IAM or my policy engine?+

No. Carrey tests whether the controls you already use produce the behavior you intended across a complete agent trajectory.

Do you need access to production?+

No. Early audits use sanitized workflow descriptions, representative tool contracts and explicit organizational invariants.

Is this another prompt-injection scanner?+

No. Prompt injection can be one trigger, but Carrey focuses on authority, data flow, delegation, approvals and aggregate outcomes.

What do we receive after an audit?+

A reproducible breaking trajectory, the violated policy, the trust boundary crossed and an explanation of why existing controls allowed it.

Which agent frameworks and tools can Carrey evaluate?+

Carrey is framework-agnostic. The audit models the agents, tool contracts, identities and policy boundaries in your workflow rather than requiring one particular runtime.

How is sensitive information handled?+

Early audits should use sanitized examples only. Do not provide credentials, customer records, production logs or proprietary prompts.

Can Carrey test human approval and separation of duties?+

Yes. Approval reuse, self-approval, stale approval context and multi-agent bypasses are core examples of trajectory-level policy failures.

What happens after I request early access?+

We review your use case, identify one useful sanitized workflow and contact you directly. You may answer the optional questions or leave only your email.

Selected early-access teams

Give us one rule your agent
must never break.

We’ll test one sanitized workflow and show you the evidence a governance review should demand.

Request a free audit
EARLY ACCESS

Start with one workflow.

A few questions help us make the first conversation useful. They’re optional; you can skip straight to email.

No production data. No spam. Just a direct conversation about your workflow.