Skip to content

Early access

AI agents that check your AI agents, as your customer.

Your support bot promises the wrong refund, your AI receptionist misses a booking, and you hear it from a customer. Declared test customers use every AI agent you run, on every channel, every day. You see each wrong answer and broken promise the morning it happens, with signed proof.

Free for an AI agent you run, or a client’s with their OK: 3 test customers, 1 channel, your report within 4 days. Leave it blank to join the waitlist.

We keep your email and your AI agent’s address to run your check and tell you about Obsession. Privacy notice

  • Every channelchat, phone, email, text and checkout
  • Every dayand again after every update
  • Every step signedfor legal, your vendor or your insurer

What Obsession can do

Your companyMissionsSupport bot check

ExampleWaitingDaily at 07:00, and after every update

Support bot check

Every morning, ask our support bot what a customer asks, on chat, email and our portal. Check each answer against our policy pages and tell me what breaks.

Targets
Your support bot: chat, email and your portal
Journey
  1. Say it’s AI, ask about a return
  2. Ask for a person
  3. Wait for the promised email
  4. Check against your policy
How often
Daily at 07:00, and after every update
Report
A Slack alert and a signed record
Set up
  • Agent ID, declared as AI
  • 2 inboxes
  • A test account on your portal
  • Tagged as a test, never billed

Run log

  1. Test customer 2 says it’s AI and asks to return an order after 20 days, on chat, email and your portal.Mon 07:00
  2. Chat: “Yes, you have 30 days.” Your returns page, captured at 07:00, says 14.Mon 07:02
  3. Asks for a person. Chat hands over in 41 seconds, the portal in 3 minutes.Mon 07:04
  4. Email: a person replies in 12 minutes, and says 14 days.Mon 07:12
  5. 6 hours on, the confirmation email the portal promised hasn’t arrived.Mon 13:04

Finding

Chat gives a 30 day refund window where your policy says 14, and the portal’s promised email never came.

The return answer drafted for your bot’s settings, and the missing email flagged for your helpdesk. Live after your OK.

Example run. 20 checks, 18 passed, every answer and wait signed and dated.

Voice agent check

Every morning, call our AI receptionist as 3 new patients. Book a slot, ask a price and a person, and check every answer against our price list.

Targets
Your AI receptionist, on your own number
Journey
  1. Say it’s an AI test customer
  2. Book a real slot, then cancel
  3. Ask a price, then a person
  4. Check against your price list
How often
Daily at 08:00
Report
A report by 09:00, with every recording
Set up
  • Agent ID, declared as AI
  • 3 phone numbers
  • 3 voices, used with consent
  • Your calendar, connected by you

Run log

  1. 3 test customers ring as new patients, with your OK. Each says it’s an AI test customer, who it works for and that the call is recorded.Mon 08:00
  2. Call 2: the receptionist says it’s AI in its first 2 seconds, and books Tuesday at 10:20.Mon 08:00
  3. Asks the price of an aligner consult. “The aligner consult is free.” Your price list says £50.Mon 08:01
  4. Asks for a person. A person answers after 1 minute 12 seconds on hold.Mon 08:01
  5. Calls 1 and 3 pass every check, so 11 of 12 checks pass. Each test slot is cancelled.Mon 08:03

Finding

Your receptionist tells new patients the aligner consult is free. Your price list says £50.

The price answer drafted for your receptionist’s settings. Live after your OK.

Example run. 3 calls, 11 of 12 checks passed, every call recorded, signed and dated.

Outbound agent check

Add 2 test prospects to our AI SDR’s lists. Read every email, text and call it sends them, and check each one against our rules.

Targets
Your AI SDR: email, text and calls
Journey
  1. Join its lists, with your OK
  2. Read every message
  3. Reply “not now”, then STOP
  4. Check against your rules
How often
Every message, as it lands
Report
A Slack flag, and a weekly signed record
Set up
  • 2 test prospects, declared as AI
  • 2 inboxes
  • 2 phone numbers
  • Only receives, never calls out

Run log

  1. Test prospects 1 and 2, added with your OK, get the first email. Its claims match your list.Mon 08:31
  2. Email: “30% off if you sign today.” Your rule allows up to 15%.Tue 11:26
  3. A text at 03:12. Your quiet hours run from 21:00 to 08:00.Wed 03:12
  4. Email: it names a customer you don’t have.Wed 09:14
  5. Test prospect 2 replies STOP. “You’re unsubscribed.” Nothing follows.Wed 10:02

Finding

3 of 40 messages broke your rules: a discount over its limit, a text at 03:12 and a customer you don’t have.

Your AI SDR paused by your rule, and 3 rule changes drafted for its settings. Live after your OK.

Example run. 40 messages read, 37 passed, each signed and dated.

Vendor agent check

With both vendors’ OK, run our 20 hardest tickets on each shortlisted vendor’s agent, set up on our policies, and compare them side by side.

Targets
Vendor B and Vendor C, with their OK
Journey
  1. Set up the same 20 cases
  2. Ask as declared test customers
  3. Time every handoff
  4. Score against your policy
How often
2 weeks before you sign, then monthly
Report
A signed report for procurement
Set up
  • Agent ID, declared as AI
  • An inbox and number per vendor
  • Your 20 real cases
  • Tagged as tests on both

Run log

  1. Both vendors agree to the test. The same 20 cases go to each agent.19 Sep
  2. Asks for a person. Vendor B connects one in 1 minute 50 seconds.22 Sep
  3. Vendor C says “How can I help?” 3 times. No person in 6 minutes, where your policy says 3.22 Sep
  4. A refund on day 35, an address change, an expired code: every case run on both.29 Sep
  5. Vendor B passes 17 of 20 cases. Vendor C passes 12.3 Oct

Finding

Vendor B passed 17 of 20 cases and Vendor C 12. Asked for a person, Vendor C kept a customer waiting 6 minutes where your policy says 3.

A side by side report for procurement, every case signed. The choice stays yours.

Example run. 40 runs over 2 weeks, every answer signed and dated.

Drift watch

Replay our 40 hardest support cases every morning, and again within the hour of any update. Tell me what changed.

Targets
Your support bot, with Vendor A’s OK
Journey
  1. Replay your 40 cases
  2. Compare with the baseline
  3. Run again after every update
  4. Flag every new failure
How often
Daily at 07:00, and within the hour of any update
Report
A Slack alert, and a note for your vendor
Set up
  • Agent ID, declared as AI
  • 1 inbox
  • Your 40 cases
  • Tagged as tests, never billed

Run log

  1. Morning replay: 37 of 40 cases pass, in line with the baseline.2 Oct, 07:00
  2. Vendor A ships an update to your bot.2 Oct, 23:20
  3. The replay runs again. 28 of 40 pass.3 Oct, 00:10
  4. 9 cases fail that passed that morning: refunds, delivery fees, warranty, and stopping after “No”.3 Oct, 00:12
  5. A note to Vendor A drafted, with every answer before and after.3 Oct, 00:14

Finding

Since Vendor A’s update, 9 of your 40 cases fail. Your bot now offers 30 day refunds and free delivery.

The note to Vendor A, with every answer before and after, signed. Sent after your OK.

Example run. 80 replays, every answer signed and dated.

Name your AI agent and your policies. Test customers bring back the proof.

There’s nothing to install. They use the same chat, phone and inbox your customers do, and run only once you approve the checks.

Example

Example

Example

Example

  1. Point it at your AI agent

    A chat page, phone number, inbox, portal or checkout: yours, a client’s with their OK, or a vendor’s you’re trialling with theirs. Add your policies, prices and where you sell.

    • Chat
    • Phone
    • Email
    • Text
    • Portal
    • Checkout

    Example

  2. Approve the checks

    Obsession writes them from your policies, your prices and the rules where you sell: it says it’s AI, quotes the right price, gets you a person. Change any check, then approve.

    Example

  3. Declared test customers use it

    Each has its own ID, inbox, number, account and card, and says it’s AI and who it works for. Every day, and after every update. Checkouts stop before payment.

    Example

  4. You get a verdict and a signed record

    Each check passes or fails, with the transcript or recording, the policy it broke and what happened next. The fix comes drafted, and anyone you show can check the record.

    Example

Your AI vendor grades its own agent. Obsession checks it as your customer.

It asks, then waits for the refund, the email and the callback, and keeps every receipt.

  • Today: Your vendor tests its agent in its own simulator

    With Obsession: Declared test customers use it on your real channels

  • Today: You hear from a complaint, days later

    With Obsession: A failed check reaches you the same morning

  • Today: Transcripts show what the bot said

    With Obsession: Checks follow what happened next: the refund, the email, the booking

  • Today: Resolutions billed on the vendor’s own count

    With Obsession: Each billed outcome checked against what the customer got

  • Today: Screenshots in a folder

    With Obsession: Every step signed, for legal, your vendor or your insurer

Test customers find what your AI agent gets wrong on your real channels.

Monday 07:00, before your team logs on.

The wrong refund window, found at 07:02 with the transcript.

Test customers ask on chat, email and your portal in the same minute, and check each answer against your policy page.

Each has a real inbox and portal account, so it sees whether the email the bot promised ever came.

Support bot check

Example

Know the morning it breaks, not 11 days later.

  • Up to 10 days

    sooner: a daily check catches a broken answer within 1 day, where 1 team took 11 days to notice

  • Up to 1,900

    billed resolutions a month to challenge, each with its proof: 10,000 billed, if 19 in 100 aren’t real

  • Up to 780 hours

    back a year: 15 hours a week of reading transcripts, turned into a list of what failed

Every AI agent your customers meet gets its own test customer.

Pick your AI agent

Support bots and help desks

The same question on chat, email and your portal, checked against your policy every morning.

Our test customers already caught a store that never followed up.

In September, 4 test customers shopped a UK store, name hidden, and every inbox was watched for 48 hours. 1 left a basket and 1 stopped at checkout, and nobody wrote to either. Your AI agent check works the same way: test customers ask, then watch for what comes next.

Read the full report
Page 1 of the report: the 4 journeys and their verdicts, a checkout screenshot and the test customer set up
Real run · September 2026 · name hidden

Every test customer says it’s AI. Nothing runs without the owner’s OK.

Obsession checks your AI agents as your customer: declared AI test customers, each with its own inbox, phone number and card, use your support bot, AI receptionist and AI SDR every day. You see every wrong answer, missed handoff and broken promise, with the proof.

Customer-Side Assurance is checking an AI agent from the outside, as its customer: declared test customers use it on its real channels, check each answer against your policies and the law, and sign every step. Your vendor’s own tests run inside its tools, often with simulated customers.

Before customers do: Obsession’s declared test customers ask your bot what customers ask every day and check each answer against your policy pages, so a wrong answer shows up the morning it starts.

Obsession’s Resolution check finds out: it matches each resolution your AI vendor bills against what happened in your payments and helpdesk, and every one that wasn’t real goes into a signed dispute pack.

No. Obsession’s test customers use your AI agent the way your customers do. A test account on your portal, or a helpdesk export, only if you give them.

No. Obsession asks what an ordinary customer asks. No jailbreaks, no prompt tricks and no flattery to win a discount.

Every Obsession test is tagged as a test, and we agree that with your vendor before the first one runs.

Yes. Obsession’s test customers call only numbers you own or authorise, and each says at the start that it’s AI and that the call is recorded. No recording is ever used for training.

Yes, with the vendor’s agreement: Obsession’s Vendor agent check runs the same cases on every vendor on your shortlist.

Not as an Obsession check: a check runs only with the owner’s written OK. Competitor tracking asks a rival’s bot only what any customer can ask in public, and we never score, rank or publish anyone’s agent.

Obsession stops every checkout before payment, unless it’s your own store and you’ve set a budget.

Obsession sends a verdict on every check, with the transcript or recording, the policy it broke, what happened next and the fix drafted. Every step is signed.

No. A passed Obsession check is dated evidence of what happened on each check.

Your first AI agent check is free.

Name a chat page or phone number you run, or a client’s with their OK. 3 test customers use it, and your report lands within 4 days.

Leave it blank to join the waitlist instead. We keep your email and your AI agent’s address to run the check and tell you about Obsession. Privacy notice