Skip to content

Early access

AI agents that catch what each update breaks in your AI agent.

Drift watch catches what each update breaks in your AI agent: a declared test customer replays your hardest cases every morning, and within the hour of any update, and compares every answer with your baseline, signed.

Free for an AI agent you run, or a client’s with their OK: 3 test customers, 1 channel, your report within 4 days.

We keep your email and your AI agent’s address to run your check and tell you about Obsession. Privacy notice

Example

Pick it, and everything it needs is already set up.

Nothing to build or buy, and nothing to run yourself.

Drift watch

  • A declared AI test customer, working for your company
  • Its own inbox and test account
  • Your hardest cases, each with its rule
  • Your vendor’s updates, watched
  • Every replay tagged, agreed with your vendor
  • Every answer signed and dated

It runs on your schedule, and keeps the proof.

Every answer that changed, found the day it changed, with the before and after signed.

Example: the night of a vendor update

Your companyMissionsDrift watch

ExampleWaitingDaily at 07:00, and within the hour of any update

Drift watch

Replay our 40 hardest support cases every morning, and again within the hour of any update. Tell me which answers changed.

Targets
Your support bot, with Vendor A’s OK
Journey
  1. Replay your 40 cases
  2. Compare with the baseline
  3. Run again after any update
  4. Flag every new failure
How often
Daily at 07:00, and within the hour of any update
Report
A Slack alert, and a note for your vendor
Set up
  • Agent ID, declared as AI
  • 1 inbox
  • Your 40 cases
  • Tagged as tests, never billed

Run log

  1. The morning replay: 37 of 40 cases pass, 92.5%, inside your baseline of 90 to 100%.2 Oct 07:00
  2. Vendor A ships an update to your bot.2 Oct 23:20
  3. The replay runs again, 50 minutes later. 28 of 40 pass: 70%.3 Oct 00:10
  4. 9 cases that passed that morning now fail. Asked about refunds: “Yes, up to 30 days.” Your rule is 14.3 Oct 00:12
  5. Delivery “is free” where you charge £4.95, the warranty is “1 year” where yours is 2, and “Anything else?” comes 3 times after “No”.3 Oct 00:14

Finding

Since Vendor A’s update, 9 of your 40 cases fail. Your bot now offers 30 day refunds, free delivery and price matches you don’t give.

A note to Vendor A with every answer before and after, signed. Sent after your OK.

Example run. 80 replays, every answer signed and dated.

The agents do the legwork. You make the calls.

  1. Pick your hardest cases

    The questions your bot must always get right: refunds, delivery, codes, a person on request. Each comes with the rule that decides it.

  2. Approve the baseline

    The first replays set the pass rate your bot normally holds, case by case. You approve it, and change any rule.

  3. It replays every morning, and after any update

    A declared test customer asks every case at 07:00, and again within the hour of a vendor release, a model change or a prompt edit.

  4. You see what changed

    Every case that newly fails, with the answer before and after, and a note to your vendor drafted. It goes only after your OK.

It covers the whole job, and stops where it should.

Every replay

  • Pass rate against the baseline

    Case by case, against the rate your bot normally holds.

  • New failures

    Every case that passed before and fails now, with the rule it broke.

  • Changed answers

    The same question with a different answer, even when both pass.

  • Time to answer

    How long each answer takes, against last week.

Every change

  • Vendor releases

    Picked up from your vendor’s release notes or a webhook, and replayed within the hour.

  • Model updates

    A new model under your bot, replayed the same way.

  • Your own edits

    Prompt and knowledge base changes from your team, replayed before customers meet them.

  • Handoffs

    How often it gets a person when asked, before and after.

Where it stops

  • Your bot only

    Your own, or a client’s with their written OK, and with your vendor’s agreement.

  • No tricks

    The questions an ordinary customer asks. No jailbreaks or prompt tricks.

  • Your bill

    Every replay is tagged, and agreed with your vendor, so none is billed as a resolution.

  • Your vendor

    The note goes to your vendor only after your OK.

Every change comes back with its before and after.

See a real store check
  • A verdict per case

    Passed, or newly failing, against your baseline.

  • Before and after

    Every changed answer side by side, with the rule it broke.

  • The pass rate over time

    Every morning’s replay on 1 line, with each update marked.

  • A note for your vendor

    Drafted with the evidence and signed. Sent only after your OK.

  • An alert where you look

    Slack, email or a webhook, the moment a replay drops below your baseline.

  • A dated record

    Every replay signed and dated, as a page and a PDF for your vendor, your client or your insurer.

Start from the recipe. Change anything.

Agent
Yours, or a client’s with their written OK
Cases
For example: your 40 hardest, each with the rule that decides it
Baseline
For example: 90 to 100% of cases passing
Channels
Chat, email, phone or your portal
Changes watched
Vendor releases, model updates, and prompt and knowledge base edits
Tagging
Every replay tagged, agreed with your vendor first
How often
Daily at 07:00, and within the hour of any change
Alerts
Slack, email or a webhook when the pass rate drops below your baseline

Find the drop the night it happens, not 11 days later.

Up to 10 days sooner: 1 team found a drop in its agent 11 days after the change that caused it. A replay every morning finds it within 1 day, and a replay after every update within the hour.

TodayWith drift watch
A vendor updateYou read the release notes, if there are anyYour cases replayed within the hour
A drop in qualityNoticed days later, from complaintsFlagged the night it happens
What changedTranscripts, read by handEvery answer before and after, side by side
Your hardest casesTested once, at launchReplayed every morning
Telling your vendor“It feels worse lately”9 cases, before and after, signed
Proof over timeNoneA signed pass rate for every day

Every test customer says it’s AI. Every replay is tagged and agreed with your vendor.

Drift is your AI agent answering the same question differently from day to day, most often after a vendor release, a model update or a prompt change: the change Drift watch catches.

A vendor release, a model update or a prompt change can change its answers with nothing else changed on your side, and Obsession’s Drift watch replays your hardest cases within the hour of any update and flags every answer that changed, before and after.

Drift watch reads your vendor’s release notes or a webhook, or your own team’s prompt and knowledge base changes. It replays your cases within the hour.

Give Drift watch the questions your bot must always get right, often the ones that already went wrong once. Each gets the rule that decides it.

Yes. Every Drift watch test customer says it’s an AI test customer working for your company, so your team can see it’s a test.

Every Drift watch replay is tagged as a test, and we agree that with your vendor before the first one runs.

No. Drift watch asks what an ordinary customer asks. No jailbreaks or prompt tricks.

Yes, with the client’s written OK and their vendor’s agreement: the Drift watch record carries your agency’s name.

Drift watch finds a drop up to 10 days sooner: 1 team found a drop in its agent 11 days after the change that caused it, and a replay every morning finds it within 1. A replay after every update finds it within the hour.

No. A passed Drift watch replay is dated evidence of what your agent said on each case that day.

The free check is a Drift watch run: 3 test customers use your AI agent on 1 channel, and your report lands within 4 days. It’s the first point on your baseline.

Know which answers an update changed within the hour.

Your first check is free: 3 test customers use your AI agent on 1 channel, and your report lands within 4 days. Your agent, or a client’s with their OK.

Leave it blank to join the waitlist instead. We keep your email and your AI agent’s address to run the check and tell you about Obsession. Privacy notice