Someone has to check the AI's replies.
Right now, nobody's tracking who.
Backstop is a Zendesk app that tracks AI-verification load: which agents and which groups are absorbing the work of checking, correcting, or overriding AI-generated replies. It's a real cost native Zendesk reporting doesn't measure, and it doesn't stay evenly spread on its own.
One click in the ticket. One dashboard for the team.
No new workflow to learn. An agent flags a ticket the moment they catch themselves verifying, correcting, or overriding an AI-drafted reply. Everything after that runs on its own.
Flag
A one-click toggle in the Zendesk ticket sidebar marks the ticket as AI-assisted, verified by a human.
Sync
Backstop pulls tagged and untagged ticket volume by group on a schedule, no manual export, ever.
Aggregate
Each cycle, verification rate is computed per group, with small groups suppressed rather than guessed at.
Surface
The dashboard flags when that load is concentrating in one group instead of spreading out evenly.
AI replies didn't remove the work. They moved it, quietly, onto whoever checks them.
Every AI-assisted reply that goes out still needs a human to have trusted it, and trust isn't free. Someone reads it, someone decides whether to send it as-is, correct it, or override it entirely. That decision is real cognitive work. It just doesn't show up anywhere a ticketing system tracks by default.
Native Zendesk reporting counts tickets closed, response time, CSAT. None of that tells you whether one group is quietly absorbing most of the verification burden while another barely touches AI output at all, or whether that split is shifting cycle over cycle.
Left unmeasured, that load concentrates on whoever's conscientious enough to actually check the AI's work, and that's exactly the group that burns out first.
Most support orgs can tell you how many tickets used AI. Almost none can tell you which group is doing the checking, or whether that's changing. You can't fix what you're not instrumenting, and right now this specific kind of work runs dark.
The "GenAI divide" shows up inside the support queue too.
MIT's researchers describe a split between companies with high AI adoption and low actual transformation: the GenAI Divide. The organizations that escape it share a consistent trait. They treated integration as a workflow and measurement problem, not a procurement decision.
Support orgs rolling out AI-assisted replies are running the same experiment at smaller scale. The tool shipped. Whether the verification work it creates is sustainable, evenly distributed, or quietly wearing out one group, almost nobody is tracking. That's the gap Backstop is built to close.
MIT's researchers trace the failure to a learning gap between systems and organizations: the inability to integrate AI into existing workflows, structures, and culture, not a shortfall in the models themselves.Summarized from MIT Media Lab, Project NANDA, "The GenAI Divide," 2025
AI-verification load is the first thing Backstop tracks. Not the last.
Zendesk's native tooling is built for ticket volume and response time. It has real blind spots around the new kinds of work AI is creating inside support orgs. Backstop is built to grow into more of them, one focused app at a time, not one app trying to do everything at once.
- Shipped: AI-verification load, tracked by group, by cycle, with a suppression guard so small groups never get singled out on too little data.
- Next: whichever blind spot the first real pilot surfaces. That's deliberate, the roadmap gets set by what support teams actually run into, not guessed at in advance.
Backstop is early. We're looking for the teams and people who help it get real faster.
A support team already running AI-assisted replies in Zendesk, willing to connect a trial account and let us build the case study.
People who've built on Zendesk's app framework, run a support org through an AI rollout, or can pressure-test where this breaks.
Funding, marketplace introductions, or just a sharp argument for why this is wrong. All useful right now.