Replay a month of tickets through the AI before it sees a single live one

Simulations run the assistant against your past resolved tickets and produce a reviewable result set — per-topic accuracy, human-review queue, per-answer traces — so you calibrate confidence thresholds on real data instead of guessing. Enable deflection gradually by queue, tag, or confidence tier once the numbers are convincing.

What Simulations do

The safe-rollout mechanism for anything customer-facing — Support Ticket AI, form deflector, and helpdesk copilot

The AI answers, you compare

Historical Ticket Replay

  • Point at your last N days of resolved tickets in Zendesk / Freshdesk / Front / Intercom / Jira SM / Linear / Salesforce
  • The assistant answers each ticket as if it were live
  • Every answer stored with its retrieved sources and confidence score
  • Compare AI response side-by-side with what the human agent actually sent
  • Filter by queue, tag, agent, product area, or confidence tier
Sign off before you ship

Reviewable Results

  • Per-answer human review — accept, reject, edit-then-accept
  • Per-topic accuracy scores aggregated across the batch
  • Groundedness and citation-quality evals scored automatically alongside human review
  • Distribution of confidence scores helps you set the auto-response threshold
  • Export the review as a report for stakeholders
Turn on by queue, tag, or tier

Gradual Rollout

  • Enable auto-response for specific queues once simulation accuracy clears your bar
  • Ship as agent-copilot drafts on the queues you're not ready to auto-respond on
  • Confidence-tier rollout: auto-respond above 0.9, draft between 0.7 and 0.9, skip below 0.7
  • Re-run simulations after every meaningful docs change to catch regressions
  • Traces on real answers keep the review loop going after launch
Why 'trust me it works' isn't enough

The rollout question every ops lead asks

The reason support automation stalls at pilot isn't the AI — it's the risk. Nobody wants to enable auto-response and discover on Monday that the assistant misquoted a refund policy 40 times over the weekend. Simulations answer the question in advance: here's what the AI would have said to your last 500 tickets, here's how it scored, here's where humans overrode it. That's the artifact you take to a launch decision.

  • Run the same simulation before every AI-facing rollout, not just Support Ticket AI
  • Human-review queue lets support leads sign off ticket-by-ticket
  • Automated evals score groundedness and citation quality alongside human review
  • Confidence-tier configuration lets you deploy safely instead of all-or-nothing
  • Re-run after any docs migration or reranker change to catch regressions
What teams say

Deployed in production, cited by the buyers who chose it

Deployed on our docs site in an afternoon. Every answer has citations, and the abstention gate means we've never had a customer complain about a made-up answer.
PN
Priya Nair
Head of Customer Support · Supabase
The knowledge base connected to our Slack, Confluence, and helpdesk in one setup. On-call teams get the same cited answer whether they ask in chat, in the widget, or from Cursor.
TR
Tom Richter
IT Operations Manager · Grafana Labs
The gap analytics turned into a real docs backlog. Deflection went up because we finally knew which pages were missing — the AI told us.
AC
Ana Castillo
VP of Customer Experience · Clerk

Frequently Asked Questions

Common questions about Simulations

The tickets you select — historical, already resolved. The assistant answers each one from your indexed knowledge base as if it were live. Nothing is sent back to customers; the ticket state in your helpdesk is not touched. PII masking still applies before any LLM call.
Depends on volume. A 500-ticket batch typically completes in 10–30 minutes; 5,000 tickets in a few hours. Results are reviewable incrementally as they land.
Two things. First, an LLM judge scores each answer for groundedness (does the response come from the retrieved sources?) and citation quality (are the sources actually the right ones?). Second, human reviewers accept or reject per answer, and the per-topic aggregate becomes your rollout signal.
Yes. Simulations work for any answer-generating surface — Support Ticket AI, form deflector, agent workflows. The pattern is the same: give it a batch of representative inputs, review the outputs, calibrate the confidence threshold, enable the surface at that threshold.
After any change that could plausibly regress retrieval: a large docs migration, a new source added, a reranker version bump, a model upgrade on the provider side, a persona/tone change. In practice, teams that ship weekly re-run a small representative simulation weekly and a full one monthly.
The automated scoring is the same LLM-judge pipeline the live-answer evals use, so the numbers are comparable. The human-review layer is what's specific to simulations — that's the sign-off step before rollout.

Ship with numbers, not with a hope

Run your first simulation against last month's tickets and see the accuracy, the traces, and the confidence distribution before you enable anything customer-facing.

Get Started Free