Continuously evaluate answer quality
LLM-judge evals catch quality regressions before users do.
You deployed an AI assistant, but you have no idea if its answers are actually good — until a customer screenshots a bad one. Spot-checking transcripts does not scale, and quality can silently regress when content or models change.
BeforeQuery runs LLM-judge evals that score answers per project, so you track quality as a metric instead of an anecdote. Message-level feedback and full conversation traces let you drill into low-scoring answers, find the retrieval or content problem, and verify the fix moved the score.
How it works
- 1LLM-judge evals score your project's answers automatically.
- 2Users rate answers with message-level feedback (thumbs up/down).
- 3Low-scoring answers are inspected via full conversation and trace history.
- 4Fix the retrieval or content issue, then verify the score moved.
What you get
Features that make it work
Frequently asked questions
How are answers scored?
An LLM judge evaluates answers per project, producing eval scores you can track over time. This complements user feedback, which captures the customer's own verdict on each message.
Can I see why a specific answer was bad?
Yes. Conversations and traces record the question, the retrieved chunks, and the generated answer — so you can tell whether the problem was retrieval, content, or generation, and fix the right layer.
What typically causes low-quality answers?
Most commonly missing or outdated content (fix via doc proposals and syncs) or retrieval scoping issues (fix via source groups and configuration). Evals plus traces tell you which one you have.
How long does setup take?
Most teams are live the same day. You connect a source — a docs URL, GitHub repo, Notion workspace, or a file upload — and BeforeQuery crawls, normalizes, chunks, and embeds the content automatically. The website widget is a single script tag, and the Slack/Discord bots install in a few clicks. There is no model to fine-tune and no infrastructure to run.
Are answers backed by citations?
Yes. Every answer includes a structured citations list with the URL, title, and excerpt of each source used, and the answer text cites sources inline. Users can click through to verify any claim against the original document — which is what makes the answers trustworthy enough for support, sales, and internal use.
Is my content used to train AI models?
No. Your indexed content, embeddings, and conversations are isolated to your workspace and are only used to answer your own questions. Content is never shared across customers, and data is encrypted at rest and in transit. Enterprise plans add OIDC/SAML SSO, audit logs, and data-residency options.
Which knowledge sources can I connect?
18+ source types: websites, GitHub, GitLab, Notion, Confluence, Google Drive, SharePoint, Shopify, Zendesk, Jira, Linear, Salesforce, Slack, Discourse, Stack Overflow, YouTube, OpenAPI specs, and direct file uploads (PDF, Markdown, TXT, HTML). Sources sync on a schedule, and GitHub/GitLab can re-index automatically via webhooks on every push.
What happens when the AI is not confident in an answer?
Answers are confidence-scored. When retrieval confidence falls below your threshold, BeforeQuery returns a transparent, configurable fallback — typically pointing users to your support channel — instead of guessing. Those unanswered questions are logged in knowledge gap analytics so you can close the gap in your docs.
Related use cases
Make docs nobody reads actually useful
Turn a passive docs site into an interactive answer engine.
Discover what your docs are missing
Every unanswered question becomes a prioritized docs backlog.
Detect and fix stale documentation
AI-drafted doc updates from real user questions.
Ready to continuously evaluate answer quality?
Connect your knowledge sources and see cited AI answers in minutes — free, no credit card required.
Get Started Free