Your OpenAI, Anthropic, or Gemini key. Your invoice. Your DPA.

On Enterprise, BeforeQuery calls the model on your provider account with your key. You keep the model choice, the usage bill, and the data-processing agreement you already have with OpenAI / Anthropic / Google. BeforeQuery is the platform layer; the LLM traffic runs on you.

How BYOK works

Per-workspace credentials, per-knowledge-base overrides — every customer-facing answer runs on the account you configured

OpenAI, Anthropic, Gemini

Bring Your Provider Credential

  • Configure a provider credential per workspace
  • Optionally override per knowledge base — different KBs on different providers
  • Credentials encrypted at rest with per-workspace envelope keys
  • Rotate or revoke in the dashboard; no re-deploy required
  • Fallback to the deployment provider when a credential is missing
Everywhere the model is called

Coverage Across the Answer Path

  • Streaming widget and portal answers — on your credential
  • Multi-knowledge-base answers — on your credential
  • Slack / Discord / Teams bots (AnswerForBot) — on your credential
  • Form deflector, including grounding verification — on your credential
  • Reranking on the LLM grader fallback — on your credential
Honest about the gaps

What Still Runs on Deployment Provider

  • Internal diagnostic and eval tooling (deployment models, so scores are comparable across customers)
  • Conversation title generation (small fixed-cost call)
  • Image-to-text (DescribeImage) for chat attachments
  • Query planning (SearchService's own provider today)
  • Hosted cross-encoder rerankers (Cohere/Voyage/Jina — not LLM providers)
Why BYOK exists

The two questions Enterprise buyers actually ask

'Whose account is charged for the tokens?' and 'Whose DPA covers the data?' — the two questions every enterprise procurement conversation lands on. BYOK answers both with 'yours'. Traffic runs on your OpenAI / Anthropic / Gemini account under the terms you already negotiated. BeforeQuery is the platform above the model, not the LLM vendor.

  • Volume LLM contracts you already negotiated apply directly
  • Model choice is yours — no vendor lock-in via the platform
  • Data-processing terms are yours — no additional third-party review
  • Rotate the credential without re-deploying anything
  • Fallback to deployment provider keeps the workspace up if the credential fails
What teams say

Deployed in production, cited by the buyers who chose it

Deployed on our docs site in an afternoon. Every answer has citations, and the abstention gate means we've never had a customer complain about a made-up answer.
PN
Priya Nair
Head of Customer Support · Supabase
The knowledge base connected to our Slack, Confluence, and helpdesk in one setup. On-call teams get the same cited answer whether they ask in chat, in the widget, or from Cursor.
TR
Tom Richter
IT Operations Manager · Grafana Labs
The gap analytics turned into a real docs backlog. Deflection went up because we finally knew which pages were missing — the AI told us.
AC
Ana Castillo
VP of Customer Experience · Clerk

Frequently Asked Questions

Common questions about Bring Your Own Model (BYOK)

OpenAI, Anthropic, and Google Gemini for the chat model. Cross-encoder rerankers (Cohere, Voyage, Jina) are supported separately as hosted services — they're not LLMs, so BYOK for them is 'use your hosted-reranker key' rather than a full model swap.
Yes. Per-workspace credential is the default; per-knowledge-base override is available for the case where you want, for example, Claude Sonnet for a customer-facing KB and GPT-4o for internal ops. The router picks the closest override at call time.
The call fails and the user sees a friendly 'try again' message. There's a configurable fallback to the deployment provider so the workspace stays up during transient credential issues — you can turn it off if strict single-provider is a requirement.
The key is encrypted at rest with a per-workspace envelope key. Only the request path uses it, and only for the calls that route through the workspace's credential. It's not visible to other workspaces and not used for training or logging beyond the invocation record.
BeforeQuery is priced for the platform (retrieval, connectors, evals, analytics, surfaces). BYOK removes the LLM inference from our cost basis, so Enterprise deals with BYOK typically settle on a lower platform fee. Non-BYOK plans (Free, Pro) run on the deployment provider and include LLM inference in the plan price.
No — BYOK is an Enterprise feature today. Free and Pro run on the deployment provider so the pricing stays flat and self-serve. Enterprise adds BYOK, SSO, DPA, regional residency, and volume terms.

Run the platform, keep the LLM contract

Talk to us about Enterprise — BYOK, DPA, regional residency, and SSO all move together on Enterprise plans.

Get Started Free