Connect your knowledge, keep it in sync
Add a URL, Confluence space, SharePoint site, or Notion workspace. BeforeQuery discovers every page, extracts clean content, and automatically re-indexes when content changes. No code, no webhooks, no manual uploads.
How the knowledge crawler works
Built for the real-world complexity of enterprise knowledge sources
Sitemap Discovery
BeforeQuery fetches sitemap.xml first, then follows links to discover every page the crawler should index. No manual URL lists required.
JavaScript Rendering
Uses a headless browser to render client-side content before extraction. SPAs, Next.js sites, and dynamically rendered intranet pages all index accurately.
Clean Markdown Extraction
Strips nav, footers, ads, and boilerplate. Preserves code blocks, tables, and heading hierarchy — the signal that matters for RAG retrieval.
Incremental Sync
Re-crawls on a schedule you control (hourly, daily, or on webhook trigger). Only re-indexes pages that changed, keeping costs and latency low.
Your index stays current automatically
Stale knowledge is worse than no knowledge — teams get wrong answers. BeforeQuery polls your sources on your chosen schedule and re-indexes only the pages that changed, keeping every answer grounded in the latest content.
Works with every knowledge platform
If it renders in a browser or exposes an API, BeforeQuery can index it
Frequently Asked Questions
Common questions about BeforeQuery knowledge source connectors
Index your first knowledge source in minutes
Connect a source, click sync, and your knowledge base is ready.
Get Started Free