eDiscovery Collection
Collections defensible and reproducible
eDiscovery collection is the specialized cousin of legal-hold execution. Once litigation-hold is in place, the actual data collection for review + production begins: pulling emails, chat, files, cloud services, endpoint data for named custodians and topics, preserving chain of custody, deduplicating, indexing for review. Manual coordination with IT + custodians takes weeks; over-collection produces terabytes of irrelevant data; under-collection risks spoliation sanctions.
An hour-by-hour walkthrough.
Step by step.
- 01
Enumerate sources per custodian
Every system custodian could have data in: email, chat, files, cloud services, endpoint. Comprehensive source list per custodian.
Google Vault · Purview · Slack · Files · Cloud services - 02
Scope by topic + timeframe with keyword + semantic search
Not just keyword matching. Topic modeling + semantic search for broader relevant content. Time-window with context extension.
Search engines · Topic modeling · Semantic search - 03
Execute collection with chain of custody
Every extraction timestamped + hashed + attributed. Original preserved; parallel copies exported for review.
eDiscovery tools · Chain-of-custody logging - 04
Deduplicate + index for review
Exact + near-duplicates identified. Full-text + metadata indexes prepared for review platform.
Deduplication engine · Indexing - 05
Handoff to review workflow
Collection delivered to Relativity / Everlaw / dedicated review platform. Review workflow initiated + tracked.
Review platform · Review workflow
What you connect to make this run.
Google Vault · Microsoft Purview · Slack Enterprise · Box
readNative enterprise-grade preservation + export. Chain-of-custody native to these platforms.
eDiscovery tools · Relativity · Everlaw · Reveal · Logikcull
read+writeReview platform ingestion + workflow management. Reviewer assignment + progress tracking.
Chain-of-custody logging
writeImmutable log of every extraction: what, when, from where, by whom (Lex), hash + timestamp. Court-referenceable.
Deduplication + indexing
read+writeDeduplication reduces review volume 40-70%. Indexing enables efficient reviewer workflow.
Before and after, honestly.
Playbooks that pair with this one.
Answers about this playbook.
What about collection across international jurisdictions (GDPR concerns)?
Jurisdictional scope respected. EU-hosted data collected per GDPR requirements; cross-border transfer only with appropriate legal basis.
How does it handle encrypted content (message-level encryption, encrypted attachments)?
Encryption boundaries respected. Content encrypted with corporate keys collectible; user-encrypted content requires user cooperation or court-ordered decryption.
What about privileged communications (attorney-client)?
Privilege detection during review. Suspected-privileged material flagged for legal review before production. Privilege log maintained.
How does it handle very large collections (millions of documents)?
Distributed collection + processing. Scales to petabyte-class collections with proportional review workflow. TAR (Technology Assisted Review) supported for efficient culling.
What about cloud-native data (Salesforce records, HubSpot data, custom apps)?
SaaS-specific connectors. Salesforce, HubSpot, Zendesk, custom apps all supported. Custody + preservation per each app's capabilities.
See it run on your data.
Free plan, no credit card. Connect the systems this playbook needs and run it against a past event first.