Legal playbook · AI Employee: Lex

eDiscovery Collection

Collections defensible and reproducible

The problem

eDiscovery collection is the specialized cousin of legal-hold execution. Once litigation-hold is in place, the actual data collection for review + production begins: pulling emails, chat, files, cloud services, endpoint data for named custodians and topics, preserving chain of custody, deduplicating, indexing for review. Manual coordination with IT + custodians takes weeks; over-collection produces terabytes of irrelevant data; under-collection risks spoliation sanctions.

At a glance
Trigger
Form
Approvals
Attorney sign-off per matter
What it does
Writes to your systems
Systems
Google Vault · M365 Purview · S3
How it feels in production

An hour-by-hour walkthrough.

Legal case CASE-2026-08-14 requires eDiscovery collection for 12 custodians on topic 'Project Titan' between 2024-06 and 2026-01. Legal-hold is already in place (via legal-hold-litigation-hold playbook). Now collection begins. Lex orchestrates collection: **Source enumeration per custodian** - Email: Google Vault / Microsoft Purview - Chat: Slack Enterprise Grid, Microsoft Teams (with attachments) - Files: Google Drive, Box, Dropbox, personal folders on OneDrive - Cloud services: any custodian-owned data in GitHub, Jira, Confluence, Notion - Endpoint: OS-level file collection for local documents **Scoping per topic + timeframe** - Keyword search: Titan + related codenames + participant names - Topic modeling: broader semantic match beyond keywords - Time-window: 2024-06 to 2026-01, extending 30 days on each side for context **Collection execution** - Chain-of-custody logging: every extraction timestamped + hashed + attributed to Lex - Preservation format: original files preserved; parallel copies exported for review - Deduplication: exact + near-duplicates identified for efficient review - Indexing: full-text index + metadata index for review platform ingestion **Handoff to review** - Collection package delivered to Relativity / Everlaw with metadata + index - Review workflow initiated with reviewers assigned - Progress tracked through review + production Collection completes in days instead of weeks. Chain-of-custody is airtight. Over-collection minimized; under-collection risk contained.
How it works

Step by step.

  1. 01

    Enumerate sources per custodian

    Every system custodian could have data in: email, chat, files, cloud services, endpoint. Comprehensive source list per custodian.

    Google Vault · Purview · Slack · Files · Cloud services
  2. 02

    Scope by topic + timeframe with keyword + semantic search

    Not just keyword matching. Topic modeling + semantic search for broader relevant content. Time-window with context extension.

    Search engines · Topic modeling · Semantic search
  3. 03

    Execute collection with chain of custody

    Every extraction timestamped + hashed + attributed. Original preserved; parallel copies exported for review.

    eDiscovery tools · Chain-of-custody logging
  4. 04

    Deduplicate + index for review

    Exact + near-duplicates identified. Full-text + metadata indexes prepared for review platform.

    Deduplication engine · Indexing
  5. 05

    Handoff to review workflow

    Collection delivered to Relativity / Everlaw / dedicated review platform. Review workflow initiated + tracked.

    Review platform · Review workflow
Systems and wiring

What you connect to make this run.

Google Vault · Microsoft Purview · Slack Enterprise · Box

read

Native enterprise-grade preservation + export. Chain-of-custody native to these platforms.

eDiscovery tools · Relativity · Everlaw · Reveal · Logikcull

read+write

Review platform ingestion + workflow management. Reviewer assignment + progress tracking.

Chain-of-custody logging

write

Immutable log of every extraction: what, when, from where, by whom (Lex), hash + timestamp. Court-referenceable.

Deduplication + indexing

read+write

Deduplication reduces review volume 40-70%. Indexing enables efficient reviewer workflow.

What changes

Before and after, honestly.

Time from collection request to review-ready
Before
3-8 weeks
After
3-7 business days
Chain-of-custody defects (audit findings)
Before
5-15% of cases
After
Zero
Legal + IT hours per collection
Before
80-200 hours
After
15-40 hours
Over-collection rate (data collected + reviewed but irrelevant)
Before
60-80%
After
20-35%
Frequently asked

Answers about this playbook.

What about collection across international jurisdictions (GDPR concerns)?

Jurisdictional scope respected. EU-hosted data collected per GDPR requirements; cross-border transfer only with appropriate legal basis.

How does it handle encrypted content (message-level encryption, encrypted attachments)?

Encryption boundaries respected. Content encrypted with corporate keys collectible; user-encrypted content requires user cooperation or court-ordered decryption.

What about privileged communications (attorney-client)?

Privilege detection during review. Suspected-privileged material flagged for legal review before production. Privilege log maintained.

How does it handle very large collections (millions of documents)?

Distributed collection + processing. Scales to petabyte-class collections with proportional review workflow. TAR (Technology Assisted Review) supported for efficient culling.

What about cloud-native data (Salesforce records, HubSpot data, custom apps)?

SaaS-specific connectors. Salesforce, HubSpot, Zendesk, custom apps all supported. Custody + preservation per each app's capabilities.

See it run on your data.

Free plan, no credit card. Connect the systems this playbook needs and run it against a past event first.