E-Discovery: The Proportionality Argument Is the One Worth Winning

The proportionality argument is the one worth winning

Most writing about e-discovery technology is about review speed. The larger lever is upstream of review: what you agree to produce and what you successfully resist producing. A firm that halves its review speed and doubles its collection scope has gone backwards.

For small and mid-sized practices in particular, the discovery fight that matters is proportionality — and it is argued with facts about volume, cost and burden that you have to be able to produce on demand.

What you need to be able to say, with numbers

When resisting a request as disproportionate, or negotiating custodians and date ranges, the assertions that carry weight are quantitative:

  • How many documents fall within the proposed scope, by custodian and by date range.
  • What the volume becomes under your narrower proposal.
  • What review will cost at each scope, on a defensible per-document basis.
  • What proportion of a sample is responsive — the strongest single argument that a broad request is a fishing expedition.

A party who can produce those numbers in a meet-and-confer is negotiating. A party who cannot is agreeing.

This is the most valuable and least discussed use of discovery tooling: not reviewing faster, but knowing the shape of the corpus early enough to argue about it.

Where machine assistance is well established

Deduplication and threading. The single largest reduction in volume, and uncontroversial. Email threads reviewed as threads rather than as individual messages remove a great deal of duplicated reading.

Clustering by concept to find the pockets worth human attention, and the large regions that are plainly irrelevant.

Privilege screening as a first pass. Names of counsel, patterns of correspondence, and terms of art surface candidates. Every one still needs human confirmation — a privilege call is not delegable — but the candidate list is mechanical.

Sampling to estimate richness. A statistically sensible sample of a proposed scope tells you the responsiveness rate, which is the number the proportionality argument runs on.

Where the risk sits

Privilege. An inadvertent production is the discovery failure with the most serious consequences. Clawback agreements help and do not undo what has been read. Any automated privilege step should be treated as a filter that narrows human review, never as a substitute for it.

Defensibility of the process. If you use technology-assisted review, you should be able to describe the method — how the classifier was trained, how the seed set was chosen, how validation was performed — because you may be asked. A process you cannot describe is a process you cannot defend.

Scope creep through convenience. Collecting broadly because collection is cheap creates review cost and privilege risk downstream. The cheapest document to review is the one never collected.

Confidentiality of the corpus itself. Where material is under a protective order, contains health data, or is subject to privilege, the question of which systems it passes through is a decision to make at the outset.

A workable sequence for a small firm

  1. Map the sources before collecting. Custodians, systems, date ranges, and what is realistically retrievable. Guessing here produces both over-collection and missed evidence.
  2. Sample early. Before agreeing scope, know the responsiveness rate. This is the fact that wins or loses the proportionality argument.
  3. Deduplicate and thread before any human reads anything.
  4. Screen for privilege, then have a human confirm every candidate.
  5. Review in priority order, richest clusters first, so that the useful material surfaces before the budget is spent.
  6. Keep a record of the method as you go. Reconstructed afterwards, it is an argument; recorded contemporaneously, it is documentation.

The question to ask a vendor

The same one that applies everywhere in this field: when the system tells you something, how long does it take to check?

For discovery specifically, that means: can you see why a document was classified as responsive or privileged, can you get from an assertion to the underlying document immediately, and can you export the process detail if you are asked to defend it. A system whose reasoning is opaque may still be useful for clustering, and should not be relied on for a privilege call.

The connection to the rest of the case

Discovery does not end at production. The documents that matter reappear in depositions, and the testimony about them is where they become evidence. Keeping documents and testimony addressable in the same way — so that an exhibit can be followed from production through every witness who was handed it — removes a class of work that is otherwise done by memory.

That is the design behind Lawnova PDF: documents and transcripts indexed together, every answer carrying a document, page and line reference you can open, and a local airgapped option for material that cannot leave your environment. For the testimony half of this, see deposition analysis is a testimony problem.

Lawnova Academy — AI certificates for legal professionals · Levels I–IV · 100% online Explore the Academy →