Long read, August 3, 2025
how review quality is actually measured in the wild
The problem
Every review team we work with has a version of this problem. How review quality is actually measured in the wild rarely announces itself.
What makes this hard is that the signal is distributed. A single clause read in isolation looks fine. Read against the definitions section and two schedules, the same clause does something different.
What we do about it
We score against a human baseline drawn from executed matters rather than a synthetic set, because synthetic contracts do not contain the drafting scars that cause real misses.
In the current build this runs as part of the standard pass, so it applies to every document in the set rather than only the ones someone thought to check.
Where it breaks
Scanned originals with poor image quality remain the weakest input. So do agreements that were assembled from three precedents and never reconciled, which is common in long-lived supplier relationships.
None of this removes the reviewer. It moves the reviewer to the part of the work where judgement actually pays.
Working notes from the Lawrs team. General information about legal technology and practice, not legal advice.