Benchmark, July 23, 2026
Benchmark: obligation extraction on long-form outsourcing contracts
The problem
Every review team we work with has a version of this problem. Obligation extraction on long-form outsourcing contracts rarely announces itself.
What makes this hard is that the signal is distributed. A single clause read in isolation looks fine. Read against the definitions section and two schedules, the same clause does something different.
How we measured it
We score against a human baseline drawn from executed matters rather than a synthetic set, because synthetic contracts do not contain the drafting scars that cause real misses.
In the current build this runs as part of the standard pass, so it applies to every document in the set rather than only the ones someone thought to check.
Where it breaks
Scanned originals with poor image quality remain the weakest input. So do agreements that were assembled from three precedents and never reconciled, which is common in long-lived supplier relationships.
We will revisit this once we have a larger sample. The current numbers are directional rather than settled.
Working notes from the Lawrs team. General information about legal technology and practice, not legal advice.