← All projects

RESEARCH·BENCHMARK·2025

Benchmarking Humans & AI in Contract Drafting

Co-authored study scoring 13 AI tools and in-house lawyers on 30 contract-drafting tasks: 450 outputs, 72 survey responses, 12 interviews.

/ Overview

A study by Anna Guo, Arthur Souza Rodrigues, Mohamed Al Mamari, Sakshi Udeshi and Marc Astbury, published by Legal Benchmarks in September 2025 as preliminary findings.

“Is the AI good enough to draft this?” gets answered with vibes far more often than evidence. This study measures it: thirteen AI tools, seven built for lawyers and six general-purpose assistants, put up against in-house lawyers on the same thirty drafting tasks, plus 72 survey responses and 12 interviews with in-house legal leaders.

The report ranks the tools, but its more useful result is the comparison: where AI first drafts were as reliable as a lawyer’s, and where lawyers were still better.

/ What it found

  • Scores 13 AI tools and a group of in-house lawyers on 30 real drafting tasks, on reliability, usefulness and workflow support.
  • Draws on 450 task outputs, 72 survey responses and 12 interviews with in-house legal leaders.
  • Lawyers’ first drafts were reliable 56.7% of the time and the AI tools’ 57% on average, with the best tool at 73.3% and the best lawyer at 70%. Lawyers took about 13 minutes a task; the tools under a minute.

/ Phase 1

Phase 1 (April 2025), co-authored with Anna Guo, tested six AI tools on 18 data-extraction tasks submitted by in-house counsel, graded entirely by lawyers. General-purpose tools matched the legal tools on accuracy; the legal tools scored higher on usefulness.

/ Approach

Mixed-methods on purpose: the scores say what happened, the interviews say why it matters. The methodology and its limitations are set out in the report’s appendix; the task set and outputs have not been released. Reliability was graded by an LLM jury with expert review of borderline cases, and usefulness was scored blind by lawyers. My later work at benchmarks.law drops the machine judge.

The report is now titled “Benchmarking Humans & AI in Contract Workflows” on legalbenchmarks.ai.