Journalism begins where hype ends

,,

If you can't explain it to a six year old, you don't understand it yourself."

—Albert Einstein

Legal AI Lacks ‘Legibility,’ Stanford Researchers Call for Public Benchmarking

Beyond technical evaluations, the researchers argue that building trustworthy legal AI requires independent, community-driven benchmarking supported by the right institutions.
Close-up of a judge's gavel, balanced scales of justice and a legal document, symbolizing the use of artificial intelligence and benchmarking in the legal profession.
July 24, 2026 09:34 AM IST | Written by Supriya Singh | Edited by Pratima O Pareek

A new study by researchers at Stanford University argues that many artificial intelligence tools used in the legal field lack independent public benchmarking, making it difficult for lawyers, judges and the public to assess how reliably they perform.

The study comes as AI tools are increasingly being adopted for legal research, contract review, litigation support and other legal tasks. In a recent annual report on the judiciary, U.S. Chief Justice John Roberts cautioned that “any use of AI requires caution and humility.”

The paper, titled “There Is No Free Benchmark: An Institutional View of Legal AI Benchmarking,” published in a special issue of the Proceedings of the National Academy of Sciences (PNAS), argues that greater public benchmarking is needed to improve transparency around legal AI systems.

The study, led by Neel Guha, a recent Stanford graduate and now an associate professor at Columbia Law School and co-authored by Daniel E. Ho, Christopher D. Manning and Julian Nyarko , said that while AI has the potential to make legal services more affordable, efficient and accessible, there is limited public information about the performance and risks of widely used legal AI tools.

The researchers said previous studies have documented hallucinations, bias and uneven performance in legal AI systems, including more than 120 documented cases involving AI-generated fake legal citations. They added that other errors, such as misinterpreting legal arguments or citing sources incorrectly, can be harder to detect.

“Legal AI lacks what we call legibility,” Daniel E. Ho said. “We know surprisingly little about the performance of legal AI systems, the kinds of mistakes they make and the likelihood of error.” He added that the consequences can be severe, citing more than 1,700 legal cases involving AI-hallucinated facts, cases and laws.

Unlike previous studies that focused primarily on technical evaluations of legal AI, the researchers argue that effective public benchmarking is also an institutional challenge. They said its success depends on who conducts the evaluations, the incentives they face and the resources available.

“In other domains, institutional stakeholders have approached similar concerns through community-driven benchmarking processes, in which multidisciplinary expert groups assess the performance of AI systems and forecast their impact,” the researchers said.

The researchers pointed to benchmarking efforts in fields such as medicine, software engineering and facial recognition as examples of how community-driven evaluation frameworks have supported responsible AI adoption.

They said greater transparency around AI system performance would help lawyers, courts and policymakers better understand the capabilities, limitations and risks of legal AI systems, supporting more informed adoption, innovation and governance.

Also Read: As AI Reaches Courts: UNESCO Issues Global Guidance for Judiciary

Authors

  • AI FrontPage Reporter Supriya Singh

    Supriya Singh is a Reporter at AI FrontPage covering the AI & Education and AI & Jobs beats. She brings six years of print and digital experience, including three years at The Asian Age, where she reported on higher education, Delhi government, and crime. She is based in Delhi-NCR.

    LinkedIn

  • Pratima Pareek, Editor and Co-founder of AI FrontPage

    Pratima O Pareek is an Editor and Co-Founder of AI FrontPage. A gold medalist in Mass Communication and Journalism, she's worked across national and international newsrooms, bringing sharp editorial instincts and a commitment to clarity. She believes in cutting through the noise to deliver stories that actually matter.
    Off the clock, she watches offbeat cinema, follows tennis, and explores new places like a traveler, not a tourist.

    LinkedIn