A new study by researchers at Stanford University argues that many artificial intelligence tools used in the legal field lack independent public benchmarking, making it difficult for lawyers, judges and the public to assess how reliably they perform.
The study comes as AI tools are increasingly being adopted for legal research, contract review, litigation support and other legal tasks. In a recent annual report on the judiciary, U.S. Chief Justice John Roberts cautioned that “any use of AI requires caution and humility.”
The paper, titled “There Is No Free Benchmark: An Institutional View of Legal AI Benchmarking,” published in a special issue of the Proceedings of the National Academy of Sciences (PNAS), argues that greater public benchmarking is needed to improve transparency around legal AI systems.
The study, led by Neel Guha, a recent Stanford graduate and now an associate professor at Columbia Law School and co-authored by Daniel E. Ho, Christopher D. Manning and Julian Nyarko , said that while AI has the potential to make legal services more affordable, efficient and accessible, there is limited public information about the performance and risks of widely used legal AI tools.
The researchers said previous studies have documented hallucinations, bias and uneven performance in legal AI systems, including more than 120 documented cases involving AI-generated fake legal citations. They added that other errors, such as misinterpreting legal arguments or citing sources incorrectly, can be harder to detect.
“Legal AI lacks what we call legibility,” Daniel E. Ho said. “We know surprisingly little about the performance of legal AI systems, the kinds of mistakes they make and the likelihood of error.” He added that the consequences can be severe, citing more than 1,700 legal cases involving AI-hallucinated facts, cases and laws.
Unlike previous studies that focused primarily on technical evaluations of legal AI, the researchers argue that effective public benchmarking is also an institutional challenge. They said its success depends on who conducts the evaluations, the incentives they face and the resources available.
“In other domains, institutional stakeholders have approached similar concerns through community-driven benchmarking processes, in which multidisciplinary expert groups assess the performance of AI systems and forecast their impact,” the researchers said.
The researchers pointed to benchmarking efforts in fields such as medicine, software engineering and facial recognition as examples of how community-driven evaluation frameworks have supported responsible AI adoption.
They said greater transparency around AI system performance would help lawyers, courts and policymakers better understand the capabilities, limitations and risks of legal AI systems, supporting more informed adoption, innovation and governance.
Also Read: As AI Reaches Courts: UNESCO Issues Global Guidance for Judiciary






