Beyond technical evaluations, the researchers argue that building trustworthy legal AI requires independent, community-driven benchmarking supported by the right institutions.
Legal AI Lacks ‘Legibility,’ Stanford Researchers Call for Public Benchmarking