Legal AI’s Legibility Problem

Increasing reliance on language models for legal tasks have led to calls for more information about the likelihood and severity of potential errors. Realizing useful and robust transparency requires an institutional perspective, say the editors of a new special section in PNAS.
In a recent annual letter on the judiciary, Chief Justice John Roberts cautioned the legal profession against too great a reliance on AI. “Any use of AI requires caution and humility,” he wrote, adding that human judges would not be unseated anytime soon, if ever. The nod was remarkable both as an acknowledgment of AI’s rising presence in the judiciary and as a spotlight on its considerable shortcomings on the law.
“AI in the practice of law is no longer a question of if, but of how,” says Daniel E. Ho, a professor of law at Stanford Law School and senior fellow at the Stanford Institute for Human-Centered AI (HAI). “That ‘how’ is the harder question – how to harness AI without compromising the professional and ethical obligations at the heart of legal practice.”
Information Is Downstream of Institutions
Ho, Stanford professors Julian Nyarko and Chris Manning, and HAI Managing Director Vanessa Parli are the editors of a special edition of Proceedings of the National Academy of Sciences (PNAS) exploring the future of AI and the judiciary. The issue includes a paper led by Neel Guha, a recent Stanford CS PhD/JD graduate now a Columbia Law School associate professor, and coauthored by Ho, Nyarko, and Manning, called “There’s No Free Benchmark: An Institutional View of Legal AI Benchmarking.” Ho says the problem stems from how legal AI systems are currently tested and validated – a stage in the AI development process well known to computer scientists as “benchmarking.”
“Legal AI lacks what we call legibility,” Ho says. “We know surprisingly little about the performance of legal AI systems, the kinds of mistakes they make, and the likelihood of error. The consequences can be severe, with over 1,700 legal cases involving hallucinated facts, cases, and laws.”
What the field needs, the piece argues, is more attention to the who, what, and how of benchmarking. Guha describes this as an “institutional view” of benchmarking. “Benchmarking is a series of decisions about metrics, data, and methodological configurations. While there have been many discussions in computer science about the design of benchmarks, there has been little analysis of the underlying institutional dynamics in settings where benchmarks are expensive and data is confidential. At the end of the day, someone has to do the benchmarking, and that someone is operating under distinct incentives and limitations.”
Guha notes that an institutional perspective goes a long way in explaining why legal AI lacks legibility. “When we take stock of the diverse set of actors in the legal AI ecosystem – developers, firms, academics, and public oversight bodies – it starts to become apparent exactly why we have widespread challenges around AI transparency.”
Institutionally Aware Benchmarking
In the paper, the authors suggest that overcoming legal AI’s legibility problems will require serious consideration of institutional design. “The key,” Ho describes, “is to recognize the institutional constraints at the beginning and work from there to identify how the benchmarking process should be structured.”
The authors separate their recommendations based on the amount of resources available for commitment to a benchmarking effort – what they delineate as high-, medium-, and low-resource settings. For example, the high-resource settings might utilize public bodies that can engage in benchmarking in an independent, neutral, and expertise-informed way. Ho and coauthors note that one candidate body – the National Institute for Standards and Technology (NIST) – has already modeled this approach in the context of facial recognition. Conversely, low-resource settings call for more targeted approaches focused on settings where the marginal benefit of any benchmarking information is greatest. This might cover, for instance, areas of the law like bankruptcy or child custody, where individuals are most likely to turn to popular general consumer chatbots for advice.
The paper is just one among many in a Special Features issue in which PNAS tackles emerging or underrepresented areas of research through interdisciplinary perspectives. Stanford HAI, as a top interdisciplinary institute, was invited to put together this special feature. Topics include the collective licensing of copyrighted works for training, whether AI regulation should target deployers or developers, how fairness in machine learning should grapple with recent changes in equal protection doctrine, using legal interpretation to align AI systems with human values, and using machine learning to map federal common law as a network. All pieces were subject to rigorous peer review and require complete transparency on conflicts of interest.