flâneur — a map of the web's best reading

AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries

hai.stanford.edu · 2,156 words · saved by 1 readers

Artificial intelligence (AI) tools are rapidly transforming the practice of law. Nearly three quarters of lawyers plan on using generative AI for their work, from sifting through mountains of case law to drafting contracts to reviewing documents to writing legal memoranda. But are these tools reliable enough for real-world use? Large language models have a documented tendency to “hallucinate,” or make up false information. In one highly-publicized case, a New York lawyer faced sanctions for citing ChatGPT-invented fictional cases in a legal brief; many similar cases have since been reported. And our previous study of general-purpose chatbots found that they hallucinated between 58% and 82% of the time on legal queries, highlighting the risks of incorporating AI into legal practice. In his 2023 annual report on the judiciary, Chief Justice Roberts took note and warned lawyers of hallucinations. Across all areas of industry, retrieval-augmented generation (RAG) is seen and promoted as th

A new study reveals the need for benchmarking and public evaluations of AI tools in law. Artificial intelligence (AI) tools are rapidly transforming the practice of law. Nearly three quarters of lawyers plan on using generative AI for their work, from sifting through mountains of case law to drafting contracts to reviewing documents to writing legal memoranda. But are these tools reliable enough for real-world use? Large language models have a documented tendency to “ hallucinate ,” or make up false information. In one highly-publicized case, a New York lawyer faced sanctions for citing ChatGP

Explore this link on the map →

related reading