Ad: BlueJ Better Tax Answers. -Accomplish hours of research in seconds -Instantly draft high-quality communications -Verify answers using a library of trusted tax content. Learn more

LLMs Provide Unstable Answers To Legal Questions

Andrew Blair-Stanek (Maryland; Google Scholar) & Benjamin Van Durme (Johns Hopkins; Google Scholar), LLMs Provide Unstable Answers to Legal Questions

An LLM is stable if it reaches the same conclusion when asked the identical question multiple times. We find leading LLMs like gpt-4o, claude-3.5, and gemini-1.5 are unstable when providing answers to hard legal questions, even when made as deterministic as possible by setting temperature to 0. We curate and release a novel dataset of 500 legal questions distilled from real cases, involving two parties, with facts, competing legal arguments, and the question of which party should prevail. 

When provided the exact same question, we observe that LLMs sometimes say one party should win, while other times saying the other party should win. This instability has implications for the increasing numbers of legal AI products, legal processes, and lawyers relying on these LLMs.

Editor's Note:  If you would like to receive a daily email with links to legal education posts on TaxProf Blog, email me here.


About the Author

Ad: BlueJ Better Tax Answers. Blue J's generative AI tax research solution is transforming how tax experts work. Learn more.
Information and rates on advertising on TaxProf Blog

Discover more from TaxProf Blog

Subscribe now to keep reading and get access to the full archive.

Continue reading