Fetching the paper…
Reading the bibliography…
The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning.
Ai in law practice? so far, not much
Anja Oskamp and Marc Lauritsen · 2002
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Legal reasoning
Phoebe C Ellsworth · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
Precedent and Analogy in Legal Reasoning
Grant Lamond · 2006
Earlier work this paper cites.
Oral arguments before the Supreme Court: An empirical approach
Lawrence Wrightsman · 2008
Earlier work this paper cites.
The lobbying manual: a complete guide to federal lobbying law and practice
William V Luneburg and Thomas M Susman · 2009
Earlier work this paper cites.
The litigation state
Sean Farhang · 2010
Earlier work this paper cites.
The no-reading problem in consumer contract law
Ian Ayres and Alan Schwartz · 2014
Earlier work this paper cites.
Does anyone read the fine print? consumer attention to standard-form contracts
Yannis Bakos, Florencia Marotta-Wurgler, and David R Trossen · 2014
Earlier work this paper cites.
From policy confusion to doctrinal clarity: successor liability from the perspective of big data
Frank Fagan · 2014
Earlier work this paper cites.
Campaign finance and american democracy
Yasmin Dawood · 2015
Earlier work this paper cites.
The creation and analysis of a website privacy policy corpus
Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N Cameron Russell, et al · 2016
Earlier work this paper cites.
Artificial intelligence and legal analytics: new tools for law practice in the digital age
Kevin D Ashley · 2017
Earlier work this paper cites.
The limitations of supply chain disclosure regimes
Adam S Chilton and Galit A Sarfaty · 2017
Earlier work this paper cites.
The Justice Gap: Measuring the Unmet Civil Legal Needs of Low-Income Americans, 2017
Legal Services Corporation · 2017
Earlier work this paper cites.
A logic for statutes
Sarah B Lawsky · 2017
Earlier work this paper cites.
Learning to predict charges for criminal cases with legal basis
Bingfeng Luo, Yansong Feng, Jianbo Xu, Xiang Zhang, and Dongyan Zhao · 2017
Earlier work this paper cites.
Predicting and understanding law-making with word vectors and an ensemble model
John J Nay · 2017
Earlier work this paper cites.
How ai can improve access to justice, 2017
Joel Tito · 2017
Earlier work this paper cites.
Opportunities and Obstacles for Deep Learning in Biology and Medicine
Travers Ching, Daniel S Himmelstein, Brett K Beaulieu-Jones, Alexandr A Kalinin, Brian T Do, Gregory P Way, Enrico Ferrero, Paul-Michael Agapow, Michael Zietz, Michael M Hoffman, et al · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A computational analysis of oral argument in the supreme court
Gregory M Dickinson · 2018
Earlier work this paper cites.
Nlp based latent semantic analysis for legal text summarization
Kaiz Merchant and Yash Pande · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Cail2018: A large-scale legal dataset for judgment prediction
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xianpei Han, Zhen Hu, Heng Wang, et al · 2018
Earlier work this paper cites.
Legal judgment prediction via topological learning
Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Chaojun Xiao, Zhiyuan Liu, and Maosong Sun · 2018
Earlier work this paper cites.
Neural legal judgment prediction in english
Ilias Chalkidis, Ion Androutsopoulos, and Nikolaos Aletras · 2019
Earlier work this paper cites.
Deep learning in law: early adaptation and legal word embeddings trained on large corpora
Ilias Chalkidis and Dimitrios Kampas · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
Text summarization from legal documents: a survey
Ambedkar Kanapala, Sukomal Pal, and Rajendra Pamula · 2019
Earlier work this paper cites.
Claudette: an automated detector of potentially unfair clauses in online terms of service
Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
Question answering for privacy policies: Combining computational and legal perspectives
Abhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, and Norman Sadeh · 2019
Earlier work this paper cites.
Breaking news: Drafting client alerts to prepare for practice
Cecilia Silver · 2019
Earlier work this paper cites.
Artificial intelligence and law: An overview
Harry Surden · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Earlier work this paper cites.
Maps: Scaling privacy compliance analysis to a million apps
Sebastian Zimmeck, Peter Story, Daniel Smullen, Abhilasha Ravichander, Ziqi Wang, Joel R Reidenberg, N Cameron Russell, and Norman Sadeh · 2019
Earlier work this paper cites.
Language models Are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Legal-bert: The muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos · 2020
Earlier work this paper cites.
Legal Tech, Civil Procedure, and the Future of Adversarialism
David Freeman Engstrom and Jonah B Gelbach · 2020
Earlier work this paper cites.
Pandemics and force majeure: How can ai help you?
Epiq · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
A dataset for statutory reasoning in tax law entailment and question answering
Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme · 2020
Earlier work this paper cites.
The biggest lie on the internet: Ignoring the privacy policies and terms of service policies of social networking services
Jonathan A Obar and Anne Oeldorf-Hirsch · 2020
Earlier work this paper cites.
Bootleg: Chasing the tail with self-supervised named entity disambiguation
Laurel Orr, Megan Leszczynski, Simran Arora, Sen Wu, Neel Guha, Xiao Ling, and Christopher Re · 2020
Earlier work this paper cites.
The ethics of artificial intelligence in law: Basic questions
Harry Surden · 2020
Earlier work this paper cites.
Raft: A real-world few-shot text classification benchmark
Neel Alex, Eli Lifland, Lewis Tunstall, Abhishek Thakur, Pegah Maham, C Jess Riedel, Emmie Hine, Carolyn Ashurst, Paul Sedille, Alexis Carlier, et al · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Eviction laws database: Local dataset
Legal Services Corporation · 2021
Cited alongside, same era.
Modeling law search as prediction
Faraz Dadgostari, Mauricio Guim, Peter A Beling, Michael A Livermore, and Daniel N Rockmore · 2021
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Cited alongside, same era.
Legal transformer models may not always help
Saibo Geng, Rémi Lebret, and Karl Aberer · 2021
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Legal prompting: Teaching a language model to think like a lawyer
Fangyi Yu, Lee Quartey, and Frank Schilder · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Falcon-40B: an open large language model with state-of-the-art performance
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Closest in time.
Introducing claude
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball · 2021
Cited alongside, same era.
Contractnli: A dataset for document-level natural language inference for contracts
Yuta Koreeda and Christopher D Manning · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Semantic segmentation of legal documents via rhetorical roles
Vijit Malik, Rishabh Sanjay, Shouvik Kumar Guha, Angshuman Hazarika, Shubham Nigam, Arnab Bhattacharya, and Ashutosh Modi · 2021
Cited alongside, same era.
Ildc for cjpe: Indian legal documents corpus for court judgment prediction and explanation
Vijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripa Ghosh, Shouvik Kumar Guha, Arnab Bhattacharya, and Ashutosh Modi · 2021
Cited alongside, same era.
Swiss-judgment-prediction: A multilingual legal judgment prediction benchmark
Joel Niklaus, Ilias Chalkidis, and Matthias Stürmer · 2021
Cited alongside, same era.
Multi-granular legal topic classification on greek legislation
Christos Papaloukas, Ilias Chalkidis, Konstantinos Athinaios, Despina-Athanasia Pantazi, and Manolis Koubarakis · 2021
Cited alongside, same era.
Anthropic · 2023
Closest in time.
How smart are smart readers? llms and the future of the no-reading problem
Yonathan A Arbel and Samuel Becher · 2023
Closest in time.
Open llm leaderboard
Edward Beeching, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf · 2023
Closest in time.
Can gpt-3 perform statutory reasoning?
Andrew Blair-Stanek, Nils Holzenberger, and Benjamin Van Durme · 2023
Closest in time.
2023 state-by-state ai legislation snapshot
Bryan, Cave, Leighton, and Paisner · 2023
Closest in time.
Chatgpt may pass the bar exam soon, but has a long way to go for the lexglue benchmark
Ilias Chalkidis · 2023
Closest in time.
Lexfiles and legallama: Facilitating english multinational legal language model development
Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz, and Anders Søgaard · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
How to use large language models for empirical legal research
Jonathan H. Choi · 2023
Closest in time.
Chatgpt goes to law school
Jonathan H Choi, Kristin E Hickman, Amy Monahan, and Daniel Schwarcz · 2023
Closest in time.
Ai assistance in legal analysis: An empirical study
Jonathan H. Choi and Daniel Schwarcz · 2023
Closest in time.
Redpajama: An open source recipe to reproduce llama training dataset, April 2023
Together Computer · 2023
Closest in time.
Simple hardware-efficient long convolutions for sequence modeling
Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré · 2023
Closest in time.
Embroid: Unsupervised prediction smoothing can improve few-shot classification
Neel Guha, Mayee F Chen, Kush Bhatia, Azalia Mirhoseini, Frederic Sala, and Christopher Ré · 2023
Closest in time.
Generative interpretation
David Hoffman and Yonathan Arbel · 2023
Closest in time.
Defeating the empire of forms
David A Hoffman · 2023
Closest in time.
Legal syllogism prompting: Teaching large language models for legal judgment prediction
Cong Jiang and Xiaolei Yang · 2023
Closest in time.
U-creat: Unsupervised case retrieval using events extraction
Abhinav Joshi, Akshat Sharma, Sai Kiran Tanikella, and Ashutosh Modi · 2023
Closest in time.
Gpt-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2023
Closest in time.
Natural language processing in the legal domain
Daniel Martin Katz, Dirk Hartung, Lauritz Gerlach, Abhik Jana, and Michael J Bommarito II · 2023
Closest in time.
Chain of reference prompting helps llm to think like a lawyer
Aditya Kuppa, Nikon Rasumov-Rahe, and Marc Voses · 2023
Closest in time.
Applying large language models for enhancing contract drafting
Kwok-Yan Lam, Victor CW Cheng, and Zee Kin Yeong · 2023
Closest in time.
Don’t use a cannon to kill a fly: An efficient cascading pipeline for long documents
Zehua Li, Neel Guha, and Julian Nyarko · 2023
Closest in time.
Rethinking the field of automatic prediction of court decisions
Masha Medvedeva, Martijn Wieling, and Michel Vols · 2023
Closest in time.
Large language models as tax attorneys: A case study in legal capabilities emergence, 2023
John J. Nay, David Karamardian, Sarah B. Lawsky, Wenting Tao, Meghana Bhat, Raghav Jain, Aaron Travis Lee, Jonathan H. Choi, and Jungo Kasai · 2023
Closest in time.
Lextreme: A multi-lingual and multi-task benchmark for the legal domain
Joel Niklaus, Veton Matoshi, Pooja Rani, Andrea Galassi, Matthias Stürmer, and Ilias Chalkidis · 2023
Closest in time.
Multilegalpile: A 689gb multilingual legal corpus
Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis, and Daniel E Ho · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Closest in time.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Closest in time.
Scale: Scaling up the complexity for advanced language model evaluation
Vishvaksenan Rasiah, Ronja Stern, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis, Daniel E Ho, and Joel Niklaus · 2023
Closest in time.
Street: A multi-task structured reasoning and explanation benchmark
Danilo Ribeiro, Shen Wang, Xiaofei Ma, Henry Zhu, Rui Dong, Deguang Kong, Juliette Burger, Anjelica Ramos, William Wang, Zhiheng Huang, et al · 2023
Closest in time.
No, ruth bader ginsburg did not dissent in obergefell — and other things chatgpt gets wrong about the supreme court
James Romoser · 2023
Closest in time.
Jaromir Savelka · 2023
Closest in time.
Can gpt-4 support analysis of textual data in tasks requiring highly specialized domain expertise?
Jaromir Savelka, Kevin D Ashley, Morgan A Gray, Hannes Westermann, and Huihui Xu · 2023
Closest in time.
Introducing mpt-30b: Raising the bar for open-source foundation models, 2023
MosaicML NLP Team · 2023
Closest in time.
Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned & chat models
Together · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Chatgpt coming to court, by way of self-represented litigants
Eugene Volokh · 2023
Closest in time.
Maud: An expert-annotated legal nlp dataset for merger agreement understanding, 2023
Steven H. Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks · 2023
Closest in time.
Xinyi Wang, Wanrong Zhu, and William Yang Wang · 2023
Closest in time.
Here’s what happens when your lawyer uses chatgpt
Benjamin Weiser · 2023
Closest in time.
Llmediator: Gpt-4 assisted online dispute resolution
Hannes Westermann, Jaromir Savelka, and Karim Benyekhlef · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions, 2023
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Closest in time.
Exploring the effectiveness of prompt engineering for legal reasoning tasks
Fangyi Yu, Lee Quartey, and Frank Schilder · 2023
Closest in time.
Private enforcement in the states
Diego Zambrano, Neel Guha, Austin Peters, and Jeffrey Xia · 2023
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang · 2023
Closest in time.
The robots are coming: Ai large language models and the legal profession
Lee B. Ziffer · 2023
Closest in time.