Fetching the paper…
Reading the bibliography…
Annotation and classification of legal text are central components of empirical legal research.
U.S. Courts of Appeals databases 1925–1996
Donald R. Songer · 1996
Earlier work this paper cites.
Dynamic ideal point estimation via markov chain monte carlo for the us supreme court, 1953–1999
Andrew D Martin and Kevin M Quinn · 2002
Earlier work this paper cites.
The Supreme Court and the attitudinal model revisited
Jeffrey A Segal and Harold J Spaeth · 2002
Earlier work this paper cites.
Systematic content analysis of judicial opinions
Mark A Hall and Ronald F Wright · 2008
Earlier work this paper cites.
Routine use of microbial whole genome sequencing in diagnostic and public health microbiology
Claudio U Köser, Matthew J Ellington, Edward JP Cartwright, Stephen H Gillespie, Nicholas M Brown, Mark Farrington, Matthew TG Holden, Gordon Dougan, Stephen D Bentley, Julian Parkhill, et al · 2012
Earlier work this paper cites.
The behavior of federal judges: a theoretical and empirical study of rational choice
Lee Epstein, William M Landes, and Richard A Posner · 2013
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
The creation and analysis of a website privacy policy corpus
Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N Cameron Russell, et al · 2016
Earlier work this paper cites.
Utilizing field collected insects for next generation sequencing: Effects of sampling, storage, and dna extraction methods
Kimberly M Ballare, Nathaniel S Pope, Antonio R Castilla, Sarah Cusser, Richard P Metz, and Shalene Jha · 2019
Earlier work this paper cites.
Ghost work: How to stop Silicon Valley from building a new global underclass
Mary L Gray and Siddharth Suri · 2019
Earlier work this paper cites.
CLAUDETTE: an automated detector of potentially unfair clauses in online terms of service
Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni · 2019
Earlier work this paper cites.
Law as Data: Computation, Text, & the Future of Legal Analysis
Michael A. Livermore and Daniel N. Rockmore, editors · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Question answering for privacy policies: Combining computational and legal perspectives
Abhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, and Norman Sadeh · 2019
Earlier work this paper cites.
Maps: Scaling privacy compliance analysis to a million apps
Sebastian Zimmeck, Peter Story, Daniel Smullen, Abhilasha Ravichander, Ziqi Wang, Joel R Reidenberg, N Cameron Russell, and Norman Sadeh · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
LEGAL-BERT: the muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos · 2020
Earlier work this paper cites.
Text classification of ideological direction in judicial opinions
Carina I Hausladen, Marcel H Schubert, and Elliott Ash · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Cleaning corporate governance
Jens Frankenreiter, Cathy Hwang, Yaron Nili, and Eric Talley · 2021
Cited alongside, same era.
The pile: an 800gb dataset of diverse text for language modeling 2020
L Gao, S Biderman, S Black, L Golding, T Hoppe, C Foster, J Phang, H He, A Thite, N Nabeshima, et al · 2021
Cited alongside, same era.
Cuad: An expert-annotated nlp dataset for legal contract review
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball · 2021
Cited alongside, same era.
Factoring statutory reasoning as language understanding challenges
MAUD: An expert-annotated legal nlp dataset for merger agreement understanding
Steven H Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks · 2023
Later among the works it cites.
Evaluating ai for law: Bridging the gap with open-source solutions
Rohan Bhambhoria, Samuel Dahan, Jonathan Li, and Xiaodan Zhu · 2024
Closest in time.
Axolotl, 2024
Axolotl AI Cloud · 2024
Closest in time.
Training on the test task confounds evaluation and emergence
Ricardo Dominguez-Olmedo, Florian E Dorner, and Moritz Hardt · 2024
Closest in time.
Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget
Florian E Dorner and Moritz Hardt · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Holzenberger and Benjamin Van Durme · 2021
Cited alongside, same era.
ContractNLI: A dataset for document-level natural language inference for contracts
Yuta Koreeda and Christopher D Manning · 2021
Cited alongside, same era.
When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings
Lucia Zheng, Neel Guha, Brandon R Anderson, Peter Henderson, and Daniel E Ho · 2021
Cited alongside, same era.
Patterns, predictions, and actions: Foundations of machine learning
Moritz Hardt and Benjamin Recht · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Harvey, which uses AI to answer legal questions, lands cash from OpenAI
Kyle Wiggers · 2022
Cited alongside, same era.
How to use large language models for empirical legal research
Jonathan H Choi · 2023
Cited alongside, same era.
AI assistance in legal analysis: An empirical study
Jonathan H Choi and Daniel Schwarcz · 2023
Cited alongside, same era.
Asking GPT for the ordinary meaning of statutory terms
Christoph Engel and Richard H Mcadams · 2024
Closest in time.
Sticky charters? the surprisingly tepid embrace of officer-protecting waivers in delaware
Jens Frankenreiter and Eric L Talley · 2024
Closest in time.
A framework for few-shot language model evaluation, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2024
Closest in time.
Empirical legal analysis simplified: reducing complexity through automatic identification and evaluation of legally relevant factors
Morgan A Gray, Jaromir Savelka, Wesley M Oliver, and Kevin D Ashley · 2024
Closest in time.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Promises and pitfalls of artificial intelligence for legal applications
Sayash Kapoor, Peter Henderson, and Arvind Narayanan · 2024
Closest in time.
GPT-4 passes the Bar Exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2024
Closest in time.
Llama 3: Advancing open foundation models, 2024
MetaAI · 2024
Closest in time.
Large language models as tax attorneys: a case study in legal capabilities emergence
John J Nay, David Karamardian, Sarah B Lawsky, Wenting Tao, Meghana Bhat, Raghav Jain, Aaron Travis Lee, Jonathan H Choi, and Jungo Kasai · 2024
Closest in time.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Closest in time.
Smollm2: When smol goes big–data-centric training of a small language model
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, et al · 2025
Closest in time.
Claude 3.7 sonnet system card, 2025
Anthropic · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.