Protected health information filter (philter): accurately and securely de-identifying free-text clinical notes
Beau Norgeot, Kathleen Muenzen, Thomas A Peterson, Xuancheng Fan, Benjamin S Glicksberg, Gundolf Schenk, Eugenia Rutenberg, Boris Oskotsky, Marina Sirota, Jinoos Yazdany, et al · 2020
Later among the works it cites.
Echr: legal corpus for argument mining
Prakash Poudyal, Jaromír Šavelka, Aagje Ieven, Marie Francine Moens, Teresa Gonçalves, and Paulo Quaresma · 2020
Later among the works it cites.
The ethical dilemmas of the office of legal counsel in the wake of a whistleblower complaint
Abigail M Reecer · 2020
Later among the works it cites.
Information Leakage in Embedding Models , page 377–390
Congzheng Song and Ananth Raghunathan · 2020
Later among the works it cites.
Availability of state voter file and confidential information, 2020
U.S. Election Assistance Commission · 2020
Later among the works it cites.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave · 2020
Later among the works it cites.
Exploring the Limits of Large Scale Pre-training
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi · 2021
Later among the works it cites.
Abusive language detection in youtube comments leveraging replies as conversational context
Noman Ashraf, Arkaitz Zubiaga, and Alexander Gelbukh · 2021
Later among the works it cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
LexNLP: Natural language processing and information extraction for legal and regulatory texts
Michael J Bommarito II, Daniel Martin Katz, and Eric M Detterman · 2021
Later among the works it cites.
Code of Professional Conduct for British Columbia, 2021
British Columbia Law Society · 2021
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Later among the works it cites.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Later among the works it cites.
Racial epithets and racial etiquette
Richard Thompson Ford · 2021
Later among the works it cites.
The Law of Large Documents: Understanding the Structure of Legal Contracts Using Visual Cues
Original
Allison Hegel, Marina Shah, Genevieve Peaslee, Brendan Roof, and Emad Elwany · 2021
Later among the works it cites.
CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
Original
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball · 2021
Later among the works it cites.
Context-aware legal citation recommendation using deep learning
Zihan Huang, Charles Low, Mengqiu Teng, Hongyi Zhang, Daniel E Ho, Mark S Krass, and Matthias Grabmair · 2021
Later among the works it cites.
Perspective, 2021
Jigsaw and Google’s Counter Abuse Technology team · 2021
Later among the works it cites.
The new taboo: Quoting epithets in the classroom and beyond
Randall Kennedy and Eugene Volokh · 2021
Later among the works it cites.
ADePT: Auto-encoder based Differentially Private Text Transformation
Satyapriya Krishna, Rahul Gupta, and Christophe Dupuy · 2021
Later among the works it cites.
Capturing covertly toxic speech via crowdsourcing
Alyssa Lees, Daniel Borkan, Ian Kivlichan, Jorge Nario, and Tesh Goyal · 2021
Later among the works it cites.
Detecting and explaining unfairness in consumer contracts through memory networks
Federico Ruggeri, Francesca Lagioia, Marco Lippi, and Paolo Torroni · 2021
Later among the works it cites.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Later among the works it cites.
Sherman Estate v. Donovan, 2021
Supreme Court of Canada · 2021
Later among the works it cites.
Pseudonymous litigation, 2021
Eugene Volokh · 2021
Later among the works it cites.
Towards generalisable hate speech detection: a review on obstacles and solutions
Wenjie Yin and Arkaitz Zubiaga · 2021
Later among the works it cites.
Differential privacy for text analytics via natural text sanitization
Xiang Yue, Minxin Du, Tianhao Wang, Yaliang Li, Huan Sun, and Sherman SM Chow · 2021
Later among the works it cites.
When does pretraining help? assessing self-supervised learning for law and the CaseHOLD dataset of 53,000+ legal holdings
Lucia Zheng, Neel Guha, Brandon R Anderson, Peter Henderson, and Daniel E Ho · 2021
Later among the works it cites.
What Does it Mean for a Language Model to Preserve Privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Closest in time.
Quantifying memorization across neural language models
Original
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Closest in time.
Lexglue: A benchmark dataset for legal language understanding in english
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, and Nikolaos Aletras · 2022
Closest in time.
Whose language counts as high quality? Measuring language ideologies in text data selection
Original
Suchin Gururangan, Dallas Card, Sarah K Drier, Emily K Gade, Leroy Z Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A Smith · 2022
Closest in time.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Original
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Closest in time.
Data governance in the age of large-scale data-driven language technology
Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, Somaieh Nikpoor, Jörg Frohberg, Aaron Gokaslan, Peter Henderson, Rishi Bommasani, and Margaret Mitchell · 2022
Closest in time.
Deduplicating training data mitigates privacy risks in language models
Original
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Closest in time.
Rules of Professional Conduct, 2022
Law Society of Ontario · 2022
Closest in time.
Do language models plagiarize?
Original
Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee · 2022
Closest in time.
Microsoft presidio
Microsoft · 2022
Closest in time.
Lifting the curse of multilinguality by pre-training modular transformers
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe · 2022
Closest in time.
OPT: Open Pre-trained Transformer Language Models
Original
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Closest in time.
Executive Control of Agency Adjudication: Capacity, Selection and Precedential Rulemaking
David Hausman, Daniel E Ho, Mark Krass, and Anne M McDonough · 2024
Closest in time.