Fetching the paper…
Reading the bibliography…
Ensuring alignment, which refers to making models behave in accordance with human intentions [1,2], has become a critical task before deploying large language models (LLMs) in real-world applications.
An optimum character recognition system using decision functions
Chi-Keung Chow · 1957
Earlier work this paper cites.
The nearest neighbor classification rule with a reject option
Martin E Hellman · 1970
Earlier work this paper cites.
Calculation of the wasserstein distance between probability distributions on the line
SS Vallender · 1974
Earlier work this paper cites.
How to generate and exchange secrets
Andrew Chi-Chih Yao · 1986
Earlier work this paper cites.
Social choice theory
Amartya Sen · 1986
Earlier work this paper cites.
“just how much did that wheelchair cost?”: Management of privacy boundaries by persons with disabilities
Dawn O Braithwaite · 1991
Earlier work this paper cites.
Envy-freeness and distributive justice
Christian Arnsperger · 1994
Earlier work this paper cites.
A method for improving classification reliability of multilayer perceptrons
Luigi Pietro Cordella, Claudio De Stefano, Francesco Tortorella, and Mario Vento · 1995
Earlier work this paper cites.
The regulation of pornography and child pornography on the internet
Yaman Akdeniz · 1997
Earlier work this paper cites.
Misuse of the internet by pedophiles: Implications for law enforcement and probation practice
Keith F Durkin · 1997
Earlier work this paper cites.
Appraising the performance of employees with disabilities: A review and model
Adrienne Colella, Angelo S DeNisi, and Arup Varma · 1997
Earlier work this paper cites.
False memories and confabulation
Marcia K Johnson and Carol L Raye · 1998
Earlier work this paper cites.
The dark side of cyberspace: Internet content regulation and child protection
David Oswell · 1999
Earlier work this paper cites.
The quest for modern manhood: Masculine stereotypes, peer culture and the social significance of homophobia
David C Plummer · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Controversies and legal issues of prescribing and dispensing medications using the internet
Constance H Fung, Hawkin E Woo, and Steven M Asch · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Internet-based mental health interventions
Michele L Ybarra and William W Eaton · 2005
Earlier work this paper cites.
Usable security: Why do we need it? how do we get it?
M Angela Sasse and Ivan Flechais · 2005
Earlier work this paper cites.
Classification with reject option
Radu Herbei and Marten H Wegkamp · 2006
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
Differential privacy
Cynthia Dwork · 2006
Earlier work this paper cites.
Not gay enough for the government: Racial and sexual stereotypes in sexual orientation asylum cases
Deborah A Morgan · 2006
Earlier work this paper cites.
Age discrimination: An historical and contemporary analysis
John Macnicol · 2006
Earlier work this paper cites.
Can machine learning be secure?
Marco Barreno, Blaine Nelson, Russell Sears, Anthony D Joseph, and J Doug Tygar · 2006
Earlier work this paper cites.
Paragraph: Thwarting signature learning by training maliciously
James Newsome, Brad Karp, and Dawn Song · 2006
Earlier work this paper cites.
Online information, extreme communities and internet therapy: Is the internet good for our mental health?
Vaughan Bell · 2007
Earlier work this paper cites.
A short introduction to computational social choice
Yann Chevaleyre, Ulle Endriss, Jérôme Lang, and Nicolas Maudet · 2007
Earlier work this paper cites.
Fighting spam on social web sites: A survey of approaches and future challenges
Paul Heymann, Georgia Koutrika, and Hector Garcia-Molina · 2007
Earlier work this paper cites.
Review spam detection
Nitin Jindal and Bing Liu · 2007
Earlier work this paper cites.
Exploiting machine learning to subvert your spam filter
Blaine Nelson, Marco Barreno, Fuching Jack Chi, Anthony D Joseph, Benjamin IP Rubinstein, Udam Saini, Charles Sutton, J Doug Tygar, and Kai Xia · 2008
Earlier work this paper cites.
Casting out demons: Sanitizing training data for anomaly sensors
Gabriela F Cretu, Angelos Stavrou, Michael E Locasto, Salvatore J Stolfo, and Angelos D Keromytis · 2008
Earlier work this paper cites.
Child protection and self-regulation in the internet industry: The uk experience
John Carr and Zoë Hilton · 2009
Earlier work this paper cites.
Religious stereotyping and voter support for evangelical candidates
Monika L McDermott · 2009
Earlier work this paper cites.
Gay stereotypes: The use of sexual orientation as a cue for gender-related attributes
Aaron J Blashill and Kimberly K Powlishta · 2009
Earlier work this paper cites.
Building classifiers with independency constraints
Toon Calders, Faisal Kamiran, and Mykola Pechenizkiy · 2009
Earlier work this paper cites.
Privacy and security usable security: how to get it
Butler Lampson · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Antidote: understanding and defending against poisoning of anomaly detectors
Benjamin IP Rubinstein, Blaine Nelson, Ling Huang, Anthony D Joseph, Shing-hon Lau, Satish Rao, Nina Taft, and J Doug Tygar · 2009
Earlier work this paper cites.
Dynamics of hate based internet user networks
Pawel Sobkowicz and Antoni Sobkowicz · 2010
Earlier work this paper cites.
On the foundations of noise-free selective classification
Ran El-Yaniv et al · 2010
Earlier work this paper cites.
Effect of pathological use of the internet on adolescent mental health: a prospective study
Lawrence T Lam and Zi-Wen Peng · 2010
Earlier work this paper cites.
Of passwords and people: measuring the effect of password-composition policies
Saranga Komanduri, Richard Shay, Patrick Gage Kelley, Michelle L Mazurek, Lujo Bauer, Nicolas Christin, Lorrie Faith Cranor, and Serge Egelman · 2011
Earlier work this paper cites.
Adversarial machine learning
Ling Huang, Anthony D Joseph, Blaine Nelson, Benjamin IP Rubinstein, and J Doug Tygar · 2011
Earlier work this paper cites.
A review of internet pornography use research: Methodology and content from the past 10 years
Mary B Short, Lora Black, Angela H Smith, Chad T Wetterneck, and Daryl E Wells · 2012
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel · 2012
Earlier work this paper cites.
Gender stereotypes and workplace bias
Madeline E Heilman · 2012
Earlier work this paper cites.
A comparative study of cyberattacks
Seung Hyun Kim, Qiu-Hong Wang, and Johannes B Ullrich · 2012
Earlier work this paper cites.
Survey on web spam detection: principles and algorithms
Nikita Spirin and Jiawei Han · 2012
Earlier work this paper cites.
Predicting fair use
Matthew Sag · 2012
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
Stochastic gradient descent with differentially private updates
Shuang Song, Kamalika Chaudhuri, and Anand D Sarwate · 2013
Earlier work this paper cites.
Racial and ethnic stereotypes and bullying victimization
Anthony A Peguero and Lisa M Williams · 2013
Earlier work this paper cites.
Writing the wrong: Can counter-stereotypes offset negative media messages about african americans?
Lanier Frush Holt · 2013
Earlier work this paper cites.
Phishing detection: a literature survey
Mahmoud Khonji, Youssef Iraqi, and Andrew Jones · 2013
Earlier work this paper cites.
Regulating the internet of things: first steps toward managing discrimination, privacy, security and consent
Scott R Peppet · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Surveying the development of biometric user authentication on mobile phones
Weizhi Meng, Duncan S Wong, Steven Furnell, and Jianying Zhou · 2014
Earlier work this paper cites.
Abductive reasoning
Douglas Walton · 2014
Earlier work this paper cites.
Robust logistic regression and classification
Jiashi Feng, Huan Xu, Shie Mannor, and Shuicheng Yan · 2014
Earlier work this paper cites.
Recent advances in robust optimization: An overview
Virginie Gabrel, Cécile Murat, and Aurélie Thiele · 2014
Earlier work this paper cites.
Online peer-to-peer support for young people with mental health problems: a systematic review
Kathina Ali, Louise Farrer, Amelia Gulliver, Kathleen M Griffiths, et al · 2015
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart · 2015
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Advanced social engineering attacks
Katharina Krombholz, Heidelinde Hobel, Markus Huber, and Edgar Weippl · 2015
Earlier work this paper cites.
User authentication schemes for wireless sensor networks: A review
Saru Kumari, Muhammad Khurram Khan, and Mohammed Atiquzzaman · 2015
Earlier work this paper cites.
Survey of review spam detection using machine learning techniques
Michael Crawford, Taghi M Khoshgoftaar, Joseph D Prusa, Aaron N Richter, and Hamzah Al Najada · 2015
Earlier work this paper cites.
Deep learning, 2016
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Boosting with abstention
Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri · 2016
Earlier work this paper cites.
Towards the science of security and privacy in machine learning
Nicolas Papernot, Patrick McDaniel, Arunesh Sinha, and Michael Wellman · 2016
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Earlier work this paper cites.
Fairness in learning: Classic and contextual bandits
Matthew Joseph, Michael Kearns, Jamie H Morgenstern, and Aaron Roth · 2016
Earlier work this paper cites.
How do us christians and atheists stereotype one another’s moral values?
Ain Simpson and Kimberly Rios · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai · 2016
Earlier work this paper cites.
A literature survey on social engineering attacks: Phishing attack
Surbhi Gupta, Abhishek Singhal, and Akanksha Kapoor · 2016
Earlier work this paper cites.
The rise of social bots
Emilio Ferrara, Onur Varol, Clayton Davis, Filippo Menczer, and Alessandro Flammini · 2016
Earlier work this paper cites.
Continuous user authentication on mobile devices: Recent progress and remaining challenges
Vishal M Patel, Rama Chellappa, Deepak Chandra, and Brandon Barbello · 2016
Earlier work this paper cites.
You are not your developer, either: A research agenda for usable security and privacy research beyond end users
Yasemin Acar, Sascha Fahl, and Michelle L Mazurek · 2016
Earlier work this paper cites.
Semi-supervised knowledge transfer for deep learning from private training data
Nicolas Papernot, Martín Abadi, Ulfar Erlingsson, Ian Goodfellow, and Kunal Talwar · 2016
Earlier work this paper cites.
Strategic classification
Moritz Hardt, Nimrod Megiddo, Christos Papadimitriou, and Mary Wootters · 2016
Earlier work this paper cites.
Data poisoning attacks on factorization-based collaborative filtering
Bo Li, Yining Wang, Aarti Singh, and Yevgeniy Vorobeychik · 2016
Earlier work this paper cites.
A systematic review of the relationship between internet use, self-harm and suicidal behaviour in young people: The good, the bad and the unknown
Amanda Marchant, Keith Hawton, Ann Stewart, Paul Montgomery, Vinod Singaravelu, Keith Lloyd, Nicola Purdy, Kate Daine, and Ann John · 2017
Earlier work this paper cites.
Is the internet causing political polarization? evidence from demographics
Levi Boxell, Matthew Gentzkow, and Jesse M Shapiro · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Fake news detection on social media: A data mining perspective
Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu · 2017
Earlier work this paper cites.
Some like it hoax: Automated fake news detection in social networks
Eugenio Tacchini, Gabriele Ballarin, Marco L Della Vedova, Stefano Moret, and Luca De Alfaro · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Machine learning models that remember too much
Congzheng Song, Thomas Ristenpart, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Calibrated fairness in bandits
Yang Liu, Goran Radanovic, Christos Dimitrakakis, Debmalya Mandal, and David C Parkes · 2017
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova · 2017
Earlier work this paper cites.
Algorithmic decision making and the cost of fairness
Sam Corbett-Davies, Emma Pierson, Avi Feller, Sharad Goel, and Aziz Huq · 2017
Earlier work this paper cites.
Automated crowdturfing attacks and defenses in online review systems
Yuanshun Yao, Bimal Viswanath, Jenna Cryan, Haitao Zheng, and Ben Y Zhao · 2017
Earlier work this paper cites.
Social network security: Issues, challenges, threats, and solutions
Shailendra Rathore, Pradip Kumar Sharma, Vincenzo Loia, Young-Sik Jeong, and Jong Hyuk Park · 2017
Earlier work this paper cites.
Malicious accounts: Dark of the social networks
Kayode Sakariyah Adewole, Nor Badrul Anuar, Amirrudin Kamsin, Kasturi Dewi Varathan, and Syed Abdul Razak · 2017
Earlier work this paper cites.
Systematization of knowledge (sok): A systematic review of software-based web phishing detection
Zuochao Dou, Issa Khalil, Abdallah Khreishah, Ala Al-Fuqaha, and Mohsen Guizani · 2017
Earlier work this paper cites.
Deep multimodal learning: A survey on recent advances and trends
Dhanesh Ramachandram and Graham W Taylor · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Sandra Wachter, Brent Mittelstadt, and Chris Russell · 2017
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran · 2017
Earlier work this paper cites.
Hate me, hate me not: Hate speech detection on facebook
Fabio Del Vigna12, Andrea Cimino23, Felice Dell’Orletta, Marinella Petrocchi, and Maurizio Tesconi · 2017
Earlier work this paper cites.
Unbiased learning-to-rank with biased feedback
Thorsten Joachims, Adith Swaminathan, and Tobias Schnabel · 2017
Earlier work this paper cites.
Normative challenges of identification in the internet of things: Privacy, profiling, discrimination, and the gdpr
Sandra Wachter · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Ensure the correctness of the summary: Incorporate entailment knowledge into abstractive sentence summarization
Haoran Li, Junnan Zhu, Jiajun Zhang, and Chengqing Zong · 2018
Earlier work this paper cites.
Faithful to the original: Fact aware neural abstractive summarization
Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li · 2018
Earlier work this paper cites.
Machine learning with membership privacy using adversarial regularization
Milad Nasr, Reza Shokri, and Amir Houmansadr · 2018
Earlier work this paper cites.
Property inference attacks on fully connected neural networks using permutation invariant representations
Karan Ganju, Qi Wang, Wei Yang, Carl A Gunter, and Nikita Borisov · 2018
Earlier work this paper cites.
Reverse engineering convolutional neural networks through side-channel information leaks
Weizhe Hua, Zhiru Zhang, and G Edward Suh · 2018
Earlier work this paper cites.
Stealing hyperparameters in machine learning
Binghui Wang and Neil Zhenqiang Gong · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha · 2018
Earlier work this paper cites.
A pragmatic introduction to secure multi-party computation
David Evans, Vladimir Kolesnikov, Mike Rosulek, et al · 2018
Earlier work this paper cites.
Aby3: A mixed protocol framework for machine learning
Payman Mohassel and Peter Rindal · 2018
Earlier work this paper cites.
Secure logistic regression based on homomorphic encryption: Design and evaluation
Miran Kim, Yongsoo Song, Shuang Wang, Yuhou Xia, Xiaoqian Jiang, et al · 2018
Earlier work this paper cites.
Decoupled classifiers for group-fair and efficient machine learning
Cynthia Dwork, Nicole Immorlica, Adam Tauman Kalai, and Max Leiserson · 2018
Earlier work this paper cites.
Gender stereotypes
Naomi Ellemers · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
The accuracy, fairness, and limits of predicting recidivism
Julia Dressel and Hany Farid · 2018
Earlier work this paper cites.
A reductions approach to fair classification
Alekh Agarwal, Alina Beygelzimer, Miroslav Dudík, John Langford, and Hanna Wallach · 2018
Earlier work this paper cites.
The spread of true and false news online
Soroush Vosoughi, Deb Roy, and Sinan Aral · 2018
Earlier work this paper cites.
False information on web and social media: A survey
Srijan Kumar and Neil Shah · 2018
Earlier work this paper cites.
Fake profile detection techniques in large-scale online social networks: A comprehensive review
Devakunchari Ramalingam and Valliyammai Chinnaiah · 2018
Earlier work this paper cites.
Multi-factor authentication: A survey
Aleksandr Ometov, Sergey Bezzateev, Niko Mäkitalo, Sergey Andreev, Tommi Mikkonen, and Yevgeni Koucheryavy · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Scalable private learning with pate
Nicolas Papernot, Shuang Song, Ilya Mironov, Ananth Raghunathan, Kunal Talwar, and Úlfar Erlingsson · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Explainable artificial intelligence: A survey
Filip Karlo Došilović, Mario Brčić, and Nikica Hlupić · 2018
Earlier work this paper cites.
Machine learning suites for online toxicity detection
David Noever · 2018
Earlier work this paper cites.
Challenges for toxic comment classification: An in-depth error analysis
Betty Van Aken, Julian Risch, Ralf Krestel, and Alexander Löser · 2018
Earlier work this paper cites.
Delayed impact of fair machine learning
Lydia T Liu, Sarah Dean, Esther Rolf, Max Simchowitz, and Moritz Hardt · 2018
Earlier work this paper cites.
Manipulating machine learning: Poisoning attacks and countermeasures for regression learning
Matthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu, Cristina Nita-Rotaru, and Bo Li · 2018
Earlier work this paper cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Ali Shafahi, W Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan · 2019
Earlier work this paper cites.
A simple recipe towards reducing hallucination in neural surface realisation
Feng Nie, Jin-Ge Yao, Jinpeng Wang, Rong Pan, and Chin-Yew Lin · 2019
Earlier work this paper cites.
Incorporating external knowledge into machine reading for generative question answering
Bin Bi, Chen Wu, Ming Yan, Wei Wang, Jiangnan Xia, and Chenliang Li · 2019
Earlier work this paper cites.
Using local knowledge graph construction to scale seq2seq models to multi-document inputs
Angela Fan, Claire Gardent, Chloé Braud, and Antoine Bordes · 2019
Earlier work this paper cites.
Examining emergent communities and social bots within the polarized online vaccination debate in twitter
Xiaoyi Yuan, Ross J Schuchard, and Andrew T Crooks · 2019
Earlier work this paper cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Cited alongside, same era.
Deep leakage from gradients
Ligeng Zhu, Zhijian Liu, and Song Han · 2019
Cited alongside, same era.
Exploiting unintended feature leakage in collaborative learning
Luca Melis, Congzheng Song, Emiliano De Cristofaro, and Vitaly Shmatikov · 2019
Cited alongside, same era.
Model inversion attacks against collaborative inference
Zecheng He, Tianwei Zhang, and Ruby B Lee · 2019
Cited alongside, same era.
Knockoff nets: Stealing functionality of black-box models
Tribhuvanesh Orekondy, Bernt Schiele, and Mario Fritz · 2019
Cited alongside, same era.
Towards reverse-engineering black-box neural networks
Seong Joon Oh, Bernt Schiele, and Mario Fritz · 2019
Cited alongside, same era.
Prompting gpt-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang · 2022
Later among the works it cites.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau · 2022
Later among the works it cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Later among the works it cites.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Later among the works it cites.
Towards improving selective prediction ability of nlp systems
Neeraj Varshney, Swaroop Mishra, and Chitta Baral · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Cited alongside, same era.
Certified data removal from machine learning models
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens Van Der Maaten · 2019
Cited alongside, same era.
Towards federated learning at scale: System design
Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloe Kiddon, Jakub Konečnỳ, Stefano Mazzocchi, Brendan McMahan, et al · 2019
Cited alongside, same era.
Agnostic federated learning
Mehryar Mohri, Gary Sivek, and Ananda Theertha Suresh · 2019
Cited alongside, same era.
A quasi-newton method based vertical federated learning framework for logistic regression
Kai Yang, Tao Fan, Tianjian Chen, Yuanming Shi, and Qiang Yang · 2019
Cited alongside, same era.
Fairness without harm: Decoupled classifiers with preference guarantees
Berk Ustun, Yang Liu, and David Parkes · 2019
Cited alongside, same era.
Later among the works it cites.
Neeraj Varshney, Swaroop Mishra, and Chitta Baral · 2022
Later among the works it cites.
Uncertainty estimation for natural language processing
Adam Fisch, Robin Jia, and Tal Schuster · 2022
Later among the works it cites.
Discovering language model behaviors with model-written evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Later among the works it cites.
Mitigating covertly unsafe text within natural language systems
Alex Mei, Anisha Kabir, Sharon Levy, Melanie Subbiah, Emily Allaway, John Judge, Desmond Patton, Bruce Bimber, Kathleen McKeown, and William Yang Wang · 2022
Later among the works it cites.
In conversation with artificial intelligence: aligning language models with human values, 2022
Atoosa Kasirzadeh and Iason Gabriel · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements, 2022
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving · 2022
Later among the works it cites.
A survey of artificial intelligence strategies for automatic detection of sexually explicit videos
Jenny Cifuentes, Ana Lucila Sandoval Orozco, and Luis Javier Garcia Villalba · 2022
Later among the works it cites.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer · 2022
Later among the works it cites.
Enhanced membership inference attacks against machine learning models
Jiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler, and Reza Shokri · 2022
Later among the works it cites.
Towards data-free model stealing in a hard label setting
Sunandini Sanyal, Sravanti Addepalli, and R Venkatesh Babu · 2022
Later among the works it cites.
Deepsteal: Advanced model extractions leveraging efficient weight stealing in memories
Adnan Siraj Rakin, Md Hafizul Islam Chowdhuryy, Fan Yao, and Deliang Fan · 2022
Later among the works it cites.
Measuring forgetting of memorized training examples
Matthew Jagielski, Om Thakkar, Florian Tramer, Daphne Ippolito, Katherine Lee, Nicholas Carlini, Eric Wallace, Shuang Song, Abhradeep Thakurta, Nicolas Papernot, et al · 2022
Later among the works it cites.
The privacy onion effect: Memorization is relative
Nicholas Carlini, Matthew Jagielski, Chiyuan Zhang, Nicolas Papernot, Andreas Terzis, and Florian Tramer · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan · 2022
Later among the works it cites.
Deduplicating training data mitigates privacy risks in language models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Later among the works it cites.
{ \{ DeepPhish } \} : Understanding user trust towards artificially generated profiles in online social networks
Jaron Mink, Licheng Luo, Natã M Barbosa, Olivia Figueira, Yang Wang, and Gang Wang · 2022
Later among the works it cites.
How large language models are transforming machine-paraphrased plagiarism
Jan Philip Wahle, Terry Ruas, Frederic Kirstein, and Bela Gipp · 2022
Later among the works it cites.
Towards artificial general intelligence via a multimodal foundation model
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, et al · 2022
Later among the works it cites.
Data isotopes for data provenance in dnns
Emily Wenger, Xiuyu Li, Ben Y Zhao, and Vitaly Shmatikov · 2022
Later among the works it cites.
Explainable, trustworthy, and ethical machine learning for healthcare: A survey
Khansa Rasheed, Adnan Qayyum, Mohammed Ghaly, Ala Al-Fuqaha, Adeel Razi, and Junaid Qadir · 2022
Later among the works it cites.
Opening the black box: the promise and limitations of explainable machine learning in cardiology
Jeremy Petch, Shuang Di, and Walter Nelson · 2022
Later among the works it cites.
Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011–2022)
Hui Wen Loh, Chui Ping Ooi, Silvia Seoni, Prabal Datta Barua, Filippo Molinari, and U Rajendra Acharya · 2022
Later among the works it cites.
Post-hoc interpretability for neural nlp: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar · 2022
Later among the works it cites.
Interpreting language models with contrastive explanations
Kayo Yin and Graham Neubig · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Later among the works it cites.
Few-shot learning with retrieval augmented language models
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave · 2022
Later among the works it cites.
Langchain
Harrison Chase · 2022
Later among the works it cites.
On the paradox of learning to reason from data
Honghua Zhang, Liunian Harold Li, Tao Meng, Kai-Wei Chang, and Guy Van den Broeck · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al · 2022
Later among the works it cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic · 2022
Later among the works it cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Later among the works it cites.
Solving math word problems with process-and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Later among the works it cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi · 2022
Later among the works it cites.
Probabilities of causation: three counterfactual interpretations and their identification
Judea Pearl · 2022
Later among the works it cites.
The ghost in the machine has an american accent: value conflict in gpt-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo · 2022
Later among the works it cites.
Who is gpt-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg · 2022
Later among the works it cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Later among the works it cites.
Fairness transferability subject to bounded distribution shift
Yatong Chen, Reilly Raab, Jialu Wang, and Yang Liu · 2022
Later among the works it cites.
Breaking feedback loops in recommender systems with causal inference
Karl Krauth, Yixin Wang, and Michael I Jordan · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Coyo-700m: Image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim · 2022
Later among the works it cites.
Gpt-4 system card, https://cdn.openai.com/papers/gpt-4-system-card.pdf
OpenAI · 2023
Closest in time.
Hate speech in the internet context: Unpacking the roles of internet penetration, online legal regulation, and online opinion polarization from a transnational perspective
Zikun Liu, Chen Luo, and Jia Lu · 2023
Closest in time.
Evaluating the social impact of generative ai systems in systems and society
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Hal Daumé III, Jesse Dodge, Ellie Evans, Sara Hooker, et al · 2023
Closest in time.
Eight things to know about large language models
Samuel R Bowman · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Closest in time.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Training socially aligned language models in simulated human society
Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M Dai, Diyi Yang, and Soroush Vosoughi · 2023
Closest in time.
Large language models and software as a medical device
Johan Ordish · 2023
Closest in time.
Are large language models ready for healthcare? a comparative study on clinical language understanding, 2023
Yuqing Wang, Yun Zhao, and Linda Petzold · 2023
Closest in time.
Bloomberggpt: A large language model for finance, 2023
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Closest in time.
Fingpt: Open-source financial large language models, 2023
Hongyang Yang, Xiao-Yang Liu, and Christina Dan Wang · 2023
Closest in time.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Closest in time.
A categorical archive of chatgpt failures
Ali Borji · 2023
Closest in time.
Chatgpt and software testing education: Promises & perils
Sajed Jalil, Suzzana Rafi, Thomas D LaToza, Kevin Moran, and Wing Lam · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
Artificial hallucinations in chatgpt: implications in scientific writing
Hussam Alkaissi and Samy I McFarlane · 2023
Closest in time.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al · 2023
Closest in time.
Why does chatgpt fall short in answering questions faithfully?
Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang · 2023
Closest in time.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Closest in time.
Consistency analysis of chatgpt
Myeongjun Jang and Thomas Lukasiewicz · 2023
Closest in time.
Evaluating task understanding through multilingual consistency: A chatgpt case study
Xenia Ohmer, Elia Bruni, and Dieuwke Hupkes · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Closest in time.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto · 2023
Closest in time.
To aggregate or not? learning with separate noisy labels
Jiaheng Wei, Zhaowei Zhu, Tianyi Luo, Ehsan Amid, Abhishek Kumar, and Yang Liu · 2023
Closest in time.
Conformal prediction with large language models for multi-choice question answering
Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu, David Bellamy, Ramesh Raskar, and Andrew Beam · 2023
Closest in time.
Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S Jaakkola, and Regina Barzilay · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Closest in time.
The risks of using chatgpt to obtain common safety-related information and advice
Oscar Oviedo-Trespalacios, Amy E Peden, Thomas Cole-Hunter, Arianna Costantini, Milad Haghani, Sage Kelly, Helma Torkamaan, Amina Tariq, James David Albert Newton, Timothy Gallagher, et al · 2023
Closest in time.
Multimodal chain-of-thought reasoning in language models
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola · 2023
Closest in time.
Generating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov · 2023
Closest in time.
Role of chat gpt in public health
Som S Biswas · 2023
Closest in time.
Chat-gpt: Opportunities and challenges in child mental healthcare
Nazish Imran, Aateqa Hashmi, and Ahad Imran · 2023
Closest in time.
Exploring ai ethics of chatgpt: A diagnostic analysis
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Closest in time.
Should chatgpt be biased? challenges and risks of bias in large language models
Emilio Ferrara · 2023
Closest in time.
The New York Times, 2023
F.t.c. opens investigation into chatgpt maker over technology’s potential harms · 2023
Closest in time.
UK Legislation, 2006
Racial and religious hatred act 2006 · 2023
Closest in time.
U.S. Government Publishing Office, 1990
Americans with disabilities act of 1990 · 2023
Closest in time.
Federal Register of Legislation, 2009
Fair work act 2009 · 2023
Closest in time.
UK Legislation, 2010
Equality act 2010 · 2023
Closest in time.
Federal Trade Commission, 2021
Federal trade commission. no fear act protections against discrimination and other prohibited practices · 2023
Closest in time.
The political biases of chatgpt
David Rozado · 2023
Closest in time.
Is chat gpt biased against conservatives? an empirical study
Robert W McGee · 2023
Closest in time.
Who were the 10 best and 10 worst us presidents? the opinion of chat gpt (artificial intelligence)
Robert W McGee · 2023
Closest in time.
The self-perception and political biases of chatgpt
Jérôme Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, and Markus Pauly · 2023
Closest in time.
Is chatgpt a highly fluent grammatical error correction system? a comprehensive evaluation
Tao Fang, Shu Yang, Kaixin Lan, Derek F Wong, Jinpeng Hu, Lidia S Chao, and Yue Zhang · 2023
Closest in time.
Does chatgpt provide appropriate and equitable medical advice?: A vignette-based, clinical evaluation across care contexts
Anthony J Nastasi, Katherine R Courtright, Scott D Halpern, and Gary E Weissman · 2023
Closest in time.
Study and analysis of chat gpt and its impact on different fields of study
Dinesh Kalla and Nathan Smith · 2023
Closest in time.
Is chatgpt a good translator? a preliminary study
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu · 2023
Closest in time.
Assessing cross-cultural alignment between chatgpt and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich · 2023
Closest in time.
Impact of big data analytics and chatgpt on cybersecurity
Pawankumar Sharma and Bibhu Dash · 2023
Closest in time.
PV Charan, Hrushikesh Chunduri, P Mohan Anand, and Sandeep K Shukla · 2023
Closest in time.
Weaponising chatgpt
Steve Mansfield-Devine · 2023
Closest in time.
New ai classifier for indicating ai-written text
OpenAI · 2023
Closest in time.
Gpt detector
Writefull X · 2023
Closest in time.
https://contentdetector.ai/, 2023
Ai content detector · 2023
Closest in time.
Exploring the mit mathematics and eecs curriculum using large language models
Sarah J Zhang, Samuel Florin, Ariel N Lee, Eamon Niknafs, Andrei Marginean, Annie Wang, Keith Tyser, Zad Chin, Yann Hicke, Nikhil Singh, et al · 2023
Closest in time.
Foundation models and fair use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang · 2023
Closest in time.
Do language models plagiarize?
Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee · 2023
Closest in time.
The New York Times, 2023
Sarah silverman sues openai and meta over copyright infringement · 2023
Closest in time.
The Wall Street Journal, 2023
Thousands of authors ask ai chatbot owners to pay for use of their work · 2023
Closest in time.
Watermarking text data on large language models for dataset copyright protection
Yixin Liu, Hongsheng Hu, Xuyun Zhang, and Lichao Sun · 2023
Closest in time.
Provable copyright protection for generative models
Nikhil Vyas, Sham Kakade, and Boaz Barak · 2023
Closest in time.
Glaze: Protecting artists from style mimicry by text-to-image models
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng, Rana Hanocka, and Ben Y Zhao · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Closest in time.
On the reliability of watermarks for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein · 2023
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality
Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan · 2023
Closest in time.
Explainable ai (xai): A systematic meta-survey of current challenges and future opportunities
Waddah Saeed and Christian Omlin · 2023
Closest in time.
Inseq: An interpretability toolkit for sequence generation models
Gabriele Sarti, Nils Feldhus, Ludwig Sickert, Oskar van der Wal, Malvina Nissim, and Arianna Bisazza · 2023
Closest in time.
Sequential integrated gradients: a simple but effective method for explaining language models
Joseph Enguehard · 2023
Closest in time.
Improving accuracy of gpt-3/4 results on biomedical data using a retrieval-augmented language model
David Soong, Sriram Sridhar, Han Si, Jan-Samuel Wagner, Ana Caroline Costa Sá, Christina Y Yu, Kubra Karagoz, Meijian Guan, Hisham Hamadeh, and Brandon W Higgs · 2023
Closest in time.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Closest in time.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Alexis Chevalier, and Julius Berner · 2023
Closest in time.
Evaluating the logical reasoning ability of chatgpt and gpt-4
Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yue Zhang · 2023
Closest in time.
Measuring inductive biases of in-context learning with underspecified demonstrations
Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng, Danqi Chen, and He He · 2023
Closest in time.
True detective: A deep abductive reasoning benchmark undoable for gpt-3 and challenging for gpt-4
Maksym Del and Mark Fishel · 2023
Closest in time.
Towards complex reasoning: the polaris of large language models, July 2023
Yao Fu · 2023
Closest in time.
Starcoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al · 2023
Closest in time.
Specializing smaller language models towards multi-step reasoning
Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot · 2023
Closest in time.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot · 2023
Closest in time.
Causal-discovery performance of chatgpt in the context of neuropathic pain diagnosis
Ruibo Tu, Chao Ma, and Cheng Zhang · 2023
Closest in time.
Can large language models infer causation from correlation?, 2023
Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona Diab, and Bernhard Schölkopf · 2023
Closest in time.
A new era in internet interventions: The advent of chat-gpt and ai-assisted therapist guidance
Per Carlbring, Heather Hadjistavropoulos, Annet Kleiboer, and Gerhard Andersson · 2023
Closest in time.
Chatgpt outperforms humans in emotional awareness evaluations
Zohar Elyoseph, Dorit Hadar-Shoval, Kfir Asraf, and Maya Lvovsky · 2023
Closest in time.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al · 2023
Closest in time.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al · 2023
Closest in time.
Terry Yue Zhuo, Zhuang Li, Yujin Huang, Yuan-Fang Li, Weiqing Wang, Gholamreza Haffari, and Fatemeh Shiri · 2023
Closest in time.
Bias and debias in recommender system: A survey and future directions
Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He · 2023
Closest in time.
Long-term fairness with unknown dynamics
Tongxin Yin, Reilly Raab, Mingyan Liu, and Yang Liu · 2023
Closest in time.
Rectifying unfairness in recommendation feedback loop
Mengyue Yang, Jun Wang, and Jean-Francois Ton · 2023
Closest in time.
Debiasing recommendation by learning identifiable latent confounders
Qing Zhang, Xiaoying Zhang, Yang Liu, Hongning Wang, Min Gao, Jiheng Zhang, and Ruocheng Guo · 2023
Closest in time.
Poisoning web-scale training datasets is practical
Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2023
Closest in time.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2023
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model, 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Closest in time.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Closest in time.
Cybercrime to cost the world $10.5 trillion annually by 2025
Steve Morgan · 2025
Closest in time.