Fetching the paper…
Reading the bibliography…
Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike.
Adversarial attacks against fact extraction and verification
J. Thorne and A. Vlachos · 1903
Earlier work this paper cites.
Sight, sound, and stereotype: The war on terrorism and its consequences for latinas/os
S. W. Bender · 2002
Earlier work this paper cites.
The enron corpus: A new dataset for email classification research
B. Klimt and Y. Yang · 2004
Earlier work this paper cites.
Uci machine learning repository, 2007
A. Asuncion and D. Newman · 2007
Earlier work this paper cites.
Black criminal stereotypes and racial profiling
K. Welch · 2007
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
S. Bird, E. Klein, and E. Loper · 2009
Earlier work this paper cites.
Immunity to popular stereotypes of aging? seniors and stereotype threat
S. Horton, J. Baker, W. Pearce, and J. M. Deakin · 2010
Earlier work this paper cites.
Opposition to pro-immigrant public policy: Symbolic racism and group threat
J. A. Berg · 2012
Earlier work this paper cites.
Fairness through awareness
C. Dwork, M. Hardt, T. Pitassi, O. Reingold, and R. Zemel · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Five stereotypes about poor families and education
Washington Post · 2013
Earlier work this paper cites.
Learning fair representations
R. Zemel, Y. Wu, K. Swersky, T. Pitassi, and C. Dwork · 2013
Earlier work this paper cites.
Bad drivers? no, just bad stereotypes
Association for Psychological Science · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
C. Dwork, A. Roth, et al · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
Big data’s disparate impact
S. Barocas and A. D. Selbst · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings, 2016
T. Bolukbasi, K.-W. Chang, J. Zou, V. Saligrama, and A. Kalai · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
M. Hardt, E. Price, E. Price, and N. Srebro · 2016
Earlier work this paper cites.
A racist stereotype is shattered: Study finds white youth are more likely to abuse hard drugs than black youth
Salon · 2016
Earlier work this paper cites.
Do immigrants “steal” jobs from american workers?
Brookings Institution · 2017
Earlier work this paper cites.
Stereotype threat among girls: Differences by gender identity and math education context
B. J. Casad, P. Hale, and F. L. Wachs · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
R. Jia and P. Liang · 2017
Earlier work this paper cites.
Counterfactual fairness
M. J. Kusner, J. Loftus, C. Russell, and R. Silva · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Yellow peril, red scare: race and communism in national review
S. D. Visco · 2017
Earlier work this paper cites.
Textworld: A learning environment for text-based games
M. Côté, Á. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, M. J. Hausknecht, L. E. Asri, M. Adada, W. Tay, and A. Trischler · 2018
Earlier work this paper cites.
Hierarchical neural story generation
A. Fan, M. Lewis, and Y. Dauphin · 2018
Earlier work this paper cites.
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford · 2018
Earlier work this paper cites.
‘you play like a woman!’ effects of gender stereotype threat on women’s performance in physical and sport activities: A meta-analysis
A. Gentile, S. Boca, and I. Giammusso · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
M. Iyyer, J. Wieting, K. Gimpel, and L. Zettlemoyer · 2018
Earlier work this paper cites.
204How Did East Asians Become Yellow?
M. Keevak · 2018
Earlier work this paper cites.
Stress test evaluation for natural language inference
A. Naik, A. Ravichander, N. M. Sadeh, C. P. Rosé, and G. Neubig · 2018
Earlier work this paper cites.
Roles for computing in social change
R. Abebe, S. Barocas, J. Kleinberg, K. Levy, M. Raghavan, and D. G. Robinson · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song · 2019
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
J. Dai, C. Chen, and Y. Li · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
MRQA 2019 shared task: Evaluating generalization in reading comprehension
A. Fisch, A. Talmor, R. Jia, M. Seo, E. Choi, and D. Chen · 2019
Earlier work this paper cites.
Openwebtext corpus
A. Gokaslan and V. Cohen · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
Learning the difference that makes a difference with counterfactually-augmented data
D. Kaushik, E. Hovy, and Z. Lipton · 2019
Earlier work this paper cites.
Feature noise induces loss discrepancy across groups
F. Khani and P. Liang · 2019
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
J. Li, S. Ji, T. Du, B. Li, and T. Wang · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
T. McCoy, E. Pavlick, and T. Linzen · 2019
Earlier work this paper cites.
Distributionally robust language modeling
Y. Oren, S. Sagawa, T. B. Hashimoto, and P. Liang · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Earlier work this paper cites.
Inherent tradeoffs in learning fair representations
H. Zhao and G. Gordon · 2019
Earlier work this paper cites.
An atlas of cultural commonsense for machine reasoning
A. Acharya, K. Talamadupula, and M. A. Finlayson · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
S. L. Blodgett, S. Barocas, H. Daumé III, and H. Wallach · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Fairness in machine learning: A survey
S. Caton and C. Haas · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, and Y. Choi · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Earlier work this paper cites.
Interactive fiction games: A colossal adventure
M. J. Hausknecht, P. Ammanabrolu, M. Côté, and X. Yuan · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
D. Hendrycks, X. Liu, E. Wallace, A. Dziedzic, R. Krishnan, and D. Song · 2020
Earlier work this paper cites.
Is BERT really robust? A strong baseline for natural language attack on text classification and entailment
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
K. Krishna, J. Wieting, and M. Iyyer · 2020
Earlier work this paper cites.
BERT-ATTACK: adversarial attack against BERT using BERT
L. Li, R. Ma, Q. Guo, X. Xue, and X. Qiu · 2020
Earlier work this paper cites.
UNQOVERing stereotyping biases via underspecified questions
T. Li, D. Khashabi, T. Khot, A. Sabharwal, and V. Srikumar · 2020
Earlier work this paper cites.
The radicalization risks of GPT-3 and advanced neural language models
K. McGuffie and A. Newhouse · 2020
Cited alongside, same era.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
N. Nangia, C. Vania, R. Bhalerao, and S. R. Bowman · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Breeds: Benchmarks for subpopulation shift
S. Santurkar, D. Tsipras, and A. Madry · 2020
Cited alongside, same era.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
An empirical analysis of memorization in fine-tuned autoregressive language models
F. Mireshghallah, A. Uniyal, T. Wang, D. K. Evans, and T. Berg-Kirkpatrick · 2022
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
S. Mishra, D. Khashabi, C. Baral, and H. Hajishirzi · 2022
Later among the works it cites.
Unsupervised text deidentification
J. X. Morris, J. T. Chiu, R. Zabih, and A. M. Rush · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Bbq: A hand-built bias benchmark for question answering, 2022
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. R. Bowman · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Cited alongside, same era.
The fox–eye trend isn’t cute—it’s racist
Teen Vogue · 2020
Cited alongside, same era.
T3: tree-autoencoder constrained adversarial text generation for targeted attack
B. Wang, H. Pei, B. Pan, Q. Chen, S. Wang, and B. Li · 2020
Cited alongside, same era.
Learning which features matter: RoBERTa acquires a preference for linguistic generalizations (eventually)
A. Warstadt, Y. Zhang, X. Li, H. Liu, and S. R. Bowman · 2020
Cited alongside, same era.
Neural text generation with unlikelihood training
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston · 2020
Cited alongside, same era.
Keep calm and explore: Language models for action generation in text-based games
S. Yao, R. Rao, M. Hausknecht, and K. Narasimhan · 2020
Cited alongside, same era.
Word-level textual adversarial attacking as combinatorial optimization
Y. Zang, F. Qi, C. Yang, Z. Liu, M. Zhang, Q. Liu, and M. Sun · 2020
Cited alongside, same era.
F. Perez and I. Ribeiro · 2022
Later among the works it cites.
Fairness in federated learning via core-stability
B. Ray Chaudhury, L. Li, M. Kang, B. Li, and R. Mehta · 2022
Later among the works it cites.
Just fine-tune twice: Selective differential privacy for large language models
W. Shi, R. Shea, S. Chen, C. Zhang, R. Jia, and Z. Yu · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Later among the works it cites.
Considerations for differentially private learning with large-scale public pretraining
F. Tram‘er, K. Gautam, and N. C. Carlini · 2022
Later among the works it cites.
SemAttack: Natural textual attacks via different semantic spaces
B. Wang, C. Xu, X. Liu, Y. Cheng, and B. Li · 2022
Later among the works it cites.
Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Y. Wang, S. Mishra, P. Alipoormolabashi, Y. Kordi, A. Mirzaei, A. Naik, A. Ashok, A. S. Dhanasekaran, A. Arunkumar, D. Stap, E. Pathak, G. Karamanolakis, H. Lai, I. Purohit, I. Mondal, J. Anderson, K. Kuznia, K. Doshi, K. K. Pal, M. Patel, M. Moradshahi, M. Parmar, M. Purohit, N. Varshney, P. R. Kaza, P. Verma, R. S. Puri, R. Karia, S. Doshi, S. K. Sampat, S. Mishra, S. Reddy A, S. Patro, T. Dixit, and X. Shen · 2022
Later among the works it cites.
Certifying out-of-domain generalization for blackbox functions
M. Weber, L. Li, B. Wang, Z. Zhao, B. Li, and C. Zhang · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2022
Later among the works it cites.
Blueprint for an ai bill of rights
White House Office of Science and Technology Policy · 2022
Later among the works it cites.
Prompt injection attacks against gpt-3
S. Willison · 2022
Later among the works it cites.
Ground-truth labels matter: A deeper look into input-label demonstrations
K. M. Yoo, J. Kim, H. J. Kim, H. Cho, H. Jo, S.-W. Lee, S.-g. Lee, and T. Kim · 2022
Later among the works it cites.
Differentially private fine-tuning of language models
D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, et al · 2022
Later among the works it cites.
Provably confidential language modelling
X. Zhao, L. Li, and Y.-X. Wang · 2022
Later among the works it cites.
Falcon-40B: an open large language model with state-of-the-art performance
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, E. Goffinet, D. Heslow, J. Launay, Q. Malartic, B. Noune, B. Pannier, and G. Penedo · 2023
Closest in time.
Do foundation model providers comply with the eu ai act?, 2023
R. Bommasani, K. Klyman, D. Zhang, and P. Liang · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Microsoft is bringing chatgpt technology to word, excel and outlook, 2023
CNN · 2023
Closest in time.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
T. Computer · 2023
Closest in time.
Lessons learned from chatgpt’s samsung leak, 2023
Cybernews · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.
Flocks of stochastic parrots: Differentially private prompt learning for large language models
H. Duan, A. Dziedzic, N. Papernot, and F. Boenisch · 2023
Closest in time.
The capacity for moral self-correction in large language models, 2023
D. Ganguli, A. Askell, N. Schiefer, T. I. Liao, K. Lukošiūtė, A. Chen, A. Goldie, A. Mirhoseini, C. Olsson, D. Hernandez, D. Drain, D. Li, E. Tran-Johnson, E. Perez, J. Kernion, J. Kerr, J. Mueller, J. Landau, K. Ndousse, K. Nguyen, L. Lovitt, M. Sellitto, N. Elhage, N. Mercado, N. DasSarma, O. Rausch, R. Lasenby, R. Larson, S. Ringer, S. Kundu, S. Kadavath, S. Johnston, S. Kravec, S. E. Showk, T. Lanham, T. Telleen-Lawton, T. Henighan, T. Hume, Y. Bai, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, C. Olah, J. Clark, S. R. Bowman, and J. Kaplan · 2023
Closest in time.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz · 2023
Closest in time.
W. Hariri · 2023
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Closest in time.
Introduction to prompt hacking
Learn Prompting · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on chatgpt
H. Li, D. Guo, W. Fan, M. Xu, and Y. Song · 2023
Closest in time.
Y. Li and Y. Zhang · 2023
Closest in time.
Analyzing leakage of personally identifiable information in language models
N. Lukas, A. Salem, R. Sim, S. Tople, L. Wutschitz, and S. Zanella-Béguelin · 2023
Closest in time.
Adversarial prompting for black box foundation models
N. Maus, P. Chao, E. Wong, and J. Gardner · 2023
Closest in time.
Capabilities of gpt-4 on medical challenge problems
H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz · 2023
Closest in time.
GPT-4 technical report
OpenAI · 2023
Closest in time.
A. Pan, J. S. Chan, A. Zou, N. Li, S. Basart, T. Woodside, J. Ng, H. Zhang, S. Emmons, and D. Hendrycks · 2023
Closest in time.
Differentially private in-context learning
A. Panda, T. Wu, J. T. Wang, and P. Mittal · 2023
Closest in time.
Amendments adopted by the european parliament on 14 june 2023 on the proposal for a regulation of the european parliament and of the council on laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts
E. Parliament · 2023
Closest in time.
H. Qiu, S. Zhang, A. Li, H. He, and Z. Lan · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2023
Closest in time.
H. Shao, J. Huang, S. Zheng, and K. C.-C. Chang · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Closest in time.
Prompting GPT-3 to be reliable
C. Si, Z. Gan, Z. Yang, S. Wang, J. Wang, J. L. Boyd-Graber, and L. Wang · 2023
Closest in time.
StableVicuna: An RLHF Fine-Tune of Vicuna-13B v0
StabilityAI · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, ly usable llms, 2023
M. N. Team · 2023
Closest in time.
Myths about hiv
The Human Rights Campaign · 2023
Closest in time.
Shall we pretrain autoregressive language models with retrieval? a comprehensive study
B. Wang, W. Ping, P. Xu, L. McAfee, Z. Liu, M. Shoeybi, Y. Dong, O. Kuchaiev, B. Li, C. Xiao, A. Anandkumar, and B. Catanzaro · 2023
Closest in time.
Larger language models do in-context learning differently
J. Wei, J. Wei, Y. Tay, D. Tran, A. Webson, Y. Lu, X. Chen, H. Liu, D. Huang, D. Zhou, et al · 2023
Closest in time.
Revisiting out-of-distribution robustness in nlp: Benchmark, analysis, and llms evaluations
L. Yuan, Y. Chen, G. Cui, H. Gao, F. Zou, X. Cheng, H. Ji, Z. Liu, and M. Sun · 2023
Closest in time.
Synthetic text generation with differential privacy: A simple and practical recipe
X. Yue, H. A. Inan, X. Li, G. Kumar, J. McAnallen, H. Sun, D. Levitan, and R. Sim · 2023
Closest in time.
Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert
Q. Zhong, L. Ding, J. Liu, B. Du, and D. Tao · 2023
Closest in time.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
K. Zhu, J. Wang, J. Zhou, Z. Wang, H. Chen, Y. Wang, L. Yang, W. Ye, N. Z. Gong, Y. Zhang, et al · 2023
Closest in time.
Exploring ai ethics of chatgpt: A diagnostic analysis
T. Y. Zhuo, Y. Huang, C. Chen, and Z. Xing · 2023
Closest in time.