Fetching the paper…
Reading the bibliography…
Although large language models (LLMs) have been touted for their ability to generate natural-sounding text, there are growing concerns around possible negative effects of LLMs such as data memorization, bias, and inappropriate language.
“cloze procedure”: A new tool for measuring readability
Taylor, W. L · 1953
Earlier work this paper cites.
A new algorithm for data compression
Gage, P · 1994
Earlier work this paper cites.
Speech recognition by composition of weighted finite automata
Pereira, F. C. N. and Riley, M. D · 1996
Earlier work this paper cites.
Finite-state transducers in language and speech processing
Mohri, M · 1997
Earlier work this paper cites.
Open-domain question-answering
Prager, J. M · 2006
Earlier work this paper cites.
Openfst: A general and efficient weighted finite-state transducer library
Allauzen, C., Riley, M., Schalkwyk, J., Skut, W., and Mohri, M · 2007
Earlier work this paper cites.
Introduction to Automata Theory, Languages, and Computation
Hopcroft, J. E., Motwani, R., and Ullman, J. D · 2007
Earlier work this paper cites.
Language independent text correction using finite state automata
Hassan, A., Noeman, S., and Hassan, H · 2008
Earlier work this paper cites.
Natural Language Processing with Python
Bird, S., Loper, E., and Klein, E · 2009
Earlier work this paper cites.
The OpenGrm open-source finite-state grammar software libraries
Roark, B., Sproat, R., Allauzen, C., Riley, M., Sorensen, J., and Tai, T · 2012
Earlier work this paper cites.
Pushdown automata in statistical machine translation
Allauzen, C., Byrne, B., de Gispert, A., Iglesias, G., and Riley, M · 2014
Earlier work this paper cites.
Pynini: A Python library for weighted finite-state grammar compilation
Gorman, K · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Reasoning about entailment with neural attention
Rocktäschel, T., Grefenstette, E., Hermann, K. M., Kočiskỳ, T., and Blunsom, P · 2016
Earlier work this paper cites.
Three ways to count walks in a digraph
Yancey, M · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, U., Kos, J., and Song, D · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Kurita, K., Vyas, N., Pareek, A., Black, A. W., and Tsvetkov, Y · 2019
Earlier work this paper cites.
Finite-State Techniques: Automata, Transducers and Bimachines
Mihov, S. and Schulz, K. U · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Cited alongside, same era.
Measuring and reducing gendered correlations in pre-trained models
What will it take to fix benchmarking in natural language understanding?
Bowman, S. R. and Dahl, G. E · 2021
Later among the works it cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., and Raffel, C · 2021
Later among the works it cites.
Autoregressive entity retrieval
De Cao, N., Izacard, G., Riedel, S., and Petroni, F · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D · 2021
Later among the works it cites.
run_generation.py
HuggingFace · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beutel, A., Chi, E. H., Pavlick, E., Pitler, E. B., Tenney, I., Chen, J., Webster, K., Petrov, S., and Wang, X · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
A snapshot of the frontiers of fairness in machine learning
Chouldechova, A. and Roth, A · 2020
Cited alongside, same era.
To test machine comprehension, start by defining comprehension
Dunietz, J., Burnham, G., Bharadwaj, A., Rambow, O., Chu-Carroll, J., and Ferrucci, D · 2020
Cited alongside, same era.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2020
Cited alongside, same era.
Later among the works it cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Kirk, H. R., Jun, Y., Volpin, F., Iqbal, H., Benussi, E., Dreyer, F., Shtedritski, A., and Asano, Y · 2021
Later among the works it cites.
Designing toxic content classification for a diversity of perspectives
Kumar, D., Kelley, P. G., Consolvo, S., Mason, J., Bursztein, E., Durumeric, Z., Thomas, K., and Bailey, M · 2021
Later among the works it cites.
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J · 2021
Later among the works it cites.
Probing toxic content in large pre-trained language models
Ousidhoum, N., Zhao, X., Fang, T., Song, Y., and Yeung, D.-Y · 2021
Later among the works it cites.
Prompt programming for large language models: Beyond the few-shot paradigm
Reynolds, L. and McDonell, K · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Schick, T. and Schütze, H · 2021
Later among the works it cites.
Prompting is programming: A query language for large language models
Beurer-Kellner, L., Fischer, M., and Vechev, M · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Closest in time.
A note on two problems in connexion with graphs
Dijkstra, E. W · 2022
Closest in time.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E · 2022
Closest in time.
bad_words_ids not working
Huggingface · 2022
Closest in time.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Closest in time.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2023
Closest in time.
Memory augmented large language models are computationally universal
Schuurmans, D · 2023
Closest in time.