Fetching the paper…
Reading the bibliography…
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
The state of sparsity in deep neural networks
T. Gale, E. Elsen, and S. Hooker · 1902
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
D. Borkan, L. Dixon, J. Sorensen, N. Thain, and L. Vasserman · 1903
Earlier work this paper cites.
Multifc: a real-world multi-domain dataset for evidence-based fact checking
I. Augenstein, C. Lioma, D. Wang, L. C. Lima, C. Hansen, and J. G. S. Christian Hansen · 1909
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 1910
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat · 1910
Earlier work this paper cites.
E. Elsen, M. Dukhan, T. Gale, and K. Simonyan · 1911
Earlier work this paper cites.
Relativ [sic] frequency of English speech sounds
G. Dewey · 1923
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Computing machinery and intelligence
A. Turing · 1950
Earlier work this paper cites.
Cramming more components onto integrated circuits, 1965
G. E. Moore et al · 1965
Earlier work this paper cites.
Marked and unmarked: A choice between unequals in semiotic structure
L. R. Waugh · 1982
Earlier work this paper cites.
Language acquisition, data compression and generalization
J. G. Wolff · 1982
Earlier work this paper cites.
A statistical approach to machine translation
P. F. Brown, J. Cocke, S. A. Della Pietra, V. J. Della Pietra, F. Jelinek, J. Lafferty, R. L. Mercer, and P. S. Roossin · 1990
Earlier work this paper cites.
On structuring probabilistic dependences in stochastic language modelling
H. Ney, U. Essen, and R. Kneser · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Statistical methods for speech recognition
F. Jelinek · 1997
Earlier work this paper cites.
The search for simplicity: A fundamental cognitive principle?
N. Chater · 1999
Earlier work this paper cites.
An improved error model for noisy channel spelling correction
E. Brill and R. C. Moore · 2000
Earlier work this paper cites.
Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
A. Griewank and A. Walther · 2000
Earlier work this paper cites.
NLTK: The natural language toolkit
E. Loper and S. Bird · 2002
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin · 2003
Earlier work this paper cites.
What is the state of neural network pruning?
D. W. Blalock, J. J. G. Ortiz, J. Frankle, and J. V. Guttag · 2003
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz · 2004
Earlier work this paper cites.
Woodfisher: Efficient second-order approximations for model compression
S. P. Singh and D. Alistarh · 2004
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
I. Dagan, O. Glickman, and B. Magnini · 2005
Earlier work this paper cites.
Sparse GPU kernels for deep learning
T. Gale, M. Zaharia, C. Young, and E. Elsen · 2006
Earlier work this paper cites.
Large language models in machine translation
T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean · 2007
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
S. Legg and M. Hutter · 2007
Earlier work this paper cites.
Decolonial AI: decolonial theory as sociotechnical foresight in artificial intelligence
S. Mohamed, M. Png, and W. Isaac · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Improving machine reading comprehension with single-choice decision and transfer learning
Y. Jiang, S. Wu, J. Gong, Y. Cheng, P. Meng, W. Lin, Z. Chen, and M. Li · 2011
Earlier work this paper cites.
Empirical evaluation and combination of advanced language modeling techniques
T. Mikolov, A. Deoras, S. Kombrink, L. Burget, and J. H. Černocký · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
T. Chen, I. Goodfellow, and J. Shlens · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Unequal representation and gender stereotypes in image search results for occupations
M. Kay, C. Matuszek, and S. A. Munson · 2015
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of african-american english
S. L. Blodgett, L. Green, and B. O’Connor · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan · 2016
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context, 2016
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
A. Parikh, O. Täckström, D. Das, and J. Uszkoreit · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Compression of neural machine translation models via pruning
A. See, M. Luong, and C. D. Manning · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
A. Caliskan, J. J. Bryson, and A. Narayanan · 2017
Earlier work this paper cites.
A survey on dialogue systems: Recent advances and new frontiers
H. Chen, X. Liu, D. Yin, and J. Tang · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Exploring sparsity in recurrent neural networks
S. Narang, G. F. Diamos, S. Sengupta, and E. Elsen · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
M. Zhu and S. Gupta · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
e-SNLI: Natural language inference with natural language explanations
O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz, and P. Blunsom · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Cited alongside, same era.
Amazon scraps secret AI recruiting tool that showed bias against women
J. Dastin · 2018
Cited alongside, same era.
Measuring and mitigating unintended bias in text classification
L. Dixon, J. Li, J. Sorensen, N. Thain, and L. Vasserman · 2018
Cited alongside, same era.
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford · 2018
Cited alongside, same era.
G. Irving, P. Christiano, and D. Amodei · 2018
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis · 2020
Later among the works it cites.
UnifiedQA: Crossing format boundaries with a single QA system
D. Khashabi, S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi · 2020
Later among the works it cites.
Language models as fact checkers?, 2020
N. Lee, B. Z. Li, S. Wang, W. tau Yih, H. Ma, and M. Khabsa · 2020
Later among the works it cites.
Fastbert: a self-distilling BERT with adaptive inference time
W. Liu, P. Zhou, Z. Wang, Z. Zhao, H. Deng, and Q. Ju · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, T. P. Lillicrap, K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, et al · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Kiritchenko and S. M. Mohammad · 2018
Cited alongside, same era.
T. Kudo and J. Richardson · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Cited alongside, same era.
Mixed precision training
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Cited alongside, same era.
Gender bias in coreference resolution
R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme · 2018
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Later among the works it cites.
Turing-NLG: A 17-billion-parameter language model by Microsoft
C. Rosset · 2020
Later among the works it cites.
Human Compatible
S. Russell · 2020
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
V. Sanh, T. Wolf, and A. M. Rush · 2020
Later among the works it cites.
BERT for evidence retrieval and claim verification
A. Soleimani, C. Monz, and M. Worring · 2020
Later among the works it cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Later among the works it cites.
MT5: A massively multilingual pre-trained text-to-text transformer
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel · 2020
Later among the works it cites.
Reasoning over semantic-level graph for fact checking
W. Zhong, J. Xu, D. Tang, Z. Xu, N. Duan, M. Zhou, J. Wang, and J. Yin · 2020
Later among the works it cites.
GitHub Copilot AI is generating and giving out functional API keys
M. Abubakar · 2021
Closest in time.
Large-scale differentially private BERT
R. Anil, B. Ghazi, V. Gupta, R. Kumar, and P. Manurangsi · 2021
Closest in time.
A general language assistant as a laboratory for alignment
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, et al · 2021
Closest in time.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
E. Ben Zaken, S. Ravfogel, and Y. Goldberg · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Closest in time.
Beyond the imitation game: Measuring and extrapolating the capabilities of language models
BIG-bench collaboration · 2021
Closest in time.
Stereotyping norwegian salmon: an inventory of pitfalls in fairness benchmark datasets
S. L. Blodgett, G. Lopez, A. Olteanu, R. Sim, and H. Wallach · 2021
Closest in time.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. van den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. de Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. W. Rae, E. Elsen, and L. Sifre · 2021
Closest in time.
Fair ML tools require problematic ML models
J. Buckman · 2021
Closest in time.
Toward gender-inclusive coreference resolution: An analysis of gender and bias throughout the machine learning lifecyle
Y. T. Cao and H. Daumé · 2021
Closest in time.
Extracting training data from large language models, 2021
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel · 2021
Closest in time.
Documenting the english colossal clean crawled corpus
J. Dodge, M. Sap, A. Marasovic, W. Agnew, G. Ilharco, D. Groeneveld, and M. Gardner · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2021
Closest in time.
Reasons, values, stakeholders: a philosophical framework for explainable artificial intelligence
A. Kasirzadeh · 2021
Closest in time.
Z. Kenton, T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik, and G. Irving · 2021
Closest in time.
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, the World’s Largest and Most Powerful Generative Language Model
P. Kharya and A. Alvi · 2021
Closest in time.
Scalable and efficient moe training for multitask multilingual models, 2021
Y. J. Kim, A. A. Awan, A. Muzio, A. F. C. Salinas, L. Lu, A. Hendy, S. Rajbhandari, Y. He, and H. H. Awadalla · 2021
Closest in time.
A multi-level attention model for evidence-based fact checking
C. Kruengkrai, J. Yamagishi, and X. Wang · 2021
Closest in time.
Pitfalls of static language modelling
A. Lazaridou, A. Kuncoro, E. Gribovskaya, D. Agrawal, A. Liska, T. Terzi, M. Gimenez, C. d. M. d’Autume, S. Ruder, D. Yogatama, et al · 2021
Closest in time.
Towards few-shot fact-checking via perplexity
N. Lee, Y. Bang, A. Madotto, and P. Fung · 2021
Closest in time.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2021
Closest in time.
BASE layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer · 2021
Closest in time.
A systematic investigation of commonsense understanding in large language models
X. L. Li, A. Kuncoro, C. d. M. d’Autume, P. Blunsom, and A. Nematzadeh · 2021
Closest in time.
Jurassic-1: Technical details and evaluation
O. Lieber, O. Sharir, B. Lenz, and Y. Shoham · 2021
Closest in time.
Accelerating sparse deep neural networks, 2021
A. Mishra, J. A. Latorre, J. Pool, D. Stosic, D. Stosic, G. Venkatesh, C. Yu, and P. Micikevicius · 2021
Closest in time.
Carbon emissions and large neural network training
D. A. Patterson, J. Gonzalez, Q. V. Le, C. Liang, L. Munguia, D. Rothchild, D. R. So, M. Texier, and J. Dean · 2021
Closest in time.
AC/DC: alternating compressed/decompressed training of deep neural networks
A. Peste, E. Iofinova, A. Vladu, and D. Alistarh · 2021
Closest in time.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. A. Smith, and M. Lewis · 2021
Closest in time.
Recipes for building an open-domain chatbot
S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y. Liu, J. Xu, M. Ott, E. M. Smith, Y.-L. Boureau, and J. Weston · 2021
Closest in time.
Hatecheck: Functional tests for hate speech detection models
P. Röttger, B. Vidgen, D. Nguyen, Z. Waseem, H. Margetts, and J. Pierrehumbert · 2021
Closest in time.
Multitask prompted training enables zero-shot task generalization, 2021
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. L. Scao, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, S. Biderman, L. Gao, T. Bers, T. Wolf, and A. M. Rush · 2021
Closest in time.
Automap: Towards ergonomic automated parallelism for ml models
M. Schaarschmidt, D. Grewe, D. Vytiniotis, A. Paszke, G. Schmid, T. Norman, J. Molloy, J. Godwin, N. A. Rink, V. Nair, and D. Belov · 2021
Closest in time.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP
T. Schick, S. Udupa, and H. Schütze · 2021
Closest in time.
Societal biases in language generation: Progress and challenges
E. Sheng, K.-W. Chang, P. Natarajan, and N. Peng · 2021
Closest in time.
Primer: Searching for efficient transformers for language modeling, 2021
D. R. So, W. Mańke, H. Liu, Z. Dai, N. Shazeer, and Q. V. Le · 2021
Closest in time.
Updates and lessons from AI forecasting, 2021
J. Steinhardt · 2021
Closest in time.
Do long-range language models actually use long-range context?
S. Sun, K. Krishna, A. Mattarella-Micke, and M. Iyyer · 2021
Closest in time.
Putting humans in the natural language processing loop: A survey
Z. J. Wang, D. Choi, S. Xu, and D. Yang · 2021
Closest in time.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Closest in time.
Ethical and social risks of harm from language models
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel · 2021
Closest in time.
Challenges in detoxifying language models
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang · 2021
Closest in time.
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Closest in time.
Detoxifying language models risks marginalizing minority voices
A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein · 2021
Closest in time.
Bot-adversarial dialogue for safe conversational agents
J. Xu, D. Ju, M. Li, Y.-L. Boureau, J. Weston, and E. Dinan · 2021
Closest in time.
Automatically exposing problems with neural dialog models
D. Yu and K. Sagae · 2021
Closest in time.
Differentially private fine-tuning of language models
D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, et al · 2021
Closest in time.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Closest in time.