Fetching the paper…
Reading the bibliography…
Scaling up the size and training of autoregressive language models has enabled novel ways of solving Natural Language Processing tasks using zero-shot and few-shot learning.
“Bleu: a method for automatic evaluation of machine translation”
Kishore Papineni, Salim Roukos, Todd Ward and Wei-Jing Zhu · 2002
Earlier work this paper cites.
“Rouge: A package for automatic evaluation of summaries”
Chin-Yew Lin · 2004
Earlier work this paper cites.
“Findings of the 2014 workshop on statistical machine translation”
Ondřej Bojar et al · 2014
Earlier work this paper cites.
“Pointer sentinel mixture models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“SQuAD: 100,000+ questions for machine comprehension of text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
Shashi Narayan, Shay Cohen and Mirella Lapata · 2018
Earlier work this paper cites.
“A Call for Clarity in Reporting BLEU Scores”
Matt Post · 2018
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Massively multilingual neural machine translation in the wild: Findings and challenges”
Naveen Arivazhagan et al · 2019
Earlier work this paper cites.
“Ccnet: Extracting high quality monolingual datasets from web crawl data”
Guillaume Wenzek et al · 2019
Earlier work this paper cites.
“Plug and play language models: A simple approach to controlled text generation”
Sumanth Dathathri et al · 2019
Earlier work this paper cites.
“Ctrl: A conditional transformer language model for controllable generation”
Nitish Keskar et al · 2019
Earlier work this paper cites.
“Flaubert: Unsupervised language model pre-training for french”
Hang Le et al · 2019
Earlier work this paper cites.
“ftfy” Version 5.5, Zenodo, 2019
Robyn Speer · 2019
Cited alongside, same era.
“CamemBERT: a tasty french language model”
Louis Martin et al · 2019
Cited alongside, same era.
“Language models are few-shot learners”
Tom Brown et al · 2020
Cited alongside, same era.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Cited alongside, same era.
“RealToxicityPrompts: Evaluating neural toxic degeneration in language models”
Samuel Gehman et al · 2020
Cited alongside, same era.
“On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”
Emily Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Later among the works it cites.
“Quality at a glance: An audit of web-crawled multilingual datasets”
Isaac Caswell et al · 2021
Later among the works it cites.
“Challenges in detoxifying language models”
Johannes Welbl et al · 2021
Later among the works it cites.
“The power of scale for parameter-efficient prompt tuning”
Brian Lester, Rami Al-Rfou and Noah Constant · 2021
Later among the works it cites.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
Ben Wang and Aran Komatsuzaki · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linting Xue et al · 2020
Cited alongside, same era.
“Detoxify”, https://github.com/unitaryai/detoxify , 2020
Laura Hanu and Unitary team · 2020
Cited alongside, same era.
“BARThez: a skilled pretrained french sequence-to-sequence model”
Moussa Eddine, Antoine-P Tixier and Michalis Vazirgiannis · 2020
Cited alongside, same era.
“FQuAD: French question answering dataset”
Martin d’Hoffschmidt et al · 2020
Cited alongside, same era.
“Very deep transformers for neural machine translation”
Xiaodong Liu, Kevin Duh, Liyuan Liu and Jianfeng Gao · 2020
Cited alongside, same era.
“Facebook AI WMT21 news translation task submission”
Chau Tran et al · 2021
Cited alongside, same era.
“Un modèle Transformer Génératif Pré-entrainé pour le _ français”
Antoine Simoulin and Benoit Crabbé · 2021
Cited alongside, same era.
Later among the works it cites.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2021
Later among the works it cites.
“Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
Ben Wang · 2021
Later among the works it cites.
“A framework for few-shot language model evaluation”
Leo Gao et al · 2021
Later among the works it cites.
“On the Sizes of OpenAI API Models”, https://blog.eleuther.ai/gpt3-model-sizes/ , 2021
Leo Gao · 2021
Later among the works it cites.
“Deduplicating training data makes language models better”
Katherine Lee et al · 2021
Later among the works it cites.
“Perplexity of fixed-length models” Accessed: 2022-02-04, https://huggingface.co/docs/transformers/perplexity
2022
Closest in time.
“Training language models to follow instructions with human feedback”, https://openai.com/blog/instruction-following/ , 2022
Long Ouyang et al · 2022
Closest in time.