Fetching the paper…
Reading the bibliography…
Generative Language Models gained significant attention in late 2022 / early 2023, notably with the introduction of models refined to act consistently with users' expectations of interactions with AI (conversational models).
Roberta: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 1907
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 1909
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 1910
Earlier work this paper cites.
Natural language processing with modular PDP networks and distributed lexicon
R. Miikkulainen and M. G. Dyer · 1991
Earlier work this paper cites.
Sequential neural text compression
J. Schmidhuber and S. Heil · 1996
Earlier work this paper cites.
Industry life cycles
S. Klepper · 1997
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, and P. Vincent · 2000
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
The emergence of emerging technologies
R. Adner and D. A. Levinthal · 2002
Earlier work this paper cites.
Continuous space language models for statistical machine translation
H. Schwenk, D. Déchelotte, and J. Gauvain · 2006
Earlier work this paper cites.
Calibration of google trends time series
R. West · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models
K. McGuffie and A. Newhouse · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. C. Courville, P. Manzagol, P. Vincent, and S. Bengio · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Technological revolutions and techno-economic paradigms
C. Perez · 2010
Earlier work this paper cites.
Diffusion of innovations
E. M. Rogers · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
F. Huszar · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. S. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
S. R. Bowman, L. Vilnis, O. Vinyals, A. M. Dai, R. Józefowicz, and S. Bengio · 2016
Earlier work this paper cites.
Machine learning with adversaries: Byzantine tolerant gradient descent
P. Blanchard, E. M. E. Mhamdi, R. Guerraoui, and J. Stainer · 2017
Earlier work this paper cites.
Pathnet: Evolution channels gradient descent in super neural networks
C. Fernando, D. Banarse, C. Blundell, Y. Zwols, D. Ha, A. A. Rusu, A. Pritzel, and D. Wierstra · 2017
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. B. Viégas, M. Wattenberg, G. Corrado, M. Hughes, and J. Dean · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, and I. Sutskever · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Ten years of research change using google trends: From the perspective of big data utilizations and applications
S.-P. Jun, H. S. Yoo, and S. Choi · 2018
Earlier work this paper cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Understanding new products’ market performance using google trends
P. Chumnumpan and X. Shi · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Defending against neural fake news
R. Zellers, A. Holtzman, H. Rashkin, Y. Bisk, A. Farhadi, F. Roesner, and Y. Choi · 2019
Cited alongside, same era.
Gender bias in contextualized word embeddings
J. Zhao, T. Wang, M. Yatskar, R. Cotterell, V. Ordonez, and K. Chang · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Digital transformations: new tools and methods for mining technological intelligence
T. Daim and H. Yalçin · 2022
Later among the works it cites.
New meta ai demo writes racist and inaccurate scientific literature, gets pulled
B. Edwards · 2022
Later among the works it cites.
Predictability and surprise in large generative models
D. Ganguli, D. Hernandez, L. Lovitt, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, N. Elhage, S. E. Showk, S. Fort, Z. Hatfield-Dodds, T. Henighan, S. Johnston, A. Jones, N. Joseph, J. Kernian, S. Kravec, B. Mann, N. Nanda, K. Ndousse, C. Olsson, D. Amodei, T. Brown, J. Kaplan, S. McCandlish, C. Olah, D. Amodei, and J. Clark · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trebacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, L. Campbell-Gillingham, J. Uesato, P. Huang, R. Comanescu, F. Yang, A. See, S. Dathathri, R. Greig, C. Chen, D. Fritz, J. S. Elias, R. Green, S. Mokrá, N. Fernando, B. Wu, R. Foley, S. Young, I. Gabriel, W. Isaac, J. Mellor, D. Hassabis, K. Kavukcuoglu, L. A. Hendricks, and G. Irving · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2020
Cited alongside, same era.
Reducing non-normative text generation from language models
X. Peng, S. Li, S. Frazier, and M. O. Riedl · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Take Blip’s us$ 100 million series A led by Warburg Pincus, October 2020
Warburg Pincus LLC · 2020
Cited alongside, same era.
Tracking activity in real time with google trends
N. Woloszko · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Cited alongside, same era.
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Later among the works it cites.
Unnatural instructions: Tuning language models with (almost) no human labor
O. Honovich, T. Scialom, O. Levy, and T. Schick · 2022
Later among the works it cites.
What you see is what you get: Principled deep learning via distributional generalization
B. Kulynych, Y.-Y. Yang, Y. Yu, J. Błasiok, and P. Nakkiran · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning, 2022
N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. L. Scao, M. S. Bari, S. Shen, Z.-X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Raff, and C. Raffel · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Red teaming language models with language models
E. Perez, S. Huang, H. F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Later among the works it cites.
Openalex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
J. Priem, H. Piwowar, and R. Orr · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
BLOOM: A 176b-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilic, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A. McMillan-Major, I. Beltagy, H. Nguyen, L. Saulnier, S. Tan, P. O. Suarez, V. Sanh, H. Laurençon, Y. Jernite, J. Launay, M. Mitchell, C. Raffel, A. Gokaslan, A. Simhi, A. Soroa, A. F. Aji, A. Alfassy, A. Rogers, A. K. Nitzav, C. Xu, C. Mou, C. Emezue, C. Klamm, C. Leong, D. van Strien, D. I. Adelani, and et al · 2022
Later among the works it cites.
Galactica: A large language model for science
R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Saravia, A. Poulton, V. Kerkez, and R. Stojnic · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, Y. Zhou, C. Chang, I. Krivokon, W. Rusch, M. Pickett, K. S. Meier-Hellstern, M. R. Morris, T. Doshi, R. D. Santos, T. Duke, J. Soraker, B. Zevenbergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. H. Chi, and Q. Le · 2022
Later among the works it cites.
Youtuber trains ai bot on 4chan’s pile o’ bile with entirely predictable results
J. Vincent · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2022
Later among the works it cites.
OPT: open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. T. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer · 2022
Later among the works it cites.
Historical company data, 2022
‘Crunchbase, Inc.’ · 2022
Later among the works it cites.
Microsoft invests $10 billion in chatgpt maker openai, January 2023
D. Bass · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi · 2023
Closest in time.
Our biggest sponsor pulled out - wan show february 10, 2023, Feb. 2023
LinusTechTips · 2023
Closest in time.
Reinventing search with a new ai-powered microsoft bing and edge, your copilot for the web, February 2023a
Y. Mehdi · 2023
Closest in time.
Jailbreaking chatgpt on release day, 2022
Z. Mowshowitz · 2023
Closest in time.
OpenAI · 2023
Closest in time.
What makes a dialog agent useful?
N. Rajani, N. Lambert, V. Sanh, and T. Wolf · 2023
Closest in time.
Introducing microsoft 365 copilot – your copilot for work, March 2023
J. Spataro · 2023
Closest in time.
Announcing openchatkit, March 2023
Together · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Microsoft’s bing is an emotionally manipulative liar, and people love it
J. Vincent · 2023
Closest in time.
Historical citation data, 2023
‘OurResearch, Org.’ · 2023
Closest in time.