Fetching the paper…
Reading the bibliography…
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H · 1905
Earlier work this paper cites.
MASS: Masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 1905
Earlier work this paper cites.
WinoGrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Le Bras, R., Bhagavatula, C., and Choi, Y · 1907
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings
Dev, S., Li, T., Phillips, J. M., and Srikumar, V · 1908
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 1911
Earlier work this paper cites.
Gender bias in coreference resolution
Rudinger, R., Naradowsky, J., Leonard, B., and Van Durme, B · 2002
Earlier work this paper cites.
GLU variants improve transformer
Shazeer, N · 2002
Earlier work this paper cites.
NLTK: The natural language toolkit
Bird, S. and Loper, E · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Unsupervised translation of programming languages
Lachaux, M., Rozière, B., Chanussot, L., and Lample, G · 2006
Earlier work this paper cites.
The invisible whiteness of being: Whiteness, white supremacy, white privilege, and racism
Sue, D. W · 2006
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-Kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Chaudhary, V., Goyal, N., Birch, T., Liptchinsky, V., Edunov, S., Grave, E., Auli, M., and Joulin, A · 2010
Earlier work this paper cites.
Complete multilingual neural machine translation
Freitag, M. and Firat, O · 2010
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
The Winograd Schema Challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Pascanu, R., Mikolov, T., and Bengio, Y · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
An analysis of patch plausibility and correctness for generate-and-validate patch generation systems
Qi, Z., Long, F., Achour, S., and Rinard, M · 2015
Earlier work this paper cites.
Jupiter rising: A decade of clos topologies and centralized control in google’s datacenter network
Singh, A., Ong, J., Agarwal, A., Anderson, G., Armistead, A., Bannon, R., Boving, S., Desai, G., Felderman, B., Germano, P., et al · 2015
Earlier work this paper cites.
MAWPS: A math word problem repository
Koncel-Kedziorski, R., Roy, S., Amini, A., Kushman, N., and Hajishirzi, H · 2016
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
Mostafazadeh, N., Chambers, N., He, X., Parikh, D., Batra, D., Vanderwende, L., Kohli, P., and Allen, J · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Creating training corpora for nlg micro-planners
Gardent, C., Shimorina, A., Narayan, S., and Perez-Beltrachini, L · 2017
Earlier work this paper cites.
Deepfix: Fixing common C language errors by deep learning
Gupta, R., Pal, S., Kanade, A., and Shevade, S. K · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
DéjàVu: a map of code duplicates on GitHub
Lopes, C. V., Maj, P., Martins, P., Saini, V., Yang, D., Zitny, J., Sajnani, H., and Vitek, J · 2017
Earlier work this paper cites.
The E2E dataset: New challenges for end-to-end generation
Novikova, J., Dušek, O., and Rieser, V · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
A survey of machine learning for big code and naturalness
Allamanis, M., Barr, E. T., Devanbu, P., and Sutton, C · 2018
Earlier work this paper cites.
JAX: Composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
QuAC : Question answering in context
Choi, E., He, H., Iyyer, M., Yatskar, M., Yih, W., Choi, Y., Liang, P., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Dixon, L., Li, J., Sorensen, J., Thain, N., and Vasserman, L · 2018
Earlier work this paper cites.
Understanding back-translation at scale
Edunov, S., Ott, M., Auli, M., and Grangier, D · 2018
Earlier work this paper cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Reddy, S., Chen, D., and Manning, C. D · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Mesh-TensorFlow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V · 2018
Earlier work this paper cites.
The adverse effects of code duplication in machine learning models of code
Allamanis, M · 2019
Earlier work this paper cites.
Tagged back-translation
Caswell, I., Chelba, C., and Grangier, D · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Neural generation for czech: Data and baselines
Dusek, O. and Jurvc’ivcek, F · 2019
Earlier work this paper cites.
Neural Generation for Czech: Data and Baselines
Dušek, O. and Jurčíček, F · 2019
Earlier work this paper cites.
Semantic Noise Matters for Neural Natural Language Generation
Dušek, O., Howcroft, D. M., and Rieser, V · 2019
Cited alongside, same era.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Cited alongside, same era.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Cited alongside, same era.
SPoC: Search-based pseudocode to code
Kulal, S., Pasupat, P., Chandra, K., Lee, M., Padon, O., Aiken, A., and Liang, P · 2019
Cited alongside, same era.
Quantifying social biases in contextual word representations
Kurita, K., Vyas, N., Pareek, A., Black, A. W., and Tsvetkov, Y · 2019
Cited alongside, same era.
Natural Questions: A benchmark for question answering research
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Later among the works it cites.
Beyond the imitation game: Measuring and extrapolating the capabilities of language models
BIG-bench collaboration · 2021
Later among the works it cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R. and et. al., D. A. H · 2021
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Driessche, G. v. d., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Cited alongside, same era.
The NiuTrans machine translation systems for WMT19
Li, B., Li, Y., Xu, C., Lin, Y., Liu, J., Liu, H., Wang, Z., Zhang, Y., Xu, N., Wang, Z., Feng, K., Chen, H., Liu, T., Li, Y., Wang, Q., Xiao, T., and Zhu, J · 2019
Cited alongside, same era.
Model cards for model reporting
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T · 2019
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Sap, M., Card, D., Gabriel, S., Choi, Y., and Smith, N. A · 2019
Cited alongside, same era.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Cited alongside, same era.
Evaluating gender bias in machine translation
Stanovsky, G., Smith, N. A., and Zettlemoyer, L · 2019
Cited alongside, same era.
Later among the works it cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Ponde, H., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Later among the works it cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Later among the works it cites.
Harms of gender exclusivity and challenges in non-binary representation in language technologies
Dev, S., Monajatipoor, M., Ovalle, A., Subramonian, A., Phillips, J., and Chang, K.-W · 2021
Later among the works it cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Later among the works it cites.
GLaM: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., III, H. D., and Crawford, K · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Gehrmann, S., Adewumi, T., Aggarwal, K., Ammanamanchi, P. S., Aremu, A., Bosselut, A., Chandu, K. R., Clinciu, M.-A., Das, D., Dhole, K., Du, W., Durmus, E., Dušek, O., Emezue, C. C., Gangal, V., Garbacea, C., Hashimoto, T., Hou, Y., Jernite, Y., Jhamtani, H., Ji, Y., Jolly, S., Kale, M., Kumar, D., Ladhak, F., Madaan, A., Maddela, M., Mahajan, K., Mahamood, S., Majumder, B. P., Martins, P. H., McMillan-Major, A., Mille, S., van Miltenburg, E., Nadeem, M., Narayan, S., Nikolaev, V., Niyongabo Rubungo, A., Osei, S., Parikh, A., Perez-Beltrachini, L., Rao, N. R., Raunak, V., Rodriguez, J. D., Santhanam, S., Sedoc, J., Sellam, T., Shaikh, S., Shimorina, A., Sobrevilla Cabezudo, M. A., Strobelt, H., Subramani, N., Xu, W., Yang, D., Yerukola, A., and Zhou, J · 2021
Later among the works it cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J · 2021
Later among the works it cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Later among the works it cites.
Measurement and fairness
Jacobs, A. Z. and Wallach, H · 2021
Later among the works it cites.
Mwptoolkit: An open-source framework for deep learning-based math word problem solvers
Lan, Y., Wang, L., Zhang, Q., Lan, Y., Dai, B. T., Wang, Y., Zhang, D., and Lim, E.-P · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2021
Later among the works it cites.
Jurassic-1: Technical details and evaluation
Lieber, O., Sharir, O., Lenz, B., and Shoham, Y · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Later among the works it cites.
Are NLP models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Later among the works it cites.
Carbon emissions and large neural network training
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., and Dean, J · 2021
Later among the works it cites.
An empirical cybersecurity evaluation of GitHub Copilot’s code contributions
Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R · 2021
Later among the works it cites.
Measuring and improving BERT’s mathematical abilities by predicting the order of reasoning
Piekos, P., Malinowski, M., and Michalewski, H · 2021
Later among the works it cites.
Learning compact metrics for MT
Pu, A., Chung, H. W., Parikh, A. P., Gehrmann, S., and Sellam, T · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training Gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, H. F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S. M., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B. A., Weidinger, L., Gabriel, I., Isaac, W. S., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Later among the works it cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y · 2021
Later among the works it cites.
AI and the everything in the whole wide world benchmark
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A · 2021
Later among the works it cites.
{ \{ ZeRO-Offload } \} : Democratizing { \{ Billion-Scale } \} model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Later among the works it cites.
Prompt programming for large language models: Beyond the few-shot paradigm
Reynolds, L. and McDonell, K · 2021
Later among the works it cites.
HateCheck: Functional tests for hate speech detection models
Röttger, P., Vidgen, B., Nguyen, D., Waseem, Z., Margetts, H., and Pierrehumbert, J · 2021
Later among the works it cites.
Re-imagining algorithmic fairness in India and beyond
Sambasivan, N., Arnesen, E., Hutchinson, B., Doshi, T., and Prabhakaran, V · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Later among the works it cites.
Societal biases in language generation: Progress and challenges
Sheng, E., Chang, K., Natarajan, P., and Peng, N · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W. S., Legassick, S., Irving, G., and Gabriel, I · 2021
Later among the works it cites.
Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Later among the works it cites.
GSPMD: general and scalable parallelization for ml computation graphs
Xu, Y., Lee, H., Chen, D., Hechtman, B., Huang, Y., Joshi, R., Krikun, M., Lepikhin, D., Ly, A., Maggioni, M., Pang, R., Shazeer, N., Wang, S., Wang, T., Wu, Y., and Chen, Z · 2021
Later among the works it cites.
Break-it-fix-it: Unsupervised learning for program repair
Yasunaga, M. and Liang, P · 2021
Later among the works it cites.
Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., Jiang, X., Yang, Z., Wang, K., Zhang, X., et al · 2021
Later among the works it cites.
Pathways: Asynchronous distributed dataflow for ML
Barham, P., Chowdhery, A., Dean, J., Ghemawat, S., Hand, S., Hurt, D., Isard, M., Lim, H., Pang, R., Roy, S., Saeta, B., Schuh, P., Sepassi, R., Shafey, L. E., Thekkath, C. A., and Wu, Y · 2022
Closest in time.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2022
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Gehrmann, S., Clark, E., and Sellam, T · 2022
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Closest in time.
Deduplicating training data mitigates privacy risks in language models
Kandpal, N., Wallace, E., and Raffel, C · 2022
Closest in time.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Closest in time.
Competition-level code generation with alphacode, Feb 2022
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D., Sutherland Robson, E., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O · 2022
Closest in time.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Closest in time.
Reasoning like program executors
Pi, X., Liu, Q., Chen, B., Ziyadi, M., Lin, Z., Gao, Y., Fu, Q., Lou, J.-G., and Chen, W · 2022
Closest in time.
Scaling up models and data with t5x
Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A · 2022
Closest in time.
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al · 2022
Closest in time.
Sustainability at Google.Carbon neutral since 2007.Carbon free by 2030., 2022
Sustainability, G · 2022
Closest in time.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Closest in time.
Designing effective sparse expert models
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Closest in time.