How neural networks extrapolate: From feedforward to graph neural networks
Keyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du, Ken-Ichi Kawarabayashi, and Stefanie Jegelka · 2021
Later among the works it cites.
Deep neural networks and tabular data: A survey
Original
Vadim Borisov, Tobias Leemann, Kathrin Seßler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci · 2021
Later among the works it cites.
Prompt programming for large language models: Beyond the few-shot paradigm
Laria Reynolds and Kyle McDonell · 2021
Later among the works it cites.
Cross-attention is all you need: Adapting pretrained transformers for machine translation
Original
Mozhdeh Gheini, Xiang Ren, and Jonathan May · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill · 2021
Later among the works it cites.
Tapex: table pre-training via learning a neural sql executor
Original
Qian Liu, Bei Chen, Jiaqi Guo, Zeqi Lin, and Jian-guang Lou · 2021
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Original
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al · 2021
Later among the works it cites.
Exploring the limits of large scale pre-training
Original
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Original
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Training verifiers to solve math word problems
Original
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Later among the works it cites.
End-to-end training of multi-document reader and retriever for open-domain question answering
Devendra Singh, Siva Reddy, Will Hamilton, Chris Dyer, and Dani Yogatama · 2021
Later among the works it cites.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Original
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Gpt understands, too
Original
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang · 2021
Later among the works it cites.
Towards table-to-text generation with numerical reasoning
Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura · 2021
Later among the works it cites.
Xgpt: Cross-modal generative pre-training for image captioning
Qiaolin Xia, Haoyang Huang, Nan Duan, Dongdong Zhang, Lei Ji, Zhifang Sui, Edward Cui, Taroon Bharti, and Ming Zhou · 2021
Later among the works it cites.
Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Original
Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
Pretrained language models are symbolic mathematics solvers too!
Original
Kimia Noorbakhsh, Modar Sulaiman, Mahdi Sharifi, Kallol Roy, and Pooyan Jamshidi · 2021
Later among the works it cites.
Tabbie: Pretrained representations of tabular data
Original
Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer · 2021
Later among the works it cites.
Investigating the limitations of transformers with simple arithmetic tasks
Original
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2021
Later among the works it cites.
Cross-lingual intermediate fine-tuning improves dialogue state tracking
Original
Nikita Moghe, Mark Steedman, and Alexandra Birch · 2021
Later among the works it cites.
What makes good in-context examples for gpt- 3 3 ?
Original
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen · 2021
Later among the works it cites.
How many data points is a prompt worth?
Original
Teven Le Scao and Alexander M Rush · 2021
Later among the works it cites.
Text generation with efficient (soft) q-learning
Original
Han Guo, Bowen Tan, Zhengzhong Liu, Eric P Xing, and Zhiting Hu · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Original
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Original
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Original
Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang · 2021
Later among the works it cites.
Deep neural networks and tabular data: A survey
Original
Vadim Borisov, Tobias Leemann, Kathrin Seßler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci · 2021
Later among the works it cites.
Tabular transformers for modeling multivariate time series
Inkit Padhi, Yair Schiff, Igor Melnyk, Mattia Rigotti, Youssef Mroueh, Pierre Dognin, Jerret Ross, Ravi Nair, and Erik Altman · 2021
Later among the works it cites.
Regularization is all you need: Simple neural nets can excel on tabular data
Original
Arlind Kadra, Marius Lindauer, Frank Hutter, and Josif Grabocka · 2021
Later among the works it cites.
Tabnet: Attentive interpretable tabular learning
Sercan O Arık and Tomas Pfister · 2021
Later among the works it cites.
Saint: Improved neural networks for tabular data via row attention and contrastive pre-training
Original
Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild, C Bayan Bruss, and Tom Goldstein · 2021
Later among the works it cites.
Fairbatch: Batch selection for model fairness
Yuji Roh, Kangwook Lee, Steven Euijong Whang, and Changho Suh · 2021
Later among the works it cites.
Can neural nets learn the same model twice? investigating reproducibility and double descent from the decision boundary perspective
Original
Gowthami Somepalli, Liam Fowl, Arpit Bansal, Ping Yeh-Chiang, Yehuda Dar, Richard Baraniuk, Micah Goldblum, and Tom Goldstein · 2022
Closest in time.
Codeparrot, 2022
Loubna Allal, Leandro Werra, Thomas Wolf, and Li Jia · 2022
Closest in time.
A conversational paradigm for program synthesis
Original
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Closest in time.
Diffusion models in vision: A survey
Original
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah · 2022
Closest in time.
Rethinking the role of demonstrations: What makes in-context learning work?
Original
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Closest in time.
Improving in-context few-shot learning via self-supervised training
Original
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srini Iyer, Veselin Stoyanov, and Zornitsa Kozareva · 2022
Closest in time.
Genlabel: Mixup relabeling using generative models
Original
Jy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon, Dimitris Papailiopoulos, and Kangwook Lee · 2022
Closest in time.
Improved input reprogramming for gan conditioning
Original
Tuan Dinh, Daewon Seo, Zhixu Du, Liang Shang, and Kangwook Lee · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2022
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, et al · 2022
Closest in time.
Unveiling transformers with lego: a synthetic reasoning task, 2022
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Closest in time.
Language models are general-purpose interfaces
Original
Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei · 2022
Closest in time.
Flamingo: a visual language model for few-shot learning
Original
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Closest in time.
A generalist agent
Original
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
Original
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Closest in time.
Large language models are zero-shot reasoners
Original
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Closest in time.
When to make exceptions: Exploring language models as accounts of human moral judgment
Original
Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf · 2022
Closest in time.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
A survey of pretrained language models based text generation
Original
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2022
Closest in time.
Pre-trained language models for interactive decision-making
Original
Shuang Li, Xavier Puig, Yilun Du, Clinton Wang, Ekin Akyurek, Antonio Torralba, Jacob Andreas, and Igor Mordatch · 2022
Closest in time.
Regression transformer: Concurrent conditional generation and regression by blending numerical and textual tokens
Original
Jannis Born and Matteo Manica · 2022
Closest in time.
Tabular data: Deep learning is not all you need
Ravid Shwartz-Ziv and Amitai Armon · 2022
Closest in time.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Closest in time.
Debiasing pre-trained language models via efficient fine-tuning
Original
Michael Gira, Ruisu Zhang, and Kangwook Lee · 2022
Closest in time.