Fetching the paper…
Reading the bibliography…
Human brains integrate linguistic and perceptual information simultaneously to understand natural language, and hold the critical ability to render imaginations.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
A dual coding view of imagery and verbal processes in reading comprehension
Mark Sadoski and A. Paivio. 1994 · 1994
Earlier work this paper cites.
Univilm: A unified video and language pre-training model for multimodal understanding and generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Xilin Chen, and Ming Zhou. 2020 · 2002
Earlier work this paper cites.
Object classification from a single example utilizing class relevance metrics
Michael Fink. 2004 · 2004
Earlier work this paper cites.
Imagery in sentence comprehension: an fmri study
M. Just, S. Newman, T. Keller, A. McEleney, and P. Carpenter. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
One-shot learning of object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona. 2006 · 2006
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
Michael Grubinger, Paul D. Clough, Henning Müller, and Thomas Deselaers. 2006 · 2006
Earlier work this paper cites.
Semantic textual similarity benchmark
Eneko Agirre, Llu’is M‘arquez, and Richard Wicentowski. 2007 · 2007
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, A. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Image-mediated learning for zero-shot cross-lingual document retrieval
Ruka Funaki and Hideki Nakayama. 2015 · 2015
Earlier work this paper cites.
Visual bilingual lexicon induction with transferred convnet features
Douwe Kiela, Ivan Vulic, and Stephen Clark. 2015 · 2015
Earlier work this paper cites.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni. 2015 · 2015
Earlier work this paper cites.
Resolving language and vision ambiguities together: Joint segmentation & prepositional attachment resolution in captioned scenes
Gordon Christie, Ankit Laddha, Aishwarya Agrawal, Stanislaw Antol, Yash Goyal, Kevin Kochersberger, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Multi-modal representations for improved bilingual lexicon learning
Ivan Vulic, Douwe Kiela, Stephen Clark, and Marie-Francine Moens. 2016 · 2016
Cited alongside, same era.
Imagined visual representations as multimodal embeddings
Guillem Collell, Ted Zhang, and Marie-Francine Moens. 2017 · 2017
Cited alongside, same era.
First quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornel Csernai. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
Jiuxiang Gu, Jianfei Cai, Shafiq R. Joty, Li Niu, and G. Wang. 2018 · 2018
Cited alongside, same era.
Learning visually grounded sentence representations
Douwe Kiela, Alexis Conneau, A. Jabri, and Maximilian Nickel. 2018 · 2018
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Vokenization: Improving language understanding via contextualized, visually-grounded supervision
Haochen Tan and Mohit Bansal. 2020 · 2020
Later among the works it cites.
CLUE: A Chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan. 2020 · 2020
Later among the works it cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Incorporating visual semantics into sentence representations within a grounded space
Patrick Bordes, Eloi Zablocki, Laure Soulier, Benjamin Piwowarski, and Patrick Gallinari. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Induction networks for few-shot text classification
Ruiying Geng, Binhua Li, Yongbin Li, Xiaodan Zhu, Ping Jian, and Jian Sun. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Meta-learning adversarial domain adaptation network for few-shot text classification
Chengcheng Han, Zeqiu Fan, Dongxiang Zhang, Minghui Qiu, Ming Gao, and Aoying Zhou. 2021 · 2021
Later among the works it cites.
Generative imagination elevates machine translation
Quanyu Long, Mingxuan Wang, and Lei Li. 2021 · 2021
Later among the works it cites.
Dreca: A general task augmentation strategy for few-shot natural language inference
Shikhar Murty, Tatsunori B. Hashimoto, and Christopher D. Manning. 2021 · 2021
Later among the works it cites.
Visual and linguistic semantic representations are aligned at the border of human visual cortex
Sara F Popham, Alexander G Huth, Natalia Y Bilenko, Fatma Deniz, James S Gao, Anwar O Nunez-Elizalde, and Jack L Gallant. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Knowledge guided metric learning for few-shot text classification
Dianbo Sui, Yubo Chen, Binjie Mao, Delai Qiu, Kang Liu, and Jun Zhao. 2021 · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Later among the works it cites.
Few-shot text classification with triplet networks, data augmentation, and curriculum learning
Jason Wei, Chengyu Huang, Soroush Vosoughi, Yu Cheng, and Shiqi Xu. 2021 · 2021
Later among the works it cites.
Imagine: An imagination-based automatic evaluation metric for natural language generation
Wanrong Zhu, Xin Eric Wang, An Yan, Miguel P. Eckstein, and William Yang Wang. 2021 · 2021
Later among the works it cites.
A robustly optimized BERT pre-training approach with post-training
Liu Zhuang, Lin Wayne, Shi Ya, and Zhao Jun. 2021 · 2021
Later among the works it cites.
Vqgan-clip: Open domain image generation and editing with natural language guidance
Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. 2022 · 2022
Closest in time.
Things not written in text: Exploring spatial commonsense from visual signals
Xiao Liu, Da Yin, Yansong Feng, and Dongyan Zhao. 2022 · 2022
Closest in time.