Fetching the paper…
Reading the bibliography…
Most humans use visual imagination to understand and reason about language, but models such as BERT reason about language using knowledge acquired during text-only pretraining.
A standard corpus of present-day edited american english, for use with digital computers
W Nelson Francis and Henry Kucera · 1964
Earlier work this paper cites.
The cloze test as a measure of english proficiency
Joseph Bartow Stubbs and G Richard Tucker · 1974
Earlier work this paper cites.
Age at onset of blindness and the development of the semantics of color names
Gloria Strauss Marmor · 1978
Earlier work this paper cites.
The cloze procedure and proficiency in english as a foreign language
J Charles Alderson · 1979
Earlier work this paper cites.
Half a century of research on the stroop effect: an integrative review
Colin M MacLeod · 1991
Earlier work this paper cites.
Representation of colors in the blind, color-blind, and normally sighted
Roger N Shepard and Lynn A Cooper · 1992
Earlier work this paper cites.
Is color an intrinsic property of object representation?
Galit Naor-Raz, Michael J Tarr, and Daniel Kersten · 2003
Earlier work this paper cites.
The role of color information on object recognition: A review and meta-analysis
Inês Bramão, Alexandra Reis, Karl Magnus Petersson, and Luís Faísca · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Function follows form: activation of shape and function features during object identification
Eiling Yee, Stacy Huffstetler, and Sharon L Thompson-Schill · 2011
Earlier work this paper cites.
Distributional semantics in technicolor
Elia Bruni, Gemma Boleda, Marco Baroni, and Nam-Khanh Tran · 2012
Earlier work this paper cites.
Decoding the yellow of a gray banana
Michael M Bannert and Andreas Bartels · 2013
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme · 2013
Earlier work this paper cites.
Why are abstract concepts hard to understand?
Paula J Schwanenflugel · 2013
Earlier work this paper cites.
Multimodal distributional semantics
Elia Bruni, Nam-Khanh Tran, and Marco Baroni · 2014
Earlier work this paper cites.
Concreteness ratings for 40 thousand generally known english word lemmas
Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman · 2014
Earlier work this paper cites.
Learning abstract concept embeddings from multi-modal data: Since you probably can’t see what i mean
Felix Hill and Anna Korhonen · 2014
Earlier work this paper cites.
Multi-modal models for concrete and abstract concept meaning
Felix Hill, Roi Reichart, and Anna Korhonen · 2014
Earlier work this paper cites.
Learning image embeddings using convolutional neural networks for improved multi-modal semantics
Douwe Kiela and Léon Bottou · 2014
Earlier work this paper cites.
The goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston · 2015
Earlier work this paper cites.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni · 2015
Earlier work this paper cites.
Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks
Xiao Lin and Devi Parikh · 2015
Earlier work this paper cites.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Visual word2vec (vis-w2v): Learning visually grounded word embeddings using abstract scenes
Satwik Kottur, Ramakrishna Vedantam, José MF Moura, and Devi Parikh · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
The shape of things: The origin of young children’s knowledge of the names and properties of geometric forms
Brian N Verdine, Kelsey R Lucca, Roberta M Golinkoff, Kathryn Hirsh-Pasek, and Nora S Newcombe · 2016
Earlier work this paper cites.
Structured matching for phrase localization
Mingzhe Wang, Mahmoud Azab, Noriyuki Kojima, Rada Mihalcea, and Jia Deng · 2016
Earlier work this paper cites.
Learning visually grounded sentence representations
Douwe Kiela, Alexis Conneau, Allan Jabri, and Maximilian Nickel · 2017
Earlier work this paper cites.
Reflections on sentiment/opinion analysis
Jiwei Li and Eduard Hovy · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Quantifying the visual concreteness of words and topics in multimodal datasets
Jack Hessel, David Mimno, and Lillian Lee · 2018
Cited alongside, same era.
Do neural language models overcome reporting bias?
Vered Shwartz and Yejin Choi · 2020
Later among the works it cites.
Ernie 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
Hao Tan and Mohit Bansal · 2020
Later among the works it cites.
Clozing the gap: How far do cloze items measure?
Jonathan Trace · 2020
Later among the works it cites.
Data efficient masked language modeling for vision and language
Yonatan Bitton, Gabriel Stanovsky, Michael Elhadad, and Roy Schwartz · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Illustrative language understanding: Large-scale visual grounding with image search
Jamie Kiros, William Chan, and Geoffrey Hinton · 2018
Cited alongside, same era.
Learning concept abstractness using weak supervision
Ella Rabinovich, Benjamin Sznajder, Artem Spector, Ilya Shnayderman, Ranit Aharonov, David Konopnicki, and Noam Slonim · 2018
Cited alongside, same era.
Colour envisioned: Concepts of colour in the blind and sighted
Armin Saysani, Michael C Corballis, and Paul M Corballis · 2018
Cited alongside, same era.
Predicting word concreteness and imagery
Jean Charbonnier and Christian Wartena · 2019
Cited alongside, same era.
Do neural language representations learn physical commonsense?
Maxwell Forbes, Ari Holtzman, and Yejin Choi · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Who’s waldo? linking people across text and images
Yuqing Cui, Apoorv Khandelwal, Yoav Artzi, Noah Snavely, and Hadar Averbuch-Elor · 2021
Later among the works it cites.
Effect of visual extensions on natural language understanding in vision-and-language models
Taichi Iki and Akiko Aizawa · 2021
Later among the works it cites.
Openclip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Later among the works it cites.
The world of an octopus: How reporting bias influences a language model’s perception of color
Cory Paik, Stéphane Aroca-Ouellette, Alessandro Roncone, and Katharina Kann · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Seeing colour through language: colour knowledge in the blind and sighted
Armin Saysani, Michael C Corballis, and Paul M Corballis · 2021
Later among the works it cites.
Visual grounding strategies for text-only natural language processing
Damien Sileo · 2021
Later among the works it cites.
How do blind people know that blue is cold? distributional semantics encode color-adjective associations
Jeroen van Paridon, Qiawen Liu, and Gary Lupyan · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao · 2021
Later among the works it cites.
What’s the best place for an ai conference, vancouver or _: Why completing comparative questions is difficult
Avishai Zagoury, Einat Minkov, Idan Szpektor, and William W Cohen · 2021
Later among the works it cites.
Vlp: A survey on vision-language pre-training
Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen, Jing Shi, Shuang Xu, and Bo Xu · 2022
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2022
Later among the works it cites.
Vision-language pre-training: Basics, recent advances, and future trends
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, and Jianfeng Gao · 2022
Later among the works it cites.
Flexible visual grounding
Yongmin Kim, Chenhui Chu, and Sadao Kurohashi · 2022
Later among the works it cites.
Weakly supervised face naming with symmetry-enhanced contrastive loss
Tingyu Qu, Tinne Tuytelaars, and Marie-Francine Moens · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
Visual commonsense in pretrained unimodal and multimodal models
Chenyu Zhang, Benjamin Van Durme, Zhuowan Li, and Elias Stengel-Eskin · 2022
Later among the works it cites.
Speech and language processing (3rd draft ed.), 2023
Dan Jurafsky and James H Martin · 2023
Closest in time.