Fetching the paper…
Reading the bibliography…
Having a rich multimodal inner language is an important component of human intelligence that enables several necessary core cognitive functions such as multimodal prediction, translation, and generation.
Distributional structure
Zellig S Harris · 1954
Earlier work this paper cites.
Procedures as a representation for data in a computer program for understanding natural language
Terry Winograd · 1971
Earlier work this paper cites.
Hearing lips and seeing voices
Harry McGurk and John MacDonald · 1976
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
The design for the wall street journal-based csr corpus
Douglas B Paul and Janet Baker · 1992
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller · 1995
Earlier work this paper cites.
Models of consciousness and memory
Morris Moscovitch · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Visual language
Robert E Horn · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The generative lexicon
James Pustejovsky · 1998
Earlier work this paper cites.
Origins of theory of mind, cognition and communication
Andrew N Meltzoff · 1999
Earlier work this paper cites.
Image retrieval: Current techniques, promising directions, and open issues
Yong Rui, Thomas S Huang, and Shih-Fu Chang · 1999
Earlier work this paper cites.
Audio-visual speech modeling for continuous speech recognition
Stéphane Dupont and Juergen Luettin · 2000
Earlier work this paper cites.
Affective computing
Rosalind W Picard · 2000
Earlier work this paper cites.
Crossmodal processing in the human brain: insights from functional neuroimaging studies
Gemma A Calvert · 2001
Earlier work this paper cites.
Audio content analysis for online audiovisual data segmentation and classification
Tong Zhang and C-C Jay Kuo · 2001
Earlier work this paper cites.
Application of haptic feedback to robotic surgery
Brian T Bethea, Allison M Okamura, Masaya Kitagawa, Torin P Fitton, Stephen M Cattaneo, Vincent L Gott, William A Baumgartner, and David D Yuh · 2004
Earlier work this paper cites.
Glossary of linguistic terms
Eugene Emil Loos · 2004
Earlier work this paper cites.
Modeling multimodal human-computer interaction
Zeljko Obrenovic and Dusan Starcevic · 2004
Earlier work this paper cites.
Encyclopedia of language and linguistics , volume 1
Keith Brown · 2005
Earlier work this paper cites.
Multisensory interaction: Real and virtual
Dinesh K Pai · 2005
Earlier work this paper cites.
Large-scale concept ontology for multimedia
Milind Naphade, John R Smith, Jelena Tesic, Shih-Fu Chang, Winston Hsu, Lyndon Kennedy, Alexander Hauptmann, and Jon Curtis · 2006
Earlier work this paper cites.
Artificial intelligence and consciousness
Antonio Chella and Riccardo Manzotti · 2007
Earlier work this paper cites.
Enhanced max margin learning on multimodal data mining in a multimedia database
Zhen Guo, Zhongfei Zhang, Eric Xing, and Christos Faloutsos · 2007
Earlier work this paper cites.
Artificial intelligence and consciousness
Drew McDermott · 2007
Earlier work this paper cites.
Models of consciousness
Anil Seth · 2007
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2008
Earlier work this paper cites.
Multisensory integration: current issues from the perspective of the single neuron
Barry E Stein and Terrence R Stanford · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Multimodal interfaces: A survey of principles, models and frameworks
Bruno Dumas, Denis Lalanne, and Sharon Oviatt · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Text-to-speech synthesis
Paul Taylor · 2009
Earlier work this paper cites.
On the classification of emotional biosignals evoked while viewing affective pictures: an integrated data-mining-based approach for healthcare applications
Christos A Frantzidis, Charalampos Bratsas, Manousos A Klados, Evdokimos Konstantinidis, Chrysa D Lithari, Ana B Vivas, Christos L Papadelis, Eleni Kaldoudi, Costas Pappas, and Panagiotis D Bamidis · 2010
Earlier work this paper cites.
Multimodal images in the brain
Stephen M Kosslyn, Giorgio Ganis, and William L Thompson · 2010
Earlier work this paper cites.
Features for content-based audio retrieval
Dalibor Mitrović, Matthias Zeppelzauer, and Christian Breiteneder · 2010
Earlier work this paper cites.
Emotion recognition and adaptation in spoken dialogue systems
Johannes Pittermann, Angela Pittermann, and Wolfgang Minker · 2010
Earlier work this paper cites.
The multifaceted interplay between attention and multisensory integration
Durk Talsma, Daniel Senkowski, Salvador Soto-Faraco, and Marty G Woldorff · 2010
Earlier work this paper cites.
Early language learning and literacy: neuroscience implications for education
Patricia K Kuhl · 2011
Earlier work this paper cites.
The neural bases of multisensory processes
Micah M Murray and Mark T Wallace · 2011
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
A normalization model of multisensory integration
Tomokazu Ohshiro, Dora E Angelaki, and Gregory C DeAngelis · 2011
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Richard Socher, Milind Ganjoo, Christopher D Manning, and Andrew Ng · 2013
Earlier work this paper cites.
Aspects of the Theory of Syntax , volume 11
Noam Chomsky · 2014
Earlier work this paper cites.
Dreams, reality and memory: confabulations in lucid dreamers implicate reality-monitoring dysfunction in dream consciousness
PR Corlett, SV Canavan, L Nahum, F Appah, and PT Morgan · 2014
Earlier work this paper cites.
Notes on noise contrastive estimation and negative sampling
Chris Dyer · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Cited alongside, same era.
Learning image embeddings using convolutional neural networks for improved multi-modal semantics
Douwe Kiela and Léon Bottou · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
The sensory construction of dreams and nightmare frequency in congenitally blind and late blind individuals
Amani Meaidi, Poul Jennum, Maurice Ptito, and Ron Kupers · 2014
Cited alongside, same era.
Multimodal mental imagery
Bence Nanay · 2018
Later among the works it cites.
Grounding language for transfer in deep reinforcement learning
Karthik Narasimhan, Regina Barzilay, and Tommi Jaakkola · 2018
Later among the works it cites.
Reptile: a scalable metalearning algorithm
Alex Nichol and John Schulman · 2018
Later among the works it cites.
Semantic correspondence: A hierarchical approach
Akila Pemasiri, Kien Nguyen, Sridha Sridharan, and Clinton Fookes · 2018
Later among the works it cites.
The tangled roots of inner speech, voices and delusions
Cherise Rosen, Simon McCarthy-Jones, Kayla A Chase, Clara S Humpston, Jennifer K Melbourne, Leah Kling, and Rajiv P Sharma · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Detecting sarcasm from students’ feedback in twitter
Nabeela Altrabsheh, Mihaela Cocea, and Sanaz Fallahkhair · 2015
Cited alongside, same era.
Multi-and cross-modal semantics beyond vision: Grounding in auditory perception
Douwe Kiela and Stephen Clark · 2015
Cited alongside, same era.
Deep convolutional inverse graphics network
Tejas D Kulkarni, William F Whitney, Pushmeet Kohli, and Josh Tenenbaum · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Cited alongside, same era.
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Later among the works it cites.
Centralnet: a multilayer approach for multimodal fusion, 2018
Valentin Vielzeuf, Alexis Lechervy, Stéphane Pateux, and Frédéric Jurie · 2018
Later among the works it cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
Later among the works it cites.
Multi-modal haptic feedback for grip force reduction in robotic surgery
Ahmad Abiri, Jake Pensa, Anna Tao, Ji Ma, Yen-Yi Juo, Syed J Askari, James Bisley, Jacob Rosen, Erik P Dutson, and Warren S Grundfest · 2019
Later among the works it cites.
Overview of artificial intelligence in medicine
Paras Malik Amisha, Monika Pathania, and Vyas Kumar Rathaur · 2019
Later among the works it cites.
Towards multimodal sarcasm detection (an _obviously_ perfect paper)
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Ur-funny: A multimodal language dataset for understanding humor
Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, and Mohammed Ehsan Hoque · 2019
Later among the works it cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Later among the works it cites.
Embedded multimodal interfaces in robotics: applications, future trends, and societal implications
Elsa A Kirchner, Stephen H Fairclough, and Frank Kirchner · 2019
Later among the works it cites.
Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks
Michelle A Lee, Yuke Zhu, Krishnan Srinivasan, Parth Shah, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Jeannette Bohg · 2019
Later among the works it cites.
Learning representations from imperfect time series data via tensor rank regularization
Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2019
Later among the works it cites.
Naturalistic audio-movies and narrative synchronize “visual” cortices across congenitally blind but not sighted individuals
Rita E Loiotile, Rhodri Cusack, and Marina Bedny · 2019
Later among the works it cites.
Vilbert: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
A survey of reinforcement learning informed by natural language
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob N Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 2019
Later among the works it cites.
Multimodal learning in health sciences and medicine: Merging technologies to enhance student learning and communication
Christian Moro, Jessica Smith, and Zane Stromberga · 2019
Later among the works it cites.
Multimodal dialog system: Generating responses via adaptive decoders
Liqiang Nie, Wenjie Wang, Richang Hong, Meng Wang, and Qi Tian · 2019
Later among the works it cites.
The human imagination: the cognitive neuroscience of visual mental imagery
Joel Pearson · 2019
Later among the works it cites.
Habitat: A platform for embodied ai research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al · 2019
Later among the works it cites.
Active haptic perception in robots: a review
Lucia Seminara, Paolo Gastaldo, Simon J Watt, Kenneth F Valyear, Fernando Zuher, and Fulvio Mastrogiovanni · 2019
Later among the works it cites.
Probabilistic neural symbolic models for interpretable visual question answering
Ramakrishna Vedantam, Karan Desai, Stefan Lee, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
Scaling autoregressive video models
Dirk Weissenborn, Oscar Täckström, and Jakob Uszkoreit · 2019
Later among the works it cites.
Adaptive cross-modal few-shot learning
Chen Xing, Negar Rostamzadeh, Boris Oreshkin, and Pedro O O. Pinheiro · 2019
Later among the works it cites.
Multimodal machine learning for automated icd coding
Keyang Xu, Mike Lam, Jingzhi Pang, Xin Gao, Charlotte Band, Piyush Mathur, Frank Papay, Ashish K Khanna, Jacek B Cywinski, Kamal Maheshwari, et al · 2019
Later among the works it cites.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Later among the works it cites.
Deep supervised cross-modal retrieval
Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng · 2019
Later among the works it cites.
Towards learning visual semantics
Haipeng Cai, Shiv Raj Pant, and Wen Li · 2020
Later among the works it cites.
See, hear, explore: Curiosity via audio-visual association
Victoria Dean, Shubham Tulsiani, and Abhinav Gupta · 2020
Later among the works it cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Later among the works it cites.
Multimodal sensor fusion with differentiable filters
Michelle A Lee, Brent Yi, Roberto Martín-Martín, Silvio Savarese, and Jeannette Bohg · 2020
Later among the works it cites.
Improving vision-and-language navigation with image-text pairs from the web
Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra · 2020
Later among the works it cites.
Comir: Contrastive multimodal image representation for registration
Nicolas Pielawski, Elisabeth Wetzer, Johan Öfverstedt, Jiahao Lu, Carolina Wählby, Joakim Lindblad, and Natasa Sladoje · 2020
Later among the works it cites.
Tresnet: High performance gpu-dedicated architecture
Tal Ridnik, Hussam Lawen, Asaf Noy, and Itamar Friedman · 2020
Later among the works it cites.
A comprehensive survey on graph neural networks
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip · 2020
Later among the works it cites.
Foundations of multimodal co-learning
Amir Zadeh, Paul Pu Liang, and Louis-Philippe Morency · 2020
Later among the works it cites.
Syntax: A generative introduction
Andrew Carnie · 2021
Later among the works it cites.
Fausto Giunchiglia, Luca Erculiani, and Andrea Passerini · 2021
Later among the works it cites.
Subcortical circuits mediate communication between primary sensory cortical areas in mice
Michael Lohse, Johannes C Dahmen, Victoria M Bajo, and Andrew J King · 2021
Later among the works it cites.
Styleptb: A compositional benchmark for fine-grained controllable text style transfer
Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Multimodal contrastive training for visual representation learning
Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, Yilin Wang, Michael Maire, Ajinkya Kale, and Baldo Faieta · 2021
Later among the works it cites.
When brains dream: Exploring the science and mystery of sleep
Antonio Zadra and Robert Stickgold · 2021
Later among the works it cites.