Fetching the paper…
Reading the bibliography…
Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI).
Multisensory integration
Nadler, S. A · 1993
Earlier work this paper cites.
Human hippocampus associates information in memory
Henke, K., Weber, B., Kneifel, S., Wieser, H. G. & Buck, A · 1999
Earlier work this paper cites.
Modulation of human visual cortex by crossmodal spatial attention
Macaluso, E., Frith, C. D. & Driver, J · 2000
Earlier work this paper cites.
Frontal lobe functions
Chayer, C. & Freedman, M · 2001
Earlier work this paper cites.
Automated anatomical labeling of activations in SPM using a macroscopic anatomical parcellation of the MNI MRI single-subject brain
Tzourio-Mazoyer, N. et al · 2002
Earlier work this paper cites.
The human amygdala and the emotional evaluation of sensory stimuli
Zald, D. H · 2003
Earlier work this paper cites.
Functional magnetic resonance imaging
Matthews, P. M. & Jezzard, P · 2004
Earlier work this paper cites.
Unraveling multisensory integration: patchy organization within human STS multisensory cortex
Beauchamp, M. S., Argall, B. D., Bodurka, J., Duyn, J. H. & Martin, A · 2004
Earlier work this paper cites.
Multisensory integration: methodological approaches and emerging principles in the human brain
Calvert, G. A. & Thesen, T · 2004
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Smith, L. & Gasser, M · 2005
Earlier work this paper cites.
Bimodal format effects in working memory
Goolkasian, P. & Foos, P. W · 2005
Earlier work this paper cites.
Is neocortex essentially multisensory?
Ghazanfar, A. A. & Schroeder, C. E · 2006
Earlier work this paper cites.
Multisensory processing via early cortical stages: connections of the primary auditory cortical field with other sensory systems
Budinger, E., Heil, P., Hess, A. & Scheich, H · 2006
Earlier work this paper cites.
A central capacity limit to the simultaneous storage of visual and auditory arrays in working memory
Saults, J. S. & Cowan, N · 2007
Earlier work this paper cites.
Semantic encoding in working memory: Is there a (multi) modality effect?
Delogu, F., Raffone, A. & Belardinelli, M. O · 2009
Earlier work this paper cites.
Supramarginal gyrus involvement in visual word recognition
Stoeckel, C., Gough, P. M., Watkins, K. E. & Devlin, J. T · 2009
Earlier work this paper cites.
Multisensory interactions in primate auditory cortex: fMRI and electrophysiology
Kayser, C., Petkov, C. I. & Logothetis, N. K · 2009
Earlier work this paper cites.
Multimodal fusion for multimedia analysis: a survey
Atrey, P. K., Hossain, M. A., El Saddik, A. & Kankanhalli, M. S · 2010
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval
Rasiwasia, N. et al · 2010
Earlier work this paper cites.
Multisensory integration affects visuo-spatial working memory
Botta, F. et al · 2011
Earlier work this paper cites.
Encoding and decoding in fMRI
Naselaris, T., Kay, K. N., Nishimoto, S. & Gallant, J. L · 2011
Earlier work this paper cites.
Modality-independent coding of spatial layout in the human brain
Wolbers, T., Klatzky, R. L., Loomis, J. M., Wutte, M. G. & Giudice, N. A · 2011
Earlier work this paper cites.
A ventral visual stream reading center independent of visual experience
Reich, L., Szwed, M., Cohen, L. & Amedi, A · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G. & Berg, T · 2011
Earlier work this paper cites.
A continuous semantic space describes the representation of thousands of object and action categories across the human brain
Huth, A. G., Nishimoto, S., Vu, A. T. & Gallant, J. L · 2012
Earlier work this paper cites.
Both the middle temporal gyrus and the ventral anterior temporal area are crucial for multimodal semantic processing: distortion-corrected fMRI evidence for a double gradient of information convergence in the temporal lobes
Visser, M., Jefferies, E., Embleton, K. V. & Lambon Ralph, M. A · 2012
Earlier work this paper cites.
Reading with sounds: sensory substitution selectively activates the visual word form area in the blind
Striem-Amit, E., Cohen, L., Dehaene, S. & Amedi, A · 2012
Earlier work this paper cites.
Brainnet viewer: a network visualization tool for human brain connectomics
Xia, M., Wang, J. & He, Y · 2013
Earlier work this paper cites.
Subregions of the human superior frontal gyrus and their connections
Li, W. et al · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y. et al · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A. et al · 2015
Cited alongside, same era.
Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream
Güçlü, U. & van Gerven, M. A · 2015
Cited alongside, same era.
Using goal-driven deep learning models to understand sensory cortex
Yamins, D. L. & DiCarlo, J. J · 2016
Cited alongside, same era.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T. & Rui, Y · 2016
Cited alongside, same era.
The multisensory function of the human primary visual cortex
Murray, M. M. et al · 2016
Cited alongside, same era.
Measuring the performance of neural models
Schoppe, O., Harper, N. S., Willmore, B. D., King, A. J. & Schnupp, J. W · 2016
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J. & Le, Q. V · 2020
Later among the works it cites.
Bioinspired multisensory neural network with crossmodal integration and recognition
Tan, H., Zhou, Y., Tao, Q., Rosen, J. & van Dijken, S · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A. et al · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C. et al · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Yang, A., Miech, A., Sivic, J., Laptev, I. & Schmid, C · 2021
Later among the works it cites.
Multimodal neurons in artificial neural networks
Goh, G. et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the role of the posterior middle temporal gyrus in semantic cognition: Integration of anterior temporal lobe with executive processes
Davey, J. et al · 2016
Cited alongside, same era.
A multisensory perspective on object memory
Matusz, P. J., Wallace, M. T. & Murray, M. M · 2017
Cited alongside, same era.
Audiovisual speech integration in the superior temporal region is dysfunctional in dyslexia
Ye, Z., Rüsseler, J., Gerth, I. & Münte, T. F · 2017
Cited alongside, same era.
Domain selectivity in the parahippocampal gyrus is predicted by the same structural connectivity patterns in blind and sighted individuals
Wang, X. et al · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R. et al · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
van den Oord, A., Li, Y. & Vinyals, O · 2018
Cited alongside, same era.
Later among the works it cites.
Limits to visual representational correspondence between convolutional neural networks and the human brain
Xu, Y. & Vaziri-Pashkam, M · 2021
Later among the works it cites.
Computational models of category-selective brain regions enable high-throughput tests of selectivity
Ratan Murty, N. A., Bashivan, P., Abate, A., DiCarlo, J. J. & Kanwisher, N · 2021
Later among the works it cites.
Unsupervised neural network models of the ventral visual stream
Zhuang, C. et al · 2021
Later among the works it cites.
Low-dimensional structure in the space of language representations is reflected in brain responses
Antonello, R., Turek, J. S., Vo, V. & Huth, A · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J. et al · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A. et al · 2021
Later among the works it cites.
ERNIE-ViL: Knowledge enhanced vision-language representations through scene graphs
Yu, F. et al · 2021
Later among the works it cites.
ViLT: Vision-and-language transformer without convolution or region supervision
Kim, W., Son, B. & Kim, I · 2021
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G. & Zisserman, A · 2021
Later among the works it cites.
TACo: Token-aware cascade contrastive learning for video-text alignment
Yang, J., Bisk, Y. & Gao, J · 2021
Later among the works it cites.
Support-set bottlenecks for video-text representation learning
Patrick, M. et al · 2021
Later among the works it cites.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Chefer, H., Gur, S. & Wolf, L · 2021
Later among the works it cites.
What can 5.17 billion regression fits tell us about artificial models of the human visual system?
Conwell, C., Prince, J. S., Alvarez, G. A. & Konkle, T · 2021
Later among the works it cites.
Cortical response to naturalistic stimuli is largely predictable with deep neural networks
Khosla, M., Ngo, G. H., Jamison, K., Kuceyeski, A. & Sabuncu, M. R · 2021
Later among the works it cites.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N. & Soricut, R · 2021
Later among the works it cites.
Zero-infinity: Breaking the GPU memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S. & He, Y · 2021
Later among the works it cites.
COTS: Collaborative two-stream vision-language pre-training model for cross-modal retrieval
Lu, H. et al · 2022
Closest in time.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C. & Chen, M · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C. et al · 2022
Closest in time.
Towards artificial general intelligence via a multimodal foundation model
Fei, N. et al · 2022
Closest in time.
Gallant lab natural short clips 3T fMRI data
Huth, A. G., Nishimoto, S., Vu, A. T. & la Tour, T. D · 2022
Closest in time.
The human language effective connectome
Rolls, E. T., Deco, G., Huang, C.-C. & Feng, J · 2022
Closest in time.
Feature-space selection with banded ridge regression
la Tour, T. D., Eickenberg, M. & Gallant, J. L · 2022
Closest in time.