Fetching the paper…
Reading the bibliography…
We are perceiving and communicating with the world in a multisensory manner, where different information sources are sophisticatedly processed and interpreted by separate parts of the human brain to constitute a complex, yet harmonious and unified sensing system.
Determining optical flow
Berthold KP Horn and Brian G Schunck · 1981
Earlier work this paper cites.
Principal component analysis
Svante Wold, Kim Esbensen, and Paul Geladi · 1987
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“what” and “where” in the human auditory system
Claude Alain, Stephen R Arnott, Stephanie Hevenor, Simon Graham, and Cheryl L Grady · 2001
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Histochemical identification of cortical areas in the auditory region of the human brain
Mark N Wallace, Peter W Johnston, and Alan R Palmer · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G Lowe · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Surf: Speeded up robust features
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Improved signal-to-noise ratio estimation for speech enhancement
Cyril Plapous, Claude Marro, and Pascal Scalart · 2006
Earlier work this paper cites.
Multimodal human–computer interaction: A survey
Alejandro Jaimes and Nicu Sebe · 2007
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A short-time objective intelligibility measure for time-frequency weighted noisy speech
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William B Dolan · 2011
Earlier work this paper cites.
A large-scale benchmark dataset for event recognition in surveillance video
Sangmin Oh, Anthony Hoogs, Amitha Perera, Naresh Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, JK Aggarwal, Hyungtae Lee, Larry Davis, et al · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Neurobiological roots of language in primate audition: common computational properties
Ina Bornkessel-Schlesewsky, Matthias Schlesewsky, Steven L Small, and Josef P Rauschecker · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Deep multimodal semantic embeddings for speech and images
David Harwath and James Glass · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Dense optical flow prediction from a static image
Jacob Walker, Abhinav Gupta, and Martial Hebert · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
An algorithm for predicting the intelligibility of speech masked by modulated noise maskers
Jesper Jensen and Cees H Taal · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung · 2016
Earlier work this paper cites.
Learning deep representations of fine-grained visual descriptions
Scott Reed, Zeynep Akata, Honglak Lee, and Bernt Schiele · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Where to look: Focus regions for visual question answering
Kevin J Shih, Saurabh Singh, and Derek Hoiem · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Video segmentation via object flow
Yi-Hsuan Tsai, Ming-Hsuan Yang, and Michael J Black · 2016
Earlier work this paper cites.
A comprehensive survey on cross-modal retrieval
Kaiye Wang, Qiyue Yin, Wei Wang, Shu Wu, and Liang Wang · 2016
Earlier work this paper cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Huijuan Xu and Kate Saenko · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Earlier work this paper cites.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Earlier work this paper cites.
Automatic speech recognition
Dong Yu and Li Deng · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Earlier work this paper cites.
Deep learning techniques for music generation–a survey
Jean-Pierre Briot, Gaëtan Hadjeres, and François-David Pachet · 2017
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua · 2017
Earlier work this paper cites.
Segflow: Joint learning for video object segmentation and optical flow
Jingchun Cheng, Yi-Hsuan Tsai, Shengjin Wang, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
Towards diverse and natural image descriptions via a conditional gan
Bo Dai, Sanja Fidler, Raquel Urtasun, and Dahua Lin · 2017
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm De Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville · 2017
Earlier work this paper cites.
Improved speech reconstruction from silent video
Ariel Ephrat, Tavi Halperin, and Shmuel Peleg · 2017
Earlier work this paper cites.
Vid2speech: speech reconstruction from silent video
Ariel Ephrat and Shmuel Peleg · 2017
Earlier work this paper cites.
audeep: Unsupervised learning of representations from audio with deep recurrent neural networks
Michael Freitag, Shahin Amiriparian, Sergey Pugachevskiy, Nicholas Cummins, and Björn Schuller · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Scene graph generation from objects, phrases and region captions
Yikang Li, Wanli Ouyang, Bolei Zhou, Kun Wang, and Xiaogang Wang · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2017
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Phrase localization and visual relationship detection with comprehensive image-language cues
Bryan A Plummer, Arun Mallya, Christopher M Cervantes, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Adversarial cross-modal retrieval
Bokun Wang, Yang Yang, Xing Xu, Alan Hanjalic, and Heng Tao Shen · 2017
Earlier work this paper cites.
Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space
Liwei Wang, Alexander Schwing, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Weakly-supervised visual grounding of phrases with linguistic structures
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee · 2017
Cited alongside, same era.
Deep multimodal representation learning from temporal data
Xitong Yang, Palghat Ramesh, Radha Chitta, Sriganesh Madhvanath, Edgar A Bernal, and Jiebo Luo · 2017
Cited alongside, same era.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas · 2017
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Later among the works it cites.
Realistic speech-driven facial animation with gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2020
Later among the works it cites.
Shih-Lun Wu and Yi-Hsuan Yang · 2020
Later among the works it cites.
A comprehensive survey on graph neural networks
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip · 2020
Later among the works it cites.
Cross-modal attention network for temporal inconsistent audio-visual event localization
Hanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang, and Yan Yan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visually indicated sound generation by perceptually optimized classification
Kan Chen, Chuanxi Zhang, Chen Fang, Zhaowen Wang, Trung Bui, and Ram Nevatia · 2018
Cited alongside, same era.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Unsupervised cross-modal alignment of speech and text embedding spaces
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass · 2018
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Cited alongside, same era.
Visual rhythm and beat
Abe Davis and Maneesh Agrawala · 2018
Cited alongside, same era.
Visual grounding via accumulated attention
Chaorui Deng, Qi Wu, Qingyao Wu, Fuyuan Hu, Fan Lyu, and Mingkui Tan · 2018
Cited alongside, same era.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Cited alongside, same era.
Improving one-stage visual grounding by recursive sub-query construction
Zhengyuan Yang, Tianlang Chen, Liwei Wang, and Jiebo Luo · 2020
Later among the works it cites.
Where does it exist: Spatio-temporal video grounding for multi-form sentences
Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, and Lianli Gao · 2020
Later among the works it cites.
In-domain gan inversion for real image editing
Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou · 2020
Later among the works it cites.
Describing unseen videos via multi-modal cooperative dialog agents
Ye Zhu, Yu Wu, Yi Yang, and Yan Yan · 2020
Later among the works it cites.
Dance2music: Automatic dance-driven music generation
Gunjan Aggarwal and Devi Parikh · 2021
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg · 2021
Later among the works it cites.
A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets
Khaled Bayoudh, Raja Knani, Fayçal Hamdaoui, and Abdellatif Mtibaa · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Later among the works it cites.
Probabilistic embeddings for cross-modal retrieval
Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio De Rezende, Yannis Kalantidis, and Diane Larlus · 2021
Later among the works it cites.
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Later among the works it cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Later among the works it cites.
Video background music generation with controllable music transformer
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan · 2021
Later among the works it cites.
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al · 2021
Later among the works it cites.
Audio-visual event localization via recursive fusion by joint co-attention
Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, and Yan Yan · 2021
Later among the works it cites.
How2sign: a large-scale multimodal dataset for continuous american sign language
Amanda Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, and Xavier Giro-i Nieto · 2021
Later among the works it cites.
CLIPScore: a reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Later among the works it cites.
Global context with discrete diffusion in vector quantised modelling for image generation
Minghui Hu, Yujie Wang, Tat-Jen Cham, Jianfei Yang, and PN Suganthan · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Lip to speech synthesis with visual context attentional gan
Minsu Kim, Joanna Hong, and Yong Man Ro · 2021
Later among the works it cites.
Variational diffusion models
Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Later among the works it cites.
Nu-wave: A diffusion probabilistic model for neural audio upsampling
Junhyeok Lee and Seungu Han · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
Ai choreographer: Music conditioned 3d dance generation with aist++
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa · 2021
Later among the works it cites.
Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang · 2021
Later among the works it cites.
Relation-aware instance refinement for weakly supervised visual grounding
Yongfei Liu, Bo Wan, Lin Ma, and Xuming He · 2021
Later among the works it cites.
Playable video generation
Willi Menapace, Stéphane Lathuilière, Sergey Tulyakov, Aliaksandr Siarohin, and Elisa Ricci · 2021
Later among the works it cites.
Symbolic music generation with diffusion models
Gautam Mittal, Jesse Engel, Curtis Hawthorne, and Ian Simon · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Minheng Ni, Haoyang Huang, Lin Su, Edward Cui, Taroon Bharti, Lijuan Wang, Dongdong Zhang, and Nan Duan · 2021
Later among the works it cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2021
Later among the works it cites.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Later among the works it cites.
Counterfactual vqa: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen · 2021
Later among the works it cites.
Audio retrieval with natural language queries
Andreea-Maria Oncescu, A Koepke, Joao F Henriques, Zeynep Akata, and Samuel Albanie · 2021
Later among the works it cites.
Localize to binauralize: Audio spatialization from visual sound source localization
Kranthi Kumar Rachavarapu, Vignesh Sundaresha, AN Rajagopalan, et al · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Sign language recognition: A deep survey
Razieh Rastgoo, Kourosh Kiani, and Sergio Escalera · 2021
Later among the works it cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Later among the works it cites.
Maximum likelihood training of score-based diffusion models
Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon · 2021
Later among the works it cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Later among the works it cites.
3d3l: Deep learned 3d keypoint detection and description for lidars
Dominc Streiff, Lukas Bernreiter, Florian Tschopp, Marius Fehr, and Roland Siegwart · 2021
Later among the works it cites.
Stvgbert: A visual-linguistic transformer based framework for spatio-temporal video grounding
Rui Su, Qian Yu, and Dong Xu · 2021
Later among the works it cites.
Discriminative triad matching and reconstruction for weakly referring expression grounding
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu, and John Y Goulermas · 2021
Later among the works it cites.
The right to talk: An audio-visual transformer approach
Thanh-Dat Truong, Chi Nhan Duong, Hoang Anh Pham, Bhiksha Raj, Ngan Le, Khoa Luu, et al · 2021
Later among the works it cites.
T2vlad: global-local sequence alignment for text-video retrieval
Xiaohan Wang, Linchao Zhu, and Yi Yang · 2021
Later among the works it cites.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Yu Wu and Yi Yang · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Speech prediction in silent videos using variational autoencoders
Ravindra Yadav, Ashish Sardana, Vinay P Namboodiri, and Rajesh M Hegde · 2021
Later among the works it cites.
Jiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng, and Yuejie Zhang · 2021
Later among the works it cites.
Multimodal contrastive training for visual representation learning
Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, Yilin Wang, Michael Maire, Ajinkya Kale, and Baldo Faieta · 2021
Later among the works it cites.
Pano-avqa: Grounded audio-visual question answering on 360deg videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim · 2021
Later among the works it cites.
Multi-stage aggregated transformer network for temporal language localization in videos
Mingxing Zhang, Yang Yang, Xinghan Chen, Yanli Ji, Xing Xu, Jingjing Li, and Heng Tao Shen · 2021
Later among the works it cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
Learning audio-visual correlations from variational cross-modal generation
Ye Zhu, Yu Wu, Hugo Latapie, Yi Yang, and Yan Yan · 2021
Later among the works it cites.
Cm3: A causal masked multimodal model of the internet
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, et al · 2022
Closest in time.
Blended diffusion for text-driven editing of natural images
Omri Avrahami, Dani Lischinski, and Ohad Fried · 2022
Closest in time.
Esb: A benchmark for multi-domain end-to-end speech recognition
Sanchit Gandhi, Patrick Von Platen, and Alexander M Rush · 2022
Closest in time.
A survey of sound source localization with deep learning methods
Pierre-Amaury Grumiaux, Srdjan Kitic, Laurent Girin, and Alexandre Guérin · 2022
Closest in time.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Closest in time.
Mugen: A playground for video-audio-text multimodal understanding and generation
Thomas Hayes, Songyang Zhang, Xi Yin, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Isabelle Hu, and Devi Parikh · 2022
Closest in time.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Closest in time.
Cascaded diffusion models for high fidelity image generation
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans · 2022
Closest in time.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Closest in time.
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet · 2022
Closest in time.
Make it move: controllable image-to-video generation with text descriptions
Yaosi Hu, Chong Luo, and Zhenzhong Chen · 2022
Closest in time.
Deconfounded visual grounding
Jianqiang Huang, Yu Qin, Jiaxin Qi, Qianru Sun, and Hanwang Zhang · 2022
Closest in time.
Towards building asr systems for the next billion users
Tahir Javed, Sumanth Doddapaneni, Abhigyan Raman, Kaushal Santosh Bhogale, Gowtham Ramesh, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh M Khapra · 2022
Closest in time.
Embracing consistency: A one-stage approach for spatio-temporal video grounding
Yang Jin, Zehuan Yuan, Yadong Mu, et al · 2022
Closest in time.
Diffusionclip: Text-guided diffusion models for robust image manipulation
Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye · 2022
Closest in time.
Learning to answer questions in dynamic audio-visual scenarios
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu · 2022
Closest in time.
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2022
Closest in time.
Compositional visual generation with composable diffusion models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum · 2022
Closest in time.
Playable environments: Video manipulation in space and time
Willi Menapace, Stéphane Lathuilière, Aliaksandr Siarohin, Christian Theobalt, Sergey Tulyakov, Vladislav Golyanik, and Elisa Ricci · 2022
Closest in time.
Svts: Scalable video-to-speech synthesis
Rodrigo Mira, Alexandros Haliassos, Stavros Petridis, Björn W Schuller, and Maja Pantic · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Closest in time.
Less can be more: Sound source localization with a classification model
Arda Senocak, Hyeonggon Ryu, Junsik Kim, and In So Kweon · 2022
Closest in time.
Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes
Zengjie Song, Yuxi Wang, Junsong Fan, Tieniu Tan, and Zhaoxiang Zhang · 2022
Closest in time.
Naturalspeech: End-to-end text to speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al · 2022
Closest in time.
Improved vector quantized diffusion models
Zhicong Tang, Shuyang Gu, Jianmin Bao, Dong Chen, and Fang Wen · 2022
Closest in time.
Fine-grained visual entailment
Christopher Thomas, Yipeng Zhang, and Shih-Fu Chang · 2022
Closest in time.
Zeyu Wang, Yu Wu, Karthik Narasimhan, and Olga Russakovsky · 2022
Closest in time.
Switchable novel object captioner
Yu Wu, Lu Jiang, and Yi Yang · 2022
Closest in time.
Tubedetr: Spatio-temporal video grounding with transformers
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Closest in time.
Avqa: A dataset for audio-visual question answering on videos
Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu · 2022
Closest in time.
Explainability in graph neural networks: A taxonomic survey
Hao Yuan, Haiyang Yu, Shurui Gui, and Shuiwang Ji · 2022
Closest in time.
Quantized gan for complex music generation from dance videos
Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, and Sergey Tulyakov · 2022
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Closest in time.
Epic-sounds: A large-scale dataset of actions that sound
Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, and Andrew Zisserman · 2023
Closest in time.
Diffusion models already have a semantic latent space
Mingi Kwon, Jaeseok Jeong, and Youngjung Uh · 2023
Closest in time.
Progressive spatio-temporal perception for audio-visual question answering
Guangyao Li, Wenxuan Hou, and Di Hu · 2023
Closest in time.
Drag your gan: Interactive point-based manipulation on the generative image manifold
Xingang Pan, Ayush Tewari, Thomas Leimkühler, Lingjie Liu, Abhimitra Meka, and Christian Theobalt · 2023
Closest in time.
Overwriting pretrained bias with finetuning data
Angelina Wang and Olga Russakovsky · 2023
Closest in time.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A Clifton · 2023
Closest in time.
Boundary guided learning-free semantic control with diffusion models
Ye Zhu, Yu Wu, Zhiwei Deng, Olga Russakovsky, and Yan Yan · 2023
Closest in time.
Discrete contrastive diffusion for cross-modal and conditional generation
Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan · 2023
Closest in time.
Saying the unseen: Video descriptions via dialog agents
Ye Zhu, Yu Wu, Yi Yang, and Yan Yan · 2023
Closest in time.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Closest in time.