Fetching the paper…
Reading the bibliography…
In this paper, we propose to make a systematic study on machines multisensory perception under attacks.
Visual contribution to speech intelligibility in noise
William H Sumby and Irwin Pollack · 1954
Earlier work this paper cites.
Hearing lips and seeing voices
Harry McGurk and John MacDonald · 1976
Earlier work this paper cites.
Learning classification with unlabeled data
Virginia R de Sa · 1994
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John R Hershey and Javier R Movellan · 2000
Earlier work this paper cites.
Pixels that sound
Einat Kidron, Yoav Y Schechner, and Michael Elad · 2005
Earlier work this paper cites.
Image denoising via sparse and redundant representations over learned dictionaries
Michael Elad and Michal Aharon · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Aero-tactile integration in speech perception
Bryan Gick and Donald Derrick · 2009
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
Image super-resolution via sparse representation
Jianchao Yang, John Wright, Thomas S Huang, and Yi Ma · 2010
Earlier work this paper cites.
The neural bases of multisensory processes
Micah M Murray and Mark T Wallace · 2011
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
A multisensory perspective on human auditory communication
Katharina von Kriegstein · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Distributional smoothing with virtual adversarial training
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Ken Nakae, and Shin Ishii · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio · 2016
Earlier work this paper cites.
Deepfool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
The limitations of deep learning in adversarial settings
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu · 2017
Earlier work this paper cites.
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Countering adversarial images using input transformations
Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens Van Der Maaten · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Cited alongside, same era.
End-to-end multi-view lipreading
Stavros Petridis, Yujiang Wang, Zuwei Li, and Maja Pantic · 2017
Cited alongside, same era.
Poster: Inaudible voice commands
Liwei Song and Prateek Mittal · 2017
Cited alongside, same era.
Ensemble adversarial training: Attacks and defenses
Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel · 2017
Cited alongside, same era.
Mitigating adversarial effects through randomization
Cihang Xie, Jianyu Wang, Zhishuai Zhang, Zhou Ren, and Alan Yuille · 2017
Rethinking softmax cross-entropy loss for adversarial robustness
Tianyu Pang, Kun Xu, Yinpeng Dong, Chao Du, Ning Chen, and Jun Zhu · 2019
Later among the works it cites.
Imperceptible, robust, and targeted adversarial examples for automatic speech recognition
Yao Qin, Nicholas Carlini, Garrison Cottrell, Ian Goodfellow, and Colin Raffel · 2019
Later among the works it cites.
Self-supervised audio-visual co-segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba · 2019
Later among the works it cites.
Adversarial defense by stratified convolutional sparse coding
Bo Sun, Nian-hsuan Tsai, Fangchen Liu, Ronald Yu, and Hao Su · 2019
Later among the works it cites.
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dolphinattack: Inaudible voice commands
Guoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang, Taimin Zhang, and Wenyuan Xu · 2017
Cited alongside, same era.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Cited alongside, same era.
Audio adversarial examples: Targeted attacks on speech-to-text
Nicholas Carlini and David Wagner · 2018
Cited alongside, same era.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li · 2018
Cited alongside, same era.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Cited alongside, same era.
Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L Yuille · 2019
Later among the works it cites.
Recursive visual sound separation using minus-plus net
Xudong Xu, Bo Dai, and Dahua Lin · 2019
Later among the works it cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Later among the works it cites.
Vision-infused deep audio inpainting
Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, and Xiaogang Wang · 2019
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
Soundspaces: Audio-visual navigation in 3d environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc, Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman · 2020
Later among the works it cites.
Towards resistant audio adversarial examples
Tom Dörr, Karla Markert, Nicolas M Müller, and Konstantin Böttinger · 2020
Later among the works it cites.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba · 2020
Later among the works it cites.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B Tenenbaum, and Antonio Torralba · 2020
Later among the works it cites.
Look, listen, and act: Towards audio-visual embodied navigation
Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, and Joshua B Tenenbaum · 2020
Later among the works it cites.
Visualechoes: Spatial image representation learning through echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah, Carl Schissler, and Kristen Grauman · 2020
Later among the works it cites.
Discriminative sounding objects localization via self-supervised audiovisual matching
Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou · 2020
Later among the works it cites.
Audio-visual speech inpainting with deep learning
Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan, and Jesper Jensen · 2020
Later among the works it cites.
Multiple sound sources localization from coarse to fine
Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin · 2020
Later among the works it cites.
What makes the sound?: A dual-modality interacting network for audio-visual event localization
Janani Ramaswamy · 2020
Later among the works it cites.
See the sound, hear the pixels
Janani Ramaswamy and Sukhendu Das · 2020
Later among the works it cites.
Unified multisensory perception: weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Later among the works it cites.
Semantic object prediction and spatial sound super-resolution with binaural sounds
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Later among the works it cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Later among the works it cites.
Cyclic co-learning of sounding object visual grounding and sound separation
Yapeng Tian, Di Hu, and Chenliang Xu · 2021
Closest in time.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Closest in time.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wu Wayne, Chen Change Loy, Xiaogang Wang, and Liu Ziwei · 2021
Closest in time.