Fetching the paper…
Reading the bibliography…
Humans can imagine a scene from a sound.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Generative adversarial nets,”
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Conditional generative adversarial nets,”
Mehdi Mirza and Simon Osindero, · 2014
Earlier work this paper cites.
“Generative adversarial text to image synthesis,”
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee, · 2016
Earlier work this paper cites.
“Conditional image synthesis with auxiliary classifier gans,”
Augustus Odena, Christopher Olah, and Jonathon Shlens, · 2016
Earlier work this paper cites.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Earlier work this paper cites.
“Improved techniques for training gans,”
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen, · 2016
Earlier work this paper cites.
“Visually indicated sounds,”
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman, · 2016
Earlier work this paper cites.
Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, and Gaurav Sharma, · 2016
Cited alongside, same era.
“Rethinking the inception architecture for computer vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna, · 2016
Cited alongside, same era.
“Seeing and hearing too: Audio representation for video captioning,”
Shun-Po Chuang, Chia-Hung Wan, Pang-Chi Huang, Chi-Yu Yang, and Hung-Yi Lee, · 2017
Cited alongside, same era.
Jae Hyun Lim and Jong Chul Ye, · 2017
Cited alongside, same era.
“Deep and hierarchical implicit models,”
Dustin Tran, Rajesh Ranganath, and David M Blei, · 2017
“Cmcgan: A uniform framework for cross-modal visual-audio mutual generation,”
Wangli Hao, Zhaoxiang Zhang, and He Guan, · 2017
Later among the works it cites.
Martin Arjovsky, Soumith Chintala, and Léon Bottou, · 2017
Later among the works it cites.
“cgans with projection discriminator,”
Takeru Miyato and Masanori Koyama, · 2018
Closest in time.
“Spectral normalization for generative adversarial networks,”
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Learning word-like units from joint audio-visual analysis,”
David Harwath and James R Glass, · 2017
Cited alongside, same era.
“Attention-based multimodal fusion for video description,”
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R Hershey, Tim K Marks, and Kazuhiko Sumi, · 2017
Cited alongside, same era.
“Deep cross-modal audio-visual generation,”
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu, · 2017
Cited alongside, same era.
Shane Barratt and Rishi Sharma, · 2018
Closest in time.
“Vision as an interlingua: Learning multilingual semantic embeddings of untranscribed speech,”
David Harwath, Galen Chuang, and James Glass, · 2018
Closest in time.
“Multimodal attention for fusion of audio and spatiotemporal features for video description,”
Chiori Hori, Takaaki Hori, Gordon Wichern, Jue Wang, Teng-yok Lee, Anoop Cherian, and Tim K Marks, · 2018
Closest in time.