Fetching the paper…
Reading the bibliography…
In the art of video editing, sound helps add character to an object and immerse the viewer within a space.
Autofoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning
Sanchita Ghose and John Jeffrey Prevost. 2020 · 1907
Earlier work this paper cites.
Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research
Sandra G Hart and Lowell E Staveland. 1988 · 1988
Earlier work this paper cites.
Efficient color histogram indexing for quadratic form distance functions
James Hafner, Harpreet S. Sawhney, William Equitz, Myron Flickner, and Wayne Niblack. 1995 · 1995
Earlier work this paper cites.
SUS-A quick and dirty usability scale
John Brooke et al · 1996
Earlier work this paper cites.
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke. 2006 · 2006
Earlier work this paper cites.
User guided audio selection from complex sound mixtures. In Proceedings of the 22nd annual ACM symposium on User interface software and technology . 89–92
Paris Smaragdis. 2009 · 2009
Earlier work this paper cites.
Content-based tools for editing audio stories. In Proceedings of the 26th annual ACM symposium on User interface software and technology . 113–122
Steve Rubin, Floraine Berthouzoz, Gautham J Mysore, Wilmot Li, and Maneesh Agrawala. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context. In European conference on computer vision . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Generating emotionally relevant musical scores for audio stories. In Proceedings of the 27th annual ACM symposium on User interface software and technology . 439–448
Steve Rubin and Maneesh Agrawala. 2014 · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research. In Proceedings of the 22nd ACM international conference on Multimedia . 1041–1044
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello. 2014 · 2014
Earlier work this paper cites.
ESC: Dataset for environmental sound classification. In Proceedings of the 23rd ACM international conference on Multimedia . 1015–1018
Karol J Piczak. 2015 · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Visually indicated sounds. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2405–2413
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman. 2016 · 2016
Earlier work this paper cites.
Grad-CAM: Why did you say that?
Ramprasaath R Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Dynamic authoring of audio with linked scripts. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology . 509–516
Hijung Valentina Shin, Wilmot Li, and Frédo Durand. 2016 · 2016
Cited alongside, same era.
Look, listen and learn. In Proceedings of the IEEE International Conference on Computer Vision . 609–617
Relja Arandjelovic and Andrew Zisserman. 2017 · 2017
Cited alongside, same era.
MidiNet: A convolutional generative adversarial network for symbolic-domain music generation
Li-Chia Yang, Szu-Yu Chou, and Yi-Hsuan Yang. 2017 · 2017
Cited alongside, same era.
Visually indicated sound generation by perceptually optimized classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops . 0–0
Kan Chen, Chuanxi Zhang, Chen Fang, Zhaowen Wang, Trung Bui, and Ram Nevatia. 2018 · 2018
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video. In Proceedings of the European Conference on Computer Vision (ECCV) . 35–53
Foley music: Learning to generate music from videos. In European Conference on Computer Vision . Springer, 758–775
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba. 2020 · 2020
Later among the works it cites.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2020 · 2020
Later among the works it cites.
Audeo: Audio generation for a silent performance video
Kun Su, Xiulong Liu, and Eli Shlizerman. 2020 · 2020
Later among the works it cites.
Crosspower: Bridging graphics and linguistics. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . 722–734
Haijun Xia. 2020 · 2020
Later among the works it cites.
Crosscast: adding visuals to audio travel podcasts. In Proceedings of the 33rd annual ACM symposium on user interface software and technology . 735–746
Haijun Xia, Jennifer Jacobs, and Maneesh Agrawala. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ruohan Gao, Rogerio Feris, and Kristen Grauman. 2018 · 2018
Cited alongside, same era.
Learning to localize sound source in visual scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4358–4366
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon. 2018 · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos. In Proceedings of the European Conference on Computer Vision (ECCV) . 247–263
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Cited alongside, same era.
The sound of pixels. In Proceedings of the European conference on computer vision (ECCV) . 570–586
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018 · 2018
Cited alongside, same era.
Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3550–3558
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg. 2018 · 2018
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts. 2019 · 2019
Cited alongside, same era.
Co-separating sounds of visual objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3879–3888
Ruohan Gao and Kristen Grauman. 2019 · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Dance2music: Automatic dance-driven music generation
Gunjan Aggarwal and Devi Parikh. 2021 · 2021
Closest in time.
Sanchita Ghose and John J Prevost. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
How Does it Sound?
Kun Su, Xiulong Liu, and Eli Shlizerman. 2021 · 2021
Closest in time.
Foley Techniques and Sound Effects: A Sound Design Guide
2018 · 2022
Closest in time.
Commits - openai/CLIP
2022 · 2022
Closest in time.
Epidemic Sound
2022 · 2022
Closest in time.
Record Once, Post Everywhere: Automatic Shortening of Audio Stories for Social Media. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology . 1–11
Bryan Wang, Zeyu Jin, and Gautham Mysore. 2022 · 2022
Closest in time.
SoundToons: Exemplar-Based Authoring of Interactive Audio-Driven Animation Sprites. In Proceedings of the 28th International Conference on Intelligent User Interfaces . 710–722
Toby Chong, Hijung Valentina Shin, Deepali Aneja, and Takeo Igarashi. 2023 · 2023
Closest in time.