Fetching the paper…
Reading the bibliography…
State of the art architectures for untrimmed video Temporal Action Localization (TAL) have only considered RGB and Flow modalities, leaving the information-rich audio modality totally unexploited.
THUMOS challenge: Action recognition with a large number of classes, 2014
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2015
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition, 2016
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Temporal activity detection in untrimmed videos with recurrent neural networks
Alberto Montes, Amaia Salvador, and Xavier Giró-i-Nieto · 2016
Earlier work this paper cites.
Action temporal localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Multi-stream multi-class fusion of deep networks for video classification
Zuxuan Wu, Yu-Gang Jiang, Xi Wang, Hao Ye, and Xiangyang Xue · 2016
Earlier work this paper cites.
Cuhk & ethz & siat submission to activitynet challenge 2016, 2016
Yuanjun Xiong, Limin Wang, Zhe Wang, Bowen Zhang, Hang Song, Wei Li, Dahua Lin, Yu Qiao, Luc Van Gool, and Xiaoou Tang · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
TURN TAP: temporal unit regression network for temporal action proposals
Jiyang Gao, Zhenheng Yang, Chen Sun, Kan Chen, and Ram Nevatia · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Single shot temporal action detection
Tianwei Lin, Xu Zhao, and Zheng Shou · 2017
Earlier work this paper cites.
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Earlier work this paper cites.
R-C3D: region convolutional 3d network for temporal activity detection
Huijuan Xu, Abir Das, and Kate Saenko · 2017
Earlier work this paper cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Dahua Lin, and Xiaoou Tang · 2017
Earlier work this paper cites.
Diagnosing error in temporal action detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia, and Bernard Ghanem · 2018
Earlier work this paper cites.
CTAP: complementary temporal action proposal generation
Jiyang Gao, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
Rethinking Fusion Baselines for Multi-modal Human Action Recognition: 19th Pacific-Rim Conference on Multimedia, Hefei, China, September 21-22, 2018, Proceedings, Part III
Hongda Jiang, Yanghao Li, Sijie Song, and Jiaying Liu · 2018
Cited alongside, same era.
BSN: boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Cited alongside, same era.
Multi-granularity generator for temporal action proposal
Yuan Liu, Lin Ma, Yifeng Zhang, Wei Liu, and Shih-Fu Chang · 2018
Cited alongside, same era.
Multimodal keyless attention fusion for video classification
Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization
Runhao Zeng, Chuang Gan, Peihao Chen, Wenbing Huang, Qingyao Wu, and Mingkui Tan · 2019
Later among the works it cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.
Fully supervised speaker diarization
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang · 2019
Later among the works it cites.
TSP: temporally-sensitive pretraining of video encoders for localization tasks
Humam Alwassel, Silvio Giancola, and Bernard Ghanem · 2020
Later among the works it cites.
Convolutional recurrent neural networks for weakly labeled semi-supervised sound event detection in domestic environments
Janek Ebbers and Reinhold Haeb-Umbach · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiang Long, Chuang Gan, Gerard De Melo, Xiao Liu, Yandong Li, Fu Li, and Shilei Wen · 2018
Cited alongside, same era.
Attention clusters: Purely attention based local feature integration for video classification
Xiang Long, Chuang Gan, Gerard de Melo, Jiajun Wu, Xiao Liu, and Shilei Wen · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A. Efros · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos, 2018
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Speaker diarization with lstm
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno · 2018
Cited alongside, same era.
End-to-end, single-stream temporal action detection in untrimmed videos
Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles · 2019
Cited alongside, same era.
Mean teacher with data augmentation for dcase 2019 task 4
Lionel Delphin-Poulat and Cyril Plapous · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Cited alongside, same era.
Cross-domain sound event detection: from synthesized audio to real audio
Junyong Hao, Zhenwei Hou, and Wang Peng · 2020
Later among the works it cites.
Fast learning of temporal action proposal via dense boundary generator
Chuming Lin, Jian Li, Yabiao Wang, Ying Tai, Donghao Luo, Zhipeng Cui, Chengjie Wang, Jilin Li, Feiyue Huang, and Rongrong Ji · 2020
Later among the works it cites.
Convolution-augmented transformer for semi-supervised sound event detection
Koichi Miyazaki, Tatsuya Komatsu, Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, and Kazuya Takeda · 2020
Later among the works it cites.
Proceedings of the Fifth Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2020)
Nobutaka Ono, Noboru Harada, Yohei Kawaguchi, Annamaria Mesaros, Keisuke Imoto, Yuma Koizumi, , and Tatsuya Komatsu · 2020
Later among the works it cites.
Haisheng Su, Weihao Gan, Wei Wu, Junjie Yan, and Yu Qiao · 2020
Later among the works it cites.
Cross-modal relation-aware networks for audio-visual event localization
Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, and Chuang Gan · 2020
Later among the works it cites.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S. Rojas, Ali Thabet, and Bernard Ghanem · 2020
Later among the works it cites.
Multimodal transformer networks with latent interaction for audio-visual event localization
Yixuan He, Xing Xu, Xin Liu, Weihua Ou, and Huimin Lu · 2021
Closest in time.
Cross-attentional audio-visual fusion for weakly-supervised action localization
Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, and Sungrack Yun · 2021
Closest in time.
Learning salient boundary feature for anchor-free temporal action localization
Chuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu · 2021
Closest in time.
Multi-shot temporal event localization: A benchmark
Xiaolong Liu, Yao Hu, Song Bai, Fei Ding, Xiang Bai, and Philip H. S. Torr · 2021
Closest in time.
Positive sample propagation along the audio-visual event line
Jinxing Zhou, Liang Zheng, Yiran Zhong, Shijie Hao, and Meng Wang · 2021
Closest in time.