Fetching the paper…
Reading the bibliography…
This paper addresses the problem of supervised video summarization by formulating it as a sequence-to-sequence learning problem, where the input is a sequence of original video frames, the output is a keyshot sequence.
Yufei Ma, Lie Lu, Hongjiang Zhang, and Mingjing Li, “A user attention model for video summarization,” in Proc. ACM Multimedia , 2002, pp. 533–542
2002
Earlier work this paper cites.
Chong Wah Ngo, Yufei Ma, and Hongjiang Zhang, “Video summarization and scene detection by graph modeling,” IEEE Trans. Circuits Syst. Video Technol. , vol. 15, no. 2, pp. 296–305, 2005
2005
Earlier work this paper cites.
Zuzana Cernekova, Ioannis Pitas, and Christophoros Nikou, “Information theory-based shot cut/fade detection and video summarization,” IEEE Trans. Circuits Syst. Video Technol. , vol. 16, no. 1, pp. 82–90, 2006
2006
Earlier work this paper cites.
Ba Tu Truong, and Svetha Venkatesh, “Video abstraction: a systematic review and classification,” ACM Trans. Multimedia Comput., Commun. Appl. , vol. 3, no. 1, pp. 1–37, 2007
2007
Earlier work this paper cites.
Arthur G. Money, and Harry Agius, “Video summarisation: a conceptual framework and survey of the state of the art,” J. Vis. Commun. Image Represent. , vol. 12, no. 2, pp. 121–143, 2008
2008
Earlier work this paper cites.
Yanwei Fu, Yanwen Guo, Yanshu Zhu, Feng Liu, Chuanming Song, and Zhihua Zhou, “Multi-View Video Summarization,” IEEE Trans. Multimedia , vol. 12, no. 7, pp. 717–729, 2010
2010
Earlier work this paper cites.
Sandra Eliza Fontes de Avila, Ana Paula Brando Lopes, Antonio da Luz Jr., and Arnaldo de Albuquerque Ara¨²jo, “VSUMM: a mechanism designed to produce static video summaries and a novel evaluation method,” Pattern Recognit. Lett. , vol. 32, no. 1, pp. 56–68, 2011
2011
Earlier work this paper cites.
Meng Wang, Richang Hong, Guangda Li, Zhengjun Zha, and Shuicheng Yan, “Event driven web video summarization by tag localization and key-shot identification,” IEEE Trans. Multimedia, vol. 14, no. 4, pp. 975–985, 2012
2012
Earlier work this paper cites.
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman, “Discovering important people and objects for egocentric video summarization,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2012, pp. 1346–1353
2012
Earlier work this paper cites.
Yang Cong, Junsong Yuan, and Jiebo Luo, “Towards scalable summarization of consumer videos via sparse dictionary selection,” IEEE Trans. Multimedia , vol. 14, no. 1, pp. 66–75, 2012
2012
Earlier work this paper cites.
Aditya Khosla, Raffay Hamid, Chih-Jen Lin and Neel Sundaresan, “Large-scale video summarization using web-image priors,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2013, pp. 2698–2705
2013
Earlier work this paper cites.
Ejaz, Naveed, Irfan Mehmood, and Sung Wook Baik, “Efficient visual attention based framework for extracting key frames from videos,” Signal Process. Image Commun. , vol. 28, no. 1, pp. 34–44, 2013
2013
Earlier work this paper cites.
Sanjay K. Kuanar, Rameswar Panda, and Ananda. S. Chowdhury, “Video key frame extraction through dynamic delaunay clustering with a structural constraint,” J. Vis. Commun. Image Represent. , vol. 24, no. 7, pp. 1212–1227, 2013
2013
Earlier work this paper cites.
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, and Luc. Van Gool, “Creating summaries from user videos,” in Proc. Eur. Conf. Comput. Vis. , 2014, pp. 505–520
2014
Earlier work this paper cites.
Danila Potapov, Matthijs Douze, Zaid Harchaoui, and Cordelia Schmid, “Category-specific video summarization,” in Proc. Eur. Conf. Comput. Vis. , 2014, pp. 540–555
2014
Cited alongside, same era.
Michael Gygli, Helmut Grabner, and Luc Van Gool, “Video summarization by learning submodular mixtures of objectives,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2015, pp. 3090–3098
2015
Cited alongside, same era.
Subhashini Venugopalan, Marcus Rohrbach, Jeff Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko, “Sequence to sequence-video to text,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2015, pp. 4534–4542
2015
Cited alongside, same era.
Dzmitry Bahdanau, Kyungh Yun Cho, and Yoshua Bengio, “Neural machine translation by jointly learning to align and translate,” in Int. Conf. Learn. Representations , 2015, pp. 1–15
2015
Cited alongside, same era.
Yunzuo Zhang, Ran Tao, and Yue Wang, “Motion-state-adaptive video summarization via spatiotemporal analysis,” IEEE Trans. Circuits Syst. Video Technol. , vol. 27 , no. 6, pp. 1340–1352, 2017
2017
Closest in time.
Xun Xu, Timothy M. Hospedales, and Shaogang Gong, “Discovery of shared semantic spaces for multiscene video query and summarization,” IEEE Trans. Circuits Syst. Video Technol. , vol. 27, no. 6, pp. 1353–1367, 2017
2017
Closest in time.
Richang Hong, Lei Li, Junjie Cai, Dapeng Tao, Meng Wang, and Qi Tian, “Coherent semantic-visual indexing for large-scale image retrieval in the cloud,” IEEE Trans. Image Process. , vol. 26, no. 9, pp. 4128–4138, 2017
2017
Closest in time.
Erkun Yang, Cheng Deng, Wei Liu, Xianglong Liu, Dacheng Tao, and Xinbo Gao, “Pairwise relationship guided deep hashing for cross-modal retrieval,” in Proc. AAAI Conf. Art. Intell. , 2017, pp. 1618–1625
2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minh Thang Luong, Hieu Pham, Christopher D. Manning, “Effective approaches to attention-based neural machine translation,” in Conf. on Empi. Meth. Natural Lan. Proc. , 2015, pp. 1412–1421
2015
Cited alongside, same era.
Huan Yang, Baoyuan Wang, Stephen Lin, David Wipf, Minyi Guo, and Baining Guo, “Unsupervised extraction of video highlights via robust recurrent auto-encoders,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2015, pp. 4633–4641
2015
Cited alongside, same era.
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes, “TVSum: summarizing web videos using titles,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2015, pp. 5179–5187
2015
Cited alongside, same era.
Shaohui Mei, Genliang Guan, Zhiyong Wang, Mingyi He, and David Dagan Feng, “Video summarization via minimum sparse reconstruction,” Pattern Recognition , vol. 48, no. 2, pp. 522–533, 2015
2015
Cited alongside, same era.
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich, “Going deeper with convolutions,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2015, pp. 1–9
2015
Cited alongside, same era.
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman, “ Video summarization with long short-term memory,” in Proc. Eur. Conf. Comput. Vis. , 2016, pp. 766–782
2016
Cited alongside, same era.
Gunnar A. Sigurdsson, Xinlei Chen, Abhinav Gupta, “Learning visual storylines with skipping recurrent neural networks,” in Proc. Eur. Conf. Comput. Vis. , 2016, pp. 71–88
2016
Cited alongside, same era.
Liqiang Nie, Richang Hong, Luming Zhang, Yingjie Xia, Dacheng Tao, and Nicu Sebe, “Perceptual attributes optimization for multivideo summarization,” IEEE Trans. Cybern. , vol. 46, no. 12, pp. 2991–3003, 2016
2016
Cited alongside, same era.
Yachuang Feng, Yuan Yuan, and Xiaoqiang Lu, “Learning deep event models for crowd anomaly detection,” Neurocomput. , vol. 219, pp. 548–556, 2017
2017
Closest in time.
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic, “Unsupervised video summarization with adversarial LSTM networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2017, pp. 1–10
2017
Closest in time.
Li Xuelong, Bin Zhao, and Xiaoqiang Lu, “A general framework for edited video and raw video summarization,” IEEE Trans. Image Process. , vol. 26, no. 8, pp. 3652–3664, 2017
2017
Closest in time.
2017
Closest in time.
Adway Mitra, Soma Biswas, and Chiranjib Bhattacharyya. “Bayesian modeling of temporal coherence in videos for entity discovery and summarization,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 3, pp. 430–443, 2017
2017
Closest in time.
Haykel Boukadida, Sid-Ahmed Berrani, and Patrick Gros, “Automatically creating adaptive video summaries using constraint satisfaction programming: Application to sport content,” IEEE Trans. Circuits Syst. Video Technol. , vol. 27, no. 4, pp. 920–934, 2017
2017
Closest in time.
Alex Graves, and J¨¹rgen Schmidhuber, “Framewise phoneme classification with bidirectional LSTM networks,” in Int. Joint Conf. on Neural Net. , vol. 4, 2005, pp. 2047–2052
2052
Closest in time.
Keivin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio, “Show, attend and tell: neural image caption generation with visual attention,” in Proc. 32nd Int. Conf. Mach. Learn. , 2015, pp. 2048–2057
2057
Closest in time.
Boqing Gong, Weilun Chao, Kristen Grauman, and Fei Sha, “Diverse sequential subset selection for supervised video summarization,” in Advances Neural Inf. Process. Syst. , 2014, pp. 2069–2077
2077
Closest in time.