Fetching the paper…
Reading the bibliography…
In this paper, we tackle the problem of egocentric action anticipation, i.e., predicting what actions the camera wearer will perform in the near future and which objects they will interact with.
M. Ma, H. Fan, and K. M. Kitani, “Going deeper into first-person activity recognition,” in Computer Vision and Pattern Recognition , 2016, pp. 1894–1903
1903
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in Computer Vision and Pattern Recognition , 2016, pp. 1933–1941
1941
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to forget: Continual prediction with lstm,” Neural Computation , vol. 12, no. 10, pp. 2451–2471, 2000
2000
Earlier work this paper cites.
I. Laptev, “On space-time interest points,” International Journal of Computer Vision , vol. 64, no. 2-3, pp. 107–123, 2005
2005
Earlier work this paper cites.
C. Zach, T. Pock, and H. Bischof, “A duality based approach for realtime tv-l 1 optical flow,” in Joint Pattern Recognition Symposium , 2007, pp. 214–223
2007
Earlier work this paper cites.
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in Computer Vision and Pattern Recognition , 2008
2008
Earlier work this paper cites.
E. H. Spriggs, F. De La Torre, and M. Hebert, “Temporal segmentation and activity classification from first-person sensing,” in Computer Vision and Pattern Recognition Workshops , 2009, pp. 17–24
2009
Earlier work this paper cites.
A. Fathi, A. Farhadi, and J. M. Rehg, “Understanding egocentric activities,” in International Conference on Computer Vision , 2011, pp. 407–414
2011
Earlier work this paper cites.
M. S. Ryoo, “Human activity prediction: Early recognition of ongoing activities from streaming videos,” in International Conference on Computer Vision , 2011, pp. 1036–1043
2011
Earlier work this paper cites.
T. Kanade and M. Hebert, “First-person vision,” Proceedings of the IEEE , vol. 100, no. 8, pp. 2442–2453, 2012
2012
Earlier work this paper cites.
A. Fathi, Y. Li, and J. Rehg, “Learning to recognize daily actions using gaze,” in European Conference on Computer Vision , 2012, pp. 314–327
2012
Earlier work this paper cites.
H. Pirsiavash and D. Ramanan, “Detecting activities of daily living in first-person camera views,” in Computer Vision and Pattern Recognition , 2012, pp. 2847–2854
2012
Earlier work this paper cites.
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert, “Activity forecasting,” in European Conference on Computer Vision . Springer, 2012, pp. 201–214
2012
Earlier work this paper cites.
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu, “Dense trajectories and motion boundary descriptors for action recognition,” International Journal of Computer Vision , vol. 103, no. 1, pp. 60–79, 2013
2013
Earlier work this paper cites.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in International Conference on Computer Vision , 2013, pp. 3551–3558
2013
Earlier work this paper cites.
Y. Cao, D. Barrett, A. Barbu, S. Narayanaswamy, H. Yu, A. Michaux, Y. Lin, S. Dickinson, J. Mark Siskind, and S. Wang, “Recognize human activities from partially observed videos,” in Computer Vision and Pattern Recognition , 2013, pp. 2658–2665
2013
Earlier work this paper cites.
D.-A. Huang and K. M. Kitani, “Action-reaction: Forecasting the dynamics of human interaction,” in European Conference on Computer Vision , 2014, pp. 489–504
2014
Earlier work this paper cites.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in Computer Vision and Pattern Recognition , 2014, pp. 1725–1732
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Advances in Neural Information Processing Systems , 2014, pp. 568–576
2014
Earlier work this paper cites.
M. Hoai and F. D. la Torre, “Max-margin early event detectors,” International Journal of Computer Vision , vol. 107, no. 2, pp. 191–202, 2014
2014
Earlier work this paper cites.
D. Huang, S. Yao, Y. Wang, and F. De La Torre, “Sequential max-margin event detectors,” in European conference on computer vision , 2014, pp. 410–424
2014
Earlier work this paper cites.
T. Lan, T.-C. Chen, and S. Savarese, “A hierarchical representation for future action prediction,” in European Conference on Computer Vision , 2014, pp. 689–704
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. Le, “Sequence to sequence learning with neural networks,” Advances in NIPS , 2014
2014
Earlier work this paper cites.
B. Soran, A. Farhadi, and L. Shapiro, “Generating notifications for missing actions: Don’t forget to turn the lights off!” in International Conference on Computer Vision , 2015, pp. 4669–4677
2015
Earlier work this paper cites.
M. S. Ryoo, T. J. Fuchs, L. Xia, J. K. Aggarwal, and L. Matthies, “Robot-centric activity prediction from first-person videos: What will they do to me?” in International Conference on Human-Robot Interaction , 2015, pp. 295–302
2015
Earlier work this paper cites.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 961–970
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in International Conference on Computer Vision , 2015, pp. 4489–4497
2015
Cited alongside, same era.
Y. Li, Z. Ye, and J. M. Rehg, “Delving into egocentric actions,” in Computer Vision and Pattern Recognition , 2015, pp. 287–295
2015
Cited alongside, same era.
M. S. Ryoo, B. Rothrock, and L. Matthies, “Pooled motion features for first-person videos,” in Computer Vision and Pattern Recognition , 2015, pp. 896–904
2015
Cited alongside, same era.
A. Jain, H. S. Koppula, B. Raghavan, S. Soh, and A. Saxena, “Car that knows before you do: Anticipating maneuvers via learning temporal driving models,” in International Conference on Computer Vision , 2015, pp. 3182–3190
2015
Cited alongside, same era.
K.-H. Zeng, W. B. Shen, D.-A. Huang, M. Sun, and J. Carlos Niebles, “Visual forecasting by imitating dynamics in natural sequences,” in International Conference on Computer Vision , 2017, pp. 2999–3008
2017
Later among the works it cites.
M. Zhang, K. T. Ma, J. H. Lim, Q. Zhao, and J. Feng, “Deep future gaze: Gaze anticipation on egocentric videos using adversarial networks.” in Computer Vision and Pattern Recognition , 2017, pp. 3539–3548
2017
Later among the works it cites.
A. Furnari, S. Battiato, K. Grauman, and G. M. Farinella, “Next-active-object prediction from egocentric videos,” Journal of Visual Communication and Image Representation , vol. 49, pp. 401–411, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning , 2015, pp. 448–456
2015
Cited alongside, same era.
R. Girshick, “Fast R-CNN,” in International Conference on Computer Vision , 2015, pp. 1440–1448
2015
Cited alongside, same era.
H. S. Koppula and A. Saxena, “Anticipating human activities using object affordances for reactive robotic response,” Pattern Analysis and Machine Intelligence , vol. 38, no. 1, pp. 14–29, 2016
2016
Cited alongside, same era.
R. De Geest, E. Gavves, A. Ghodrati, Z. Li, C. Snoek, and T. Tuytelaars, “Online action detection,” in European Conference on Computer Vision , 2016, pp. 269–284
2016
Cited alongside, same era.
S. Ma, L. Sigal, and S. Sclaroff, “Learning activity progression in lstms for activity detection and early detection,” in Computer Vision and Pattern Recognition , 2016
2016
Cited alongside, same era.
C. Vondrick, H. Pirsiavash, and A. Torralba, “Anticipating visual representations from unlabeled video,” in Computer Vision and Pattern Recognition , 2016, pp. 98–106
2016
Cited alongside, same era.
A. Jain, A. Singh, H. S. Koppula, S. Soh, and A. Saxena, “Recurrent neural networks for driver activity anticipation via sensory-fusion architecture,” in International Conference on Robotics and Automation . IEEE, 2016, pp. 3118–3125
2016
Cited alongside, same era.
N. Rhinehart and K. M. Kitani, “First-person activity forecasting with online inverse reinforcement learning,” in International Conference on Computer Vision , 2017
2017
Later among the works it cites.
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray, “Scaling egocentric vision: The epic-kitchens dataset,” in European Conference on Computer Vision , 2018, pp. 720–736
2018
Later among the works it cites.
A. Furnari, S. Battiato, and G. M. Farinella, “Leveraging uncertainty to rethink loss functions and evaluation measures for egocentric action anticipation,” in European Conference on Computer Vision Workshops , 2018
2018
Later among the works it cites.
Y. Abu Farha, A. Richard, and J. Gall, “When will you do what?-anticipating temporal occurrences of activities,” in Computer Vision and Pattern Recognition , 2018, pp. 5343–5352
2018
Later among the works it cites.
Y. Li, M. Liu, and J. M. Rehg, “In the eye of beholder: Joint learning of gaze and actions in first person video,” in European Conference on Computer Vision , 2018
2018
Later among the works it cites.
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” in European Conference on Computer Vision , 2018, pp. 803–818
2018
Later among the works it cites.
2018
Later among the works it cites.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6546–6555
2018
Later among the works it cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6450–6459
2018
Later among the works it cites.
S. Sudhakaran, S. Escalera, and O. Lanz, “Lsta: Long short-term attention for egocentric action recognition,” in Computer Vision and Pattern Recognition , 2018
2018
Later among the works it cites.
S. Sudhakaran and O. Lanz, “Attention is all we need: Nailing down object-centric attention for egocentric activity recognition,” British Machine Vision Conference , 2018
2018
Later among the works it cites.
R. D. Geest and T. Tuytelaars, “Modeling temporal structure with lstm for online action detection,” in Winter Conference on Applications in Computer Vision , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
P. Zhang, J. Xue, C. Lan, W. Zeng, Z. Gao, and N. Zheng, “Adding attentiveness to the neurons in recurrent neural networks,” in European Conference on Computer Vision , 2018, pp. 135–151
2018
Later among the works it cites.
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He, “Detectron,” https://github.com/facebookresearch/detectron, 2018
2018
Later among the works it cites.
Y. Ma, X. Zhu, S. Zhang, R. Yang, W. Wang, and D. Manocha, “Trafficpredict: Trajectory prediction for heterogeneous traffic-agents,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 6120–6127
2019
Later among the works it cites.
A. Miech, I. Laptev, J. Sivic, H. Wang, L. Torresani, and D. Tran, “Leveraging the present to anticipate the future in videos,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2019, pp. 0–0
2019
Later among the works it cites.
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 7083–7093
2019
Later among the works it cites.
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Epic-fusion: Audio-visual temporal binding for egocentric action recognition,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 5492–5501
2019
Later among the works it cites.
D. Ghadiyaram, D. Tran, and D. Mahajan, “Large-scale weakly-supervised pre-training for video action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 12 046–12 055
2019
Later among the works it cites.
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick, “Long-term feature banks for detailed video understanding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 284–293
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems , 2019, pp. 8024–8035
2019
Later among the works it cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning , 2015, pp. 2048–2057
2057
Closest in time.