Fetching the paper…
Reading the bibliography…
Given a video captured from a first person perspective and the environment context of where the video is recorded, can we recognize what the person is doing and identify where the action occurs in the 3D space? We address this challenging problem of jointly recognizing and localizing actions of a mobile user on a known 3D map from egocentric videos.
Frahm, J.M., Fite-Georgel, P., Gallup, D., Johnson, T., Raguram, R., Wu, C., Jen, Y.H., Dunn, E., Clipp, B., Lazebnik, S., Pollefeys, M.: Building rome on a cloudless day. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) Computer Vision – ECCV 2010. pp. 368–381. Springer Berlin Heidelberg, Berlin, Heidelberg (2010)
2010
Earlier work this paper cites.
Grabner, H., Gall, J., Van Gool, L.: What makes a chair a chair? In: CVPR (2011)
2011
Earlier work this paper cites.
Gupta, A., Satkin, S., Efros, A.A., Hebert, M.: From 3d scene geometry to human workspace. In: CVPR (2011)
2011
Earlier work this paper cites.
Alahi, A., Ortiz, R., Vandergheynst, P.: Freak: Fast retina keypoint. In: CVPR (2012)
2012
Earlier work this paper cites.
Delaitre, V., Fouhey, D.F., Laptev, I., Sivic, J., Gupta, A., Efros, A.A.: Scene semantics from long-term observation of people. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) Computer Vision – ECCV 2012. pp. 284–298. Springer Berlin Heidelberg, Berlin, Heidelberg (2012)
2012
Earlier work this paper cites.
Jiang, Y., Lim, M., Saxena, A.: Learning object arrangements in 3d scenes using human context. In: ICML (2012)
2012
Earlier work this paper cites.
Park, H., Jain, E., Sheikh, Y.: 3d social saliency from head-mounted cameras. NeurIPS (2012)
2012
Earlier work this paper cites.
Pirsiavash, H., Ramanan, D.: Detecting activities of daily living in first-person camera views. In: CVPR (2012)
2012
Earlier work this paper cites.
Sattler, T., Leibe, B., Kobbelt, L.: Improving image-based localization by active correspondence search. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) Computer Vision – ECCV 2012. pp. 752–765. Springer Berlin Heidelberg, Berlin, Heidelberg (2012)
2012
Earlier work this paper cites.
Jiang, Y., Koppula, H., Saxena, A.: Hallucinated humans as the hidden context for labeling 3d scenes. In: CVPR (2013)
2013
Earlier work this paper cites.
Koppula, H.S., Gupta, R., Saxena, A.: Learning human activities and object affordances from rgb-d videos. The International Journal of Robotics Research 32
2013
Earlier work this paper cites.
Li, Y., Fathi, A., Rehg, J.M.: Learning to predict gaze in egocentric video. In: ICCV (2013)
2013
Earlier work this paper cites.
Serra, G., Camurri, M., Baraldi, L., Benedetti, M., Cucchiara, R.: Hand segmentation for gesture recognition in ego-vision. In: Proceedings of the 3rd ACM international workshop on Interactive multimedia on mobile & portable devices. pp. 31–36 (2013)
2013
Earlier work this paper cites.
Fouhey, D.F., Delaitre, V., Gupta, A., Efros, A.A., Laptev, I., Sivic, J.: People watching: Human actions as a cue for single view geometry. IJCV (2014)
2014
Earlier work this paper cites.
Kundu, A., Li, Y., Dellaert, F., Li, F., Rehg, J.M.: Joint semantic segmentation and 3d reconstruction from monocular video. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) Computer Vision – ECCV 2014. pp. 703–718. Springer International Publishing, Cham (2014)
2014
Earlier work this paper cites.
Poleg, Y., Arora, C., Peleg, S.: Head motion signatures from egocentric videos. In: ACCV (2014)
2014
Earlier work this paper cites.
Poleg, Y., Arora, C., Peleg, S.: Temporal segmentation of egocentric videos. In: CVPR (2014)
2014
Earlier work this paper cites.
Savva, M., Chang, A.X., Hanrahan, P., Fisher, M., Nießner, M.: Scenegrok: Inferring action maps in 3d environments. TOG (2014)
2014
Earlier work this paper cites.
Song, S., Xiao, J.: Sliding shapes for 3d object detection in depth images. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) Computer Vision – ECCV 2014. pp. 634–651. Springer International Publishing, Cham (2014)
2014
Earlier work this paper cites.
Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: ICML (2015)
2015
Earlier work this paper cites.
Li, Y., Ye, Z., Rehg, J.M.: Delving into egocentric actions. In: CVPR (2015)
2015
Earlier work this paper cites.
Wang, D.Z., Posner, I.: Voting for voting in online point cloud object detection. In: TSS (2015)
2015
Earlier work this paper cites.
Ma, M., Fan, H., Kitani, K.M.: Going deeper into first-person activity recognition. In: CVPR (2016)
2016
Earlier work this paper cites.
Rhinehart, N., Kitani, K.M.: Learning action maps of large environments via first-person vision. In: CVPR (2016)
2016
Earlier work this paper cites.
Schönberger, J.L., Frahm, J.M.: Structure-from-motion revisited. In: CVPR (2016)
2016
Cited alongside, same era.
Song, S., Xiao, J.: Deep sliding shapes for amodal 3d object detection in rgb-d images. In: CVPR (2016)
2016
Cited alongside, same era.
Soo Park, H., Hwang, J.J., Niu, Y., Shi, J.: Egocentric future localization. In: CVPR (2016)
2016
Cited alongside, same era.
Zhou, Y., Ni, B., Hong, R., Yang, X., Tian, Q.: Cascaded interactional targeting network for egocentric video analysis. In: CVPR (2016)
2016
Cited alongside, same era.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: CVPR (2017)
2017
Cited alongside, same era.
Furnari, A., Farinella, G.M.: What would you expect? anticipating egocentric actions with rolling-unrolling LSTMs and modality attention. In: ICCV (2019)
2019
Later among the works it cites.
Gordon, D., Kadian, A., Parikh, D., Hoffman, J., Batra, D.: Splitnet: Sim2sim and task2task transfer for embodied visual navigation. In: ICCV (2019)
2019
Later among the works it cites.
Hassan, M., Choutas, V., Tzionas, D., Black, M.J.: Resolving 3D human pose ambiguities with 3D scene constraints. In: ICCV (2019)
2019
Later among the works it cites.
Kazakos, E., Nagrani, A., Zisserman, A., Damen, D.: Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In: ICCV (2019)
2019
Later among the works it cites.
Ke, Q., Fritz, M., Schiele, B.: Time-conditioned action anticipation in one shot. In: CVPR (2019)
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Engelcke, M., Rao, D., Wang, D.Z., Tong, C.H., Posner, I.: Vote3deep: Fast object detection in 3d point clouds using efficient convolutional neural networks. In: ICRA (2017)
2017
Cited alongside, same era.
Jang, E., Gu, S., Poole, B.: Categorical reparameterization with gumbel-softmax. In: ICLR (2017)
2017
Cited alongside, same era.
Karthika, S., Praveena, P., GokilaMani, M.: Hololens. International Journal of Computer Science and Mobile Computing 6
2017
Cited alongside, same era.
Maddison, C.J., Mnih, A., Teh, Y.W.: The concrete distribution: A continuous relaxation of discrete random variables. In: ICLR (2017)
2017
Cited alongside, same era.
Moltisanti, D., Wray, M., Mayol-Cuevas, W., Damen, D.: Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video. In: ICCV (2017)
2017
Cited alongside, same era.
Qi, C.R., Su, H., Mo, K., Guibas, L.J.: Pointnet: Deep learning on point sets for 3d classification and segmentation. In: CVPR (2017)
2017
Cited alongside, same era.
Nagarajan, T., Feichtenhofer, C., Grauman, K.: Grounded human-object interaction hotspots from video. In: ICCV (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Wu, W., Qi, Z., Fuxin, L.: Pointconv: Deep convolutional networks on 3d point clouds. In: CVPR (2019)
2019
Later among the works it cites.
Wu, Y., Kirillov, A., Massa, F., Lo, W.Y., Girshick, R.: Detectron2. https://github.com/facebookresearch/detectron2 (2019)
2019
Later among the works it cites.
Damen, D., Doughty, H., Farinella, G., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.: The epic-kitchens dataset: Collection, challenges and baselines. IEEE Computer Architecture Letters (01), 1–1 (2020)
2020
Later among the works it cites.
Fan, H., Li, Y., Xiong, B., Lo, W.Y., Feichtenhofer, C.: Pyslowfast. https://github.com/facebookresearch/slowfast (2020)
2020
Later among the works it cites.
Guan, J., Yuan, Y., Kitani, K.M., Rhinehart, N.: Generative hybrid representations for activity forecasting with no-regret learning. In: CVPR (2020)
2020
Later among the works it cites.
Li, J., Wang, X., Tang, S., Shi, H., Wu, F., Zhuang, Y., Wang, W.Y.: Unsupervised reinforcement learning of transferable meta-skills for embodied navigation. In: CVPR (2020)
2020
Later among the works it cites.
Liu, M., Tang, S., Li, Y., Rehg, J.M.: Forecasting human-object interaction: Joint prediction of motor attention and actions in first person video. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J.M. (eds.) Computer Vision – ECCV 2020. pp. 704–721. Springer International Publishing, Cham (2020)
2020
Later among the works it cites.
Nagarajan, T., Li, Y., Feichtenhofer, C., Grauman, K.: Ego-topo: Environment affordances from egocentric video. In: CVPR (2020)
2020
Later among the works it cites.
Ng, E., Xiang, D., Joo, H., Grauman, K.: You2me: Inferring body pose in egocentric video via first and second person interactions. In: CVPR (2020)
2020
Later among the works it cites.
Sulaiman, M.Z., Aziz, M.N.A., Bakar, M.H.A., Halili, N.A., Azuddin, M.A.: Matterport: virtual tour as a new marketing approach in real estate business during pandemic covid-19. In: Proceedings of the International Conference of Innovation in Media and Visual Design (IMDES 2020). Atlantis Press. pp. 221–226 (2020)
2020
Later among the works it cites.
Wu, W., Wang, Z.Y., Li, Z., Liu, W., Fuxin, L.: Pointpwc-net: Cost volume on point clouds for (self-)supervised scene flow estimation. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J.M. (eds.) Computer Vision – ECCV 2020. pp. 88–107. Springer International Publishing, Cham (2020)
2020
Later among the works it cites.
Zhang, S., Zhang, Y., Ma, Q., Black, M.J., Tang, S.: PLACE: Proximity learning of articulation and contact in 3D environments. In: 3DV (2020)
2020
Later among the works it cites.
Zhang, Y., Hassan, M., Neumann, H., Black, M.J., Tang, S.: Generating 3d people in scenes without people. In: CVPR (2020)
2020
Later among the works it cites.
2021
Closest in time.
Li, Y., Liu, M., Rehg, J.M.: In the eye of the beholder: Gaze and actions in first person video. TPAMI (2021)
2021
Closest in time.
Liu, M., Yang, D., Zhang, Y., Cui, Z., Rehg, J.M., Tang, S.: 4d human body capture from egocentric video via 3d scene grounding. 3DV (2021)
2021
Closest in time.
Zhou, Y., Tuzel, O.: Voxelnet: End-to-end learning for point cloud based 3d object detection. In: CVPR (2018)
2022
Closest in time.