Fetching the paper…
Reading the bibliography…
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge.
Flavell, J.H., Flavell, E.R., Green, F.L., Wilcox, S.A.: The development of three spatial perspective-taking rules. Child Development (1981)
1981
Earlier work this paper cites.
Newcombe, N.: The development of spatial perspective taking. Advances in child development and behavior (1989)
1989
Earlier work this paper cites.
Weinland, D., Ronfard, R., Boyer, E.: Free viewpoint action recognition using motion history volumes. Computer Vision and Image Understanding (CVIU) (2006)
2006
Earlier work this paper cites.
Torre, F.D., Hodgins, J., Montano, J., Valcarcel, S., Forcada, R., Macey, J.: Guide to the carnegie mellon university multimodal activity (cmu-mmac) database. In: Tech. Report CMU-RI-TR-08-22, Robotics Institute, Carnegie Mellon University (2009)
2009
Earlier work this paper cites.
Brodersen, K.H., Ong, C.S., Stephan, K.E., Buhmann, J.M.: The balanced accuracy and its posterior distribution. In: 2010 20th International Conference on Pattern Recognition, pp. 3121–3124 (2010). IEEE
2010
Earlier work this paper cites.
Hore, A., Ziou, D.: Image quality metrics: Psnr vs. ssim. In: 2010 20th International Conference on Pattern Recognition, pp. 2366–2369 (2010). IEEE
2010
Earlier work this paper cites.
Vicente, S., Rother, C., Kolmogorov, V.: Object cosegmentation. In: CVPR 2011, pp. 2217–2224 (2011). IEEE
2011
Earlier work this paper cites.
Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N.K., Hannemann, M., Motlícek, P., Qian, Y., Schwarz, P., Silovský, J., Stemmer, G., Veselý, K.: The kaldi speech recognition toolkit. (2011). https://api.semanticscholar.org/CorpusID:1774023
2011
Earlier work this paper cites.
Soomro, K., Zamir, A.R., Shah, M.: Ucf101: A dataset of 101 human action classes from videos in the wild. In: CRCV-TR-12-01 (2012)
2012
Earlier work this paper cites.
Lee, Y.J., Ghosh, J., Grauman, K.: Discovering important people and objects for egocentric video summarization. In: CVPR (2012)
2012
Earlier work this paper cites.
Pirsiavash, H., Ramanan, D.: Detecting activities of daily living in first-person camera views. In: CVPR (2012)
2012
Earlier work this paper cites.
Xiao, J., Owens, A., Torralba, A.: Sun3d: A database of big spaces reconstructed using sfm and object labels. In: ICCV (2013)
2013
Earlier work this paper cites.
Zhang, Q., Li, B.: Relative hidden markov models for evaluating motion skill. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2013)
2013
Earlier work this paper cites.
Pirsiavash, H., Vondrick, C., Torralba, A.: Assessing the quality of actions. In: ECCV (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Horowitz, M.: 1.1 computing’s energy problem (and what we can do about it). In: 2014 IEEE International Solid-state Circuits Conference Digest of Technical Papers (ISSCC) (2014)
2014
Earlier work this paper cites.
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., Xiao, J.: 3d shapenets: A deep representation for volumetric shapes. In: Computer Vision and Pattern Recognition, IEEE Conference On (2015)
2015
Earlier work this paper cites.
Lin, T., Cui, Y., Belongie, S., Hays, J.: Learning deep representations for ground-to-aerial geolocalization. In: CVPR (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Soran, B., Farhadi, A., Shapiro, L.: Generating notifications for missing actions: Don’t forget to turn the lights off! In: ICCV, pp. 4669–4677 (2015)
2015
Earlier work this paper cites.
Singh, K.K., Fatahalian, K., Efros, A.A.: Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks. In: WACV (2016)
2016
Earlier work this paper cites.
Alayrac, J.-B., Bojanowski, P., Agrawal, N., Sivic, J., Laptev, I., Lacoste-Julien, S.: Unsupervised learning from narrated instruction videos. In: CVPR (2016)
2016
Earlier work this paper cites.
Ardeshir, S., Borji, A.: Ego2top: Matching viewers in egocentric and top-view videos. In: ECCV (2016)
2016
Earlier work this paper cites.
Kukelova, Z., Heller, J., Fitzgibbon, A.: Efficient intersection of three quadrics and applications in computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1799–1808 (2016)
2016
Earlier work this paper cites.
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., Sorkine-Hornung, A.: A benchmark dataset and evaluation methodology for video object segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 724–732 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
De Geest, R., Gavves, E., Ghodrati, A., Li, Z., Snoek, C., Tuytelaars, T.: Online action detection. In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, pp. 269–284 (2016). Springer
2016
Earlier work this paper cites.
Rhodin, H., Richardt, C., Casas, D., Insafutdinov, E., Shafiei, M., Seidel, H.-P., Schiele, B., Theobalt, C.: Egocap: egocentric marker-less motion capture with two fisheye cameras. ACM Transactions on Graphics (TOG) 35
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Chang, A., Dai, A., Funkhouser, T., Nießner, M., Savva, M., Song, S., Zeng, A., Zhang, Y.: Matterport3d: Learning from rgb-d data in indoor environments. In: Proceedings of the International Conference on 3D Vision (3DV) (2017). MatterPort3D dataset license available at: http://kaldir.vc.in.tum.de/matterport/MP_TOS.pdf
2017
Earlier work this paper cites.
Joo, H., Simon, T., Li, X., Liu, H., Tan, L., Gui, L., Banerjee, S., Godisart, T.S., Nabbe, B., Matthews, I., Kanade, T., Nobuhara, S., Sheikh, Y.: Panoptic studio: A massively multiview system for social interaction capture. IEEE Transactions on Pattern Analysis and Machine Intelligence (2017)
2017
Earlier work this paper cites.
Bertasius, G., Park, H.S., Yu, S., Shi, J.: Am i a baller? basketball performance assessment from first-person videos. In: ICCV (2017)
2017
Earlier work this paper cites.
Fan, C., Lee, J., Xu, M., Kumar Singh, K., Jae Lee, Y., Crandall, D.J., Ryoo, M.S.: Identifying first-person camera wearers in third-person videos. In: CVPR (2017)
2017
Earlier work this paper cites.
Isola, P., Zhu, J.-Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. CVPR (2017)
2017
Earlier work this paper cites.
Sudre, C.H., Li, W., Vercauteren, T., Ourselin, S., Jorge Cardoso, M.: Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations. In: Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: Third International Workshop, DLMIA 2017, and 7th International Workshop, ML-CDS 2017, Held in Conjunction with MICCAI 2017, Québec City, QC, Canada, September 14, Proceedings 3, pp. 240–248 (2017). Springer
2017
Earlier work this paper cites.
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al
2017
Earlier work this paper cites.
Sermanet, P., Lynch, C., Hsu, J., Levine, S.: Time-contrastive networks: Self-supervised learning from multi-view observation. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 486–487 (2017). IEEE
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Parmar, P., Morris, B.T.: Learning To Score Olympic Events (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Jiang, H., Grauman, K.: Seeing invisible poses: Estimating 3d body pose from egocentric video. In: CVPR (2017)
2017
Earlier work this paper cites.
Simon, T., Joo, H., Matthews, I., Sheikh, Y.: Hand keypoint detection in single images using multiview bootstrapping. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1145–1153 (2017)
2017
Earlier work this paper cites.
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2117–2125 (2017)
2017
Earlier work this paper cites.
Romero, J., Tzionas, D., Black, M.J.: Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia) 36
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., Schmid, C., Malik, J.: Ava: A video dataset of spatio-temporally localized atomic visual actions. In: CVPR (2018)
2018
Earlier work this paper cites.
Zhou, L., Louis, N., Corso, J.: Weakly-supervised video object grounding from text by loss weighting and object interaction. In: BMVC (2018)
2018
Earlier work this paper cites.
Li, Y., Liu, M., Rehg, J.M.: In the eye of beholder: Joint learning of gaze and actions in first person video. In: ECCV (2018)
2018
Earlier work this paper cites.
Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., Alahari, K.: Actor and observer: Joint modeling of first and third-person videos. In: CVPR (2018)
2018
Earlier work this paper cites.
Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: Scaling egocentric vision: The epic-kitchens dataset. In: European Conference on Computer Vision (ECCV) (2018)
2018
Earlier work this paper cites.
Xia, F., R. Zamir, A., He, Z.-Y., Sax, A., Malik, J., Savarese, S.: Gibson Env: real-world perception for embodied agents. In: CVPR (2018). IEEE. Gibson license is available at http://svl.stanford.edu/gibson2/assets/GDS_agreement.pdf
2018
Earlier work this paper cites.
Doughty, H., Damen, D., Mayol-Cuevas, W.: Who’s better? who’s best? pairwise deep ranking for skill determination. In: CVPR (2018)
2018
Earlier work this paper cites.
Zhou, L., Xu, C., Corso, J.J.: Towards automatic learning of procedures from web instructional videos. In: AAAI (2018)
2018
Earlier work this paper cites.
Ardeshir, S., Borji, A.: Egocentric meets top-view. IEEE transactions on pattern analysis and machine intelligence 41
2018
Earlier work this paper cites.
Xu, M., Fan, C., Wang, Y., Ryoo, M.S., Crandall, D.J.: Joint person segmentation and identification in synchronized first-and third-person videos. In: ECCV (2018)
2018
Earlier work this paper cites.
Ardeshir, S., Borji, A.: An exocentric look at egocentric actions and vice versa. Computer Vision and Image Understanding 171
2018
Earlier work this paper cites.
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S.: Time-contrastive networks: Self-supervised learning from video. Proceedings of International Conference in Robotics and Automation (ICRA) (2018)
2018
Earlier work this paper cites.
Regmi, K., Borji, A.: Cross-view image synthesis using conditional gans. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Earlier work this paper cites.
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: CVPR (2018)
2018
Earlier work this paper cites.
Perez, E., Strub, F., De Vries, H., Dumoulin, V., Courville, A.: Film: Visual reasoning with a general conditioning layer. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Zhang, X., Zhou, X., Lin, M., Sun, J.: Shufflenet: An extremely efficient convolutional neural network for mobile devices. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6848–6856 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Wu, Z., Nagarajan, T., Kumar, A., Rennie, S., Davis, L.S., Grauman, K., Feris, R.: Blockdrop: Dynamic inference paths in residual networks. In: CVPR (2018)
2018
Earlier work this paper cites.
Ismail Fawaz, H., Forestier, G., Weber, J., Idoumghar, L., Muller, P.-A.: Evaluating surgical skills from kinematic data using convolutional neural networks. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2018, pp. 214–221 (2018)
2018
Earlier work this paper cites.
Yuan, Y., Kitani, K.: 3d ego-pose estimation via imitation learning. In: Proceedings of the European Conference on Computer Vision (ECCV) (2018)
2018
Earlier work this paper cites.
Monfort, M., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S.A., Yan, T., Brown, L., Fan, Q., Gutfreund, D., Vondrick, C., Oliva, A.: Moments in time dataset: one million videos for event understanding. PAMI (2019)
2019
Earlier work this paper cites.
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., Sivic, J.: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. In: ICCV (2019)
2019
Earlier work this paper cites.
Zhukov, D., Alayrac, J.-B., Cinbis, R.G., Fouhey, D., Laptev, I., Sivic, J.: Cross-task weakly supervised learning from instructional videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Parmar, P., Morris, B.: Action quality assessment across multiple actions. In: WACV (2019)
2019
Cited alongside, same era.
Doughty, H., Mayol-Cuevas, W., Damen, D.: The Pros and Cons: Rank-aware Temporal Attention for Skill Determination in Long Videos (2019)
2019
Cited alongside, same era.
Yu, H., Cai, M., Liu, Y., Lu, F.: What i see is what you see: Joint attention learning for first and third person video co-analysis. In: ACM MM (2019)
2019
Cited alongside, same era.
Regmi, K., Borji, A.: Cross-view image synthesis using geometry-guided conditional gans. Computer Vision and Image Understanding (2019) https://doi.org/10.1016/j.cviu.2019.07.008
2019
Cited alongside, same era.
Tang, H., Xu, D., Sebe, N., Wang, Y., Corso, J.J., Yan, Y.: Multi-channel attention selection gan with cascaded semantic guidance for cross-view image translation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2417–2426 (2019)
Lin, K., Wang, L., Liu, Z.: End-to-end human pose and mesh reconstruction with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1954–1963 (2021)
2021
Later among the works it cites.
Xu, M., Xiong, Y., Chen, H., Li, X., Xia, W., Tu, Z., Soatto, S.: Long short-term transformer for online action detection. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Li, J., Bian, S., Zeng, A., Wang, C., Pang, B., Liu, W., Lu, C.: Human pose regression with residual log-likelihood estimation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11025–11034 (2021)
2021
Later among the works it cites.
Sener, F., Chatterjee, D., Shelepov, D., He, K., Singhania, D., Wang, R., Yao, A.: Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 21096–21106 (2022)
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Regmi, K., Shah, M.: Bridging the domain gap for ground-to-aerial image matching. In: ICCV (2019)
2019
Cited alongside, same era.
Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., Le, Q.V.: Mnasnet: Platform-aware neural architecture search for mobile. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2820–2828 (2019)
2019
Cited alongside, same era.
Korbar, B., Tran, D., Torresani, L.: Scsampler: Sampling salient clips from video for efficient action recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Chen, Y.-H., Yang, T.-J., Emer, J., Sze, V.: Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices. IEEE Journal on Emerging and Selected Topics in Circuits and Systems (2019)
2019
Cited alongside, same era.
Parmar, P., Tran Morris, B.: What and how well you performed? a multitask learning approach to action quality assessment. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 304–313 (2019)
2019
Cited alongside, same era.
Yuan, Y., Kitani, K.: Ego-pose estimation and forecasting as real-time pd control. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2019)
2019
Cited alongside, same era.
Later among the works it cites.
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., Martin, M., Nagarajan, T., Radosavovic, I., Ramakrishnan, S.K., Ryan, F., Sharma, J., Wray, M., Xu, M., Xu, E.Z., Zhao, C., Bansal, S., Batra, D., Cartillier, V., Crane, S., Do, T., Doulaty, M., Erapalli, A., Feichtenhofer, C., Fragomeni, A., Fu, Q., Gebreselasie, A., González, C., Hillis, J., Huang, X., Huang, Y., Jia, W., Khoo, W., Kolář, J., Kottur, S., Kumar, A., Landini, F., Li, C., Li, Y., Li, Z., Mangalam, K., Modhugu, R., Munro, J., Murrell, T., Nishiyasu, T., Price, W., Ruiz, P., Ramazanova, M., Sari, L., Somasundaram, K., Southerland, A., Sugano, Y., Tao, R., Vo, M., Wang, Y., Wu, X., Yagi, T., Zhao, Z., Zhu, Y., Arbeláez, P., Crandall, D., Damen, D., Farinella, G.M., Fuegen, C., Ghanem, B., Ithapu, V.K., Jawahar, C.V., Joo, H., Kitani, K., Li, H., Newcombe, R., Oliva, A., Park, H.S., Rehg, J.M., Sato, Y., Shi, J., Shou, M.Z., Torralba, A., Torresani, L., Yan, M., Malik, J.: Ego4D: Around the world in 3,000 hours of egocentric video. In: CVPR (2022)
2022
Later among the works it cites.
Damen, D., Doughty, H., Farinella, G.M., Furnari, A., Kazakos, E., Ma, J., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al
2022
Later among the works it cites.
Wong, B., Chen, J., Wu, Y., Lei, S.W., Mao, D., Gao, D., Shou, M.Z.: Assistq: Affordance-centric question-driven task completion for egocentric assistant. In: European Conference on Computer Vision (2022)
2022
Later among the works it cites.
Bansal, S., Arora, C., Jawahar, C.V.: My view is the best view: Procedure learning from egocentric videos. In: European Conference on Computer Vision (ECCV) (2022)
2022
Later among the works it cites.
Zhang, S., Ma, Q., Zhang, Y., Qian, Z., Kwon, T., Pollefeys, M., Bogo, F., Tang, S.: Egobody: Human body shape and motion of interacting people from head-mounted devices. In: ECCV (2022)
2022
Later among the works it cites.
Dvornik, N., Hadji, I., Pham, H., Bhatt, D., Martinez, B., Fazly, A., Jepson, A.D.: Flow graph to video grounding for weakly-supervised multi-step localization. In: ECCV, pp. 319–335 (2022). Springer
2022
Later among the works it cites.
Lin, X., Petroni, F., Bertasius, G., Rohrbach, M., Chang, S.-F., Torresani, L.: Learning to recognize procedural activities with distant supervision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13853–13863 (2022)
2022
Later among the works it cites.
Zhao, H., Hadji, I., Dvornik, N., Derpanis, K.G., Wildes, R.P., Jepson, A.D.: P3iv: Probabilistic procedure planning from instructional videos with weak supervision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2938–2948 (2022)
2022
Later among the works it cites.
Shvetsova, N., Chen, B., Rouditchenko, A., Thomas, S., Kingsbury, B., Feris, R.S., Harwath, D., Glass, J., Kuehne, H.: Everything at once-multi-modal fusion transformer for video retrieval. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20020–20029 (2022)
2022
Later among the works it cites.
Ko, D., Choi, J., Ko, J., Noh, S., On, K.-W., Kim, E.-S., Kim, H.J.: Video-text representation learning via differentiable weak temporal alignment. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5016–5025 (2022)
2022
Later among the works it cites.
Cao, M., Yang, T., Weng, J., Zhang, C., Wang, J., Zou, Y.: Locvtp: Video-text pre-training for temporal localization. In: European Conference on Computer Vision (2022)
2022
Later among the works it cites.
Ren, X., Wang, X.: Look outside the room: Synthesizing a consistent long-term 3d scene video from a single image. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3563–3573 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Lin, K.Q., Wang, A.J., Soldan, M., Wray, M., Yan, R., Xu, E.Z., Gao, D., Tu, R., Zhao, W., Kong, W., et al.: Egocentric video-language pretraining. NeurIPS (2022)
2022
Later among the works it cites.
Shen, X., Efros, A.A., Joulin, A., Aubry, M.: Learning co-segmentation by segment swapping for retrieval and discovery. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5082–5092 (2022)
2022
Later among the works it cites.
Cheng, H.K., Schwing, A.G.: Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model. In: European Conference on Computer Vision, pp. 640–658 (2022). Springer
2022
Later among the works it cites.
2022
Later among the works it cites.
Mavroudi, E., Afouras, T., Torresani, L.: Learning to ground instructional articles in videos through narrations. (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Yang, L., Radway, R.M., Chen, Y.-H., Wu, T.F., Liu, H., Ansari, E., Chandra, V., Mitra, S., Beigné, E.: Three-dimensional stacked neural network accelerator architectures for ar/vr applications. IEEE Micro (2022)
2022
Later among the works it cites.
Zhang, C.-L., Wu, J., Li, Y.: Actionformer: Localizing moments of actions with transformers. In: European Conference on Computer Vision. LNCS, vol. 13664, pp. 492–510 (2022)
2022
Later among the works it cites.
Girdhar, R., Singh, M., Ravi, N., Maaten, L., Joulin, A., Misra, I.: Omnivore: A single model for many visual modalities. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16102–16112 (2022)
2022
Later among the works it cites.
Jiang, J., Streli, P., Qiu, H., Fender, A., Laich, L., Snape, P., Holz, C.: Avatarposer: Articulated full-body pose tracking from sparse motion sensing. In: European Conference on Computer Vision, pp. 443–460 (2022). Springer
2022
Later among the works it cites.
Zhao, W., Wang, W., Tian, Y.: Graformer: Graph-oriented transformer for 3d pose estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20438–20447 (2022)
2022
Later among the works it cites.
Park, J., Oh, Y., Moon, G., Choi, H., Lee, K.M.: Handoccnet: Occlusion-robust 3d hand mesh estimation network. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1496–1505 (2022)
2022
Later among the works it cites.
Zhao, Y., Krähenbühl, P.: Real-time online video detection with temporal smoothing transformers. In: European Conference on Computer Vision (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Engel, J., Somasundaram, K., Goesele, M., Sun, A., Gamino, A., Turner, A., Talattof, A., Yuan, A., Souti, B., Meredith, B., Peng, C., Sweeney, C., Wilson, C., Barnes, D., DeTone, D., Caruso, D., Valleroy, D., Ginjupalli, D., Frost, D., Miller, E., Mueggler, E., Oleinik, E., Zhang, F., Somasundaram, G., Solaira, G., Lanaras, H., Howard-Jenkins, H., Tang, H., Kim, H.J., Rivera, J., Luo, J., Dong, J., Straub, J., Bailey, K., Eckenhoff, K., Ma, L., Pesqueira, L., Schwesinger, M., Monge, M., Yang, N., Charron, N., Raina, N., Parkhi, O., Borschowa, P., Moulon, P., Gupta, P., Mur-Artal, R., Pennington, R., Kulkarni, S., Miglani, S., Gondi, S., Solanki, S., Diener, S., Cheng, S., Green, S., Saarinen, S., Patra, S., Mourikis, T., Whelan, T., Singh, T., Balntas, V., Baiyya, V., Dreewes, W., Pan, X., Lou, Y., Zhao, Y., Mansour, Y., Zou, Y., Lv, Z., Wang, Z., Yan, M., Ren, C., Nardi, R.D., Newcombe, R.: Project Aria: A New Tool for Egocentric Multi-Modal AI Research (2023)
2023
Closest in time.
Wang, X., Kwon, T., Rad, M., Pan, B., Chakraborty, I., Andrist, S., Bohus, D., Feniello, A., Tekin, B., Frujeri, F.V., et al
2023
Closest in time.
Ohkawa, T., He, K., Sener, F., Hodan, T., Tran, L., Keskin, C.: Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12999–13008 (2023)
2023
Closest in time.
Tschernezki, V., Darkhalil, A., Zhu, Z., Fouhey, D., Larina, I., Larlus, D., Damen, D., Vedaldi, A.: EPIC Fields: Marrying 3D Geometry and Video Understanding. In: Proceedings of the Neural Information Processing Systems (NeurIPS) (2023)
2023
Closest in time.
Li, J., Liu, K., Wu, J.: Ego-body pose estimation via ego-head pose estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17142–17151 (2023)
2023
Closest in time.
Khirodkar, R., Bansal, A., Ma, L., Newcombe, R., Vo, M., Kitani, K.: Egohumans: An egocentric 3d multi-human benchmark. In: ICCV (2023)
2023
Closest in time.
Zhang, S., Dai, W., Wang, S., Shen, X., Lu, J., Zhou, J., Tang, Y.: Logo: A long-form video dataset for group action quality assessment. In: CVPR (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Ashutosh, K., Ramakrishnan, S.K., Afouras, T., Grauman, K.: Video-mined task graphs for keystep recognition in instructional videos. In: NeurIPS (2023)
2023
Closest in time.
Zhou, H., Martin-Martin, R., Kapadia, M., Savarese, S., Niebles, J.C.: Procedure-aware pretraining for instructional video understanding. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2023)
2023
Closest in time.
Xue, Z., Grauman, K.: Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment. In: NeurIPS (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Pramanick, S., Song, Y., Nag, S., Lin, K.Q., Shah, H., Shou, M.Z., Chellappa, R., Zhang, P.: Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5285–5297 (2023)
2023
Closest in time.
Zhao, Y., Misra, I., Krähenbühl, P., Girdhar, R.: Learning video representations from large language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6586–6597 (2023)
2023
Closest in time.
Radford, A., Kim, J.W., Xu, T., Brockman, G., McLeavey, C., Sutskever, I.: Robust speech recognition via large-scale weak supervision. In: International Conference on Machine Learning, pp. 28492–28518 (2023). PMLR
2023
Closest in time.
Ashutosh, K., Girdhar, R., Torresani, L., Grauman, K.: Hiervl: Learning hierarchical video-language embeddings. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2023)
2023
Closest in time.
Tang, H., Liang, K., Grauman, K., Feiszli, M., Wang, W.: Egotracks: A long-term egocentric visual object tracking dataset. Advances in Neural Information Processing Systems (2023)
2023
Closest in time.
Varma, M., Wang, P., Chen, X., Chen, T., Venugopalan, S., Wang, Z.: Is attention all that neRF needs? In: The Eleventh International Conference on Learning Representations (2023). https://openreview.net/forum?id=xE-LtsE-xx
2023
Closest in time.
Song, Y., Byrne, E., Nagarajan, T., Wang, H., Martin, M., Torresani, L.: Ego4d goal-step: Toward hierarchical understanding of procedural activities. In: NeurIPS (2023)
2023
Closest in time.
Vasu, P.K.A., Gabriel, J., Zhu, J., Tuzel, O., Ranjan, A.: Mobileone: An improved one millisecond mobile backbone. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7907–7917 (2023)
2023
Closest in time.
Tan, S., Nagarajan, T., Grauman, K.: Egodistill: Egocentric head motion distillation for efficient video understanding. NeurIPS (2023)
2023
Closest in time.
Desislavov, R., Martínez-Plumed, F., Hernández-Orallo, J.: Trends in ai inference energy consumption: Beyond the performance-vs-parameter laws of deep learning. Sustainable Computing: Informatics and Systems 38
2023
Closest in time.
Liao, J., Duan, H., Feng, K., Zhao, W., Yang, Y., Chen, L.: A light weight model for active speaker detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22932–22941 (2023)
2023
Closest in time.
Zhou, H., Martín-Martín, R., Kapadia, M., Savarese, S., Niebles, J.C.: Procedure-aware pretraining for instructional video understanding. In: CVPR, pp. 10727–10738 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Castillo, A., Escobar, M., Jeanneret, G., Pumarola, A., Arbeláez, P., Thabet, A., Sanakoyeu, A.: Bodiffusion: Diffusing sparse observations for full-body human motion synthesis. CV4Metaverse workshop, International Conference on Computer Vision (2023)
2023
Closest in time.
Aboukhadra, A.T., Malik, J., Elhayek, A., Robertini, N., Stricker, D.: Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1001–1010 (2023)
2023
Closest in time.
Zheng, C., Liu, X., Qi, G.-J., Chen, C.: Potter: Pooling attention transformer for efficient human mesh recovery. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1611–1620 (2023)
2023
Closest in time.
Tendulkar, P., Surís, D., Vondrick, C.: Flex: Full-body grasping without full-body grasps. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023)
2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.-Y., Dollar, P., Girshick, R.: Segment anything. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4015–4026 (2023)
2023
Closest in time.
Huang, Y., Chen, G., Xu, J., Zhang, M., Yang, L., Pei, B., Zhang, H., Dong, L., Wang, Y., Wang, L., et al
2024
Closest in time.
Luo, M., Xue, Z., Dimakis, A., Grauman, K.: Put myself in your shoes: Lifting the egocentric perspective from exocentric videos. In: ECCV (2024)
2024
Closest in time.
Cheng, F., Luo, M., Wang, H., Dimakis, A., Torresani, L., Bertasius, G., Grauman, K.: 4DIFF: 3d-aware diffusion model for third-to-first viewpoint translation. In: ECCV (2024)
2024
Closest in time.
2024
Closest in time.
Seminara, L., Farinella, G.M., Furnari, A.: Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos (2024)
2024
Closest in time.
Plizzari, C., Goletto, G., Furnari, A., Bansal, S., Ragusa, F., Farinella, G.M., Damen, D., Tommasi, T.: An outlook into the future of egocentric vision. International Journal of Computer Vision (2024)
2024
Closest in time.
Ashutosh, K., Nagarajan, T., Pavlakos, G., Kitani, K., Grauman, K.: ExpertAF: Expert Actionable Feedback from Video (2024)
2024
Closest in time.
2024
Closest in time.
Pavlakos, G., Shan, D., Radosavovic, I., Kanazawa, A., Fouhey, D., Malik, J.: Reconstructing hands in 3d with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9826–9836 (2024)
2024
Closest in time.