Fetching the paper…
Reading the bibliography…
This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS.
Sample Selection Bias as a Specification Error
Heckman, J.J.: · 1979
Earlier work this paper cites.
Domain adaptation with clustered language models
Ueberla, J.P.: · 1997
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
Järvelin, K., Kekäläinen, J.: · 2002
Earlier work this paper cites.
A duality based approach for realtime TV-L1 optical flow
Zach, C., Pock, T., Bischof, H.: · 2007
Earlier work this paper cites.
Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) database
De La Torre, F., Hodgins, J., Bargteil, A., Martin, X., Macey, J., Collado, A., Beltran, P.: · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
Actions in context
Marszalek, M., Laptev, I., Schmid, C.: · 2009
Earlier work this paper cites.
Adapting Visual Category Models to New Domains
Saenko, K., Kulis, B., Fritz, M., Darrell, T.: · 2010
Earlier work this paper cites.
High Five: Recognising human interactions in TV shows
Patron-Perez, A., Marszalek, M., Zisserman, A., Reid, I.: · 2010
Earlier work this paper cites.
Adapting visual category models to new domains
Saenko, K., Kulis, B., Fritz, M., Darrell, T.: · 2010
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A., Efros, A.A.: · 2011
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: · 2011
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D., Dolan, W.: · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? The KITTI vision benchmark suite
Geiger, A., Lenz, P., Urtasun, R.: · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Silberman, N., Hoiem, D., Kohli, P., Fergus, R.: · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Rohrbach, M., Amin, S., Andriluka, M., Schiele, B.: · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
Pirsiavash, H., Ramanan, D.: · 2012
Earlier work this paper cites.
Learning to recognize daily actions using gaze
Fathi, A., Li, Y., Rehg, J.: · 2012
Earlier work this paper cites.
Geodesic Flow Kernel for Unsupervised Domain Adaptation
Gong, B., Shi, Y., Sha, F., Grauman, K.: · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M.: · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., Dean, J.: · 2013
Earlier work this paper cites.
Combining Embedded Accelerometers with Computer Vision for Recognizing Food Preparation Activities
Stein, S., McKenna, S.J.: · 2013
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Stein, S., McKenna, S.: · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Earlier work this paper cites.
You-do, I-learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video
Damen, D., Leelasawassuk, T., Haines, O., Calway, A., Mayol-Cuevas, W.: · 2014
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Jiang, Y.G., Liu, J., Zamir, A.R., Toderici, G., Laptev, I., Shah, M., Sukthankar, R.: · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Kuehne, H., Arslan, A., Serre, T.: · 2014
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
Bojanowski, P., Lajugie, R., Bach, F., Laptev, I., Ponce, J., Schmid, C., Sivic, J.: · 2014
Earlier work this paper cites.
Imageclef 2014: Overview and analysis of the results
Caputo, B., Müller, H., Martinez-Gomez, J., Villegas, M., Acar, B., Patricia, N., Marvasti, N., Üsküdarlı, S., Paredes, R., Cazorla, M., et al.: · 2014
Earlier work this paper cites.
Cluster canonical correlation analysis
Rasiwasia, N., Mahajan, D., Mahadevan, V., Aggarwal, G.: · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D.P., Ba, J.: · 2014
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Karpathy, A., Fei-Fei, L.: · 2015
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
Heilbron, F.C., Escorcia, V., Ghanem, B., Niebles, J.C.: · 2015
Earlier work this paper cites.
A dataset for movie description
Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: · 2015
Earlier work this paper cites.
Delving into egocentric actions
Li, Y., Ye, Z., Rehg, J.M.: · 2015
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Gorban, A., Idrees, H., Jiang, Y.G., Zamir, A.R., Laptev, I., Shah, M., Sukthankar, R.: · 2015
Earlier work this paper cites.
Learning consistent feature representation for cross-modal multimedia retrieval
Kang, C., Xiang, S., Liao, S., Xu, C., Pan, C.: · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., Szegedy, C.: · 2015
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., Rui, Y.: · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F., Pont-Tuset, J., McWilliams, B., Gool, L.V., Gross, M., Sorkine-Hornung, A.: · 2016
Earlier work this paper cites.
The ImageNet shuffle: Reorganized pre-training for video event detection
Mettes, P., Koelma, D.C., Snoek, C.G.M.: · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M., Favaro, P.: · 2016
Earlier work this paper cites.
Visual semantic role labeling
Gupta, S., Malik, J.: · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: · 2016
Earlier work this paper cites.
University of Michigan North Campus long-term vision and lidar dataset
Carlevaris-Bianco, N., Ushani, A.K., Eustice, R.M.: · 2016
Cited alongside, same era.
Hollywood in Homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: · 2016
Cited alongside, same era.
Temporal Segment Networks: Towards good practices for deep action recognition
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Gool, L.V.: · 2016
Cited alongside, same era.
Online action detection
De Geest, R., Gavves, E., Ghodrati, A., Li, Z., Snoek, C., Tuytelaars, T.: · 2016
Cited alongside, same era.
Human action localization with sparse spatial supervision
Weinzaepfel, P., Martin, X., Schmid, C.: · 2016
Cited alongside, same era.
Connectionist temporal modeling for weakly supervised action labeling
Huang, D.A., Fei-Fei, L., Niebles, J.C.: · 2016
NeuralNetwork-Viterbi: A framework for weakly supervised video learning
Richard, A., Kuehne, H., Iqbal, A., Gall, J.: · 2018
Later among the works it cites.
A flexible model for training action localization with varying levels of supervision
Chéron, G., Alayrac, J., Laptev, I., Schmid, C.: · 2018
Later among the works it cites.
Leveraging uncertainty to rethink loss functions and evaluation measures for egocentric action anticipation
Furnari, A., Battiato, S., Farinella, G.M.: · 2018
Later among the works it cites.
Deep domain adaptation in action space
Jamal, A., Namboodiri, V.P., Deodhare, D., Venkatesh, K.: · 2018
Later among the works it cites.
A unified framework for multimodal domain adaptation
Qi, F., Yang, X., Xu, C.: · 2018
Later among the works it cites.
UMAP: Uniform manifold approximation and projection for dimension reduction
McInnes, L., Healy, J., Melville, J.: · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What’s the point: Semantic segmentation with point supervision
Bearman, A., Russakovsky, O., Ferrari, V., Fei-Fei, L.: · 2016
Cited alongside, same era.
Spot on: Action localization from pointly-supervised proposals
Mettes, P., Van Gemert, J.C., Snoek, C.G.: · 2016
Cited alongside, same era.
Anticipating human activities using object affordances for reactive robotic response
Koppula, H.S., Saxena, A.: · 2016
Cited alongside, same era.
Dual many-to-one-encoder-based transfer learning for cross-dataset human action recognition
Xu, T., Zhu, F., Wong, E.K., Fang, Y.: · 2016
Cited alongside, same era.
Domain-adversarial training of neural networks
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., Lempitsky, V.: · 2016
Cited alongside, same era.
Quo Vadis, action recognition? A new model and the Kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Cited alongside, same era.
Later among the works it cites.
Rethinking ImageNet pre-training
He, K., Girshick, R., Dollár, P.: · 2019
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A.S., Neumann, M., Dosovitskiy, A., Beyer, L., Bachem, O., Tschannen, M., Michalski, M., Bousquet, O., Gelly, S., Houlsby, N.: · 2019
Later among the works it cites.
Does computer vision matter for action?
Zhou, B., Krähenbühl, P., Koltun, V.: · 2019
Later among the works it cites.
nuScenes: A multimodal dataset for autonomous driving
Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: · 2019
Later among the works it cites.
Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving
Yogamani, S., Hughes, C., Horgan, J., Sistu, G., Varley, P., O’Dea, D., Uricár, M., Milz, S., Simon, M., Amende, K., et al.: · 2019
Later among the works it cites.
Grounded video description
Zhou, L., Kalantidis, Y., Chen, X., Corso, J.J., Rohrbach, M.: · 2019
Later among the works it cites.
AVA-ActiveSpeaker: An audio-visual dataset for active speaker detection
Roth, J., Chaudhuri, S., Klejch, O., Marvin, R., Gallagher, A., Kaver, L., Ramaswamy, S., Stopczynski, A., Schmid, C., Xi, Z., et al.: · 2019
Later among the works it cites.
Efficient object annotation via speaking and pointing
Gygli, M., Ferrari, V.: · 2019
Later among the works it cites.
A short note on the Kinetics-700 human action dataset
Carreira, J., Noland, E., Hillier, C., Zisserman, A.: · 2019
Later among the works it cites.
HACS: Human action clips and segments dataset for recognition and temporal localization
Zhao, H., Yan, Z., Torresani, L., Torralba, A.: · 2019
Later among the works it cites.
EPIC-Fusion: Audio-visual temporal binding for egocentric action recognition
Kazakos, E., Nagrani, A., Zisserman, A., Damen, D.: · 2019
Later among the works it cites.
TSM: Temporal shift module for efficient video understanding
Lin, J., Gan, C., Han, S.: · 2019
Later among the works it cites.
SlowFast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., He, K.: · 2019
Later among the works it cites.
Action recognition from single timestamp supervision in untrimmed videos
Moltisanti, D., Fidler, S., Damen, D.: · 2019
Later among the works it cites.
Completeness modeling and context separation for weakly supervised temporal action localization
Liu, D., Jiang, T., Wang, Y.: · 2019
Later among the works it cites.
Weakly-supervised action localization with background modeling
Nguyen, P., Ramanan, D., Fowlkes, C.: · 2019
Later among the works it cites.
3C-Net: Category count and center loss for weakly-supervised action localization
Narayan, S., Cholakkal, H., Khan, F., Shao, L.: · 2019
Later among the works it cites.
D3TW: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation
Chang, C., Huang, D.A., Sui, Y., Fei-Fei, L., Niebles, J.C.: · 2019
Later among the works it cites.
Weakly supervised energy-based learning for action segmentation
Li, J., Lei, P., Todorovic, S.: · 2019
Later among the works it cites.
BMN: Boundary-matching network for temporal action proposal generation
Lin, T., Liu, X., Li, X., Ding, E., Wen, S.: · 2019
Later among the works it cites.
Bayesian prediction of future street scenes using synthetic likelihoods
Bhattacharyya, A., Fritz, M., Schiele, B.: · 2019
Later among the works it cites.
Moment matching for multi-source domain adaptation
Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., Wang, B.: · 2019
Later among the works it cites.
Temporal attentive alignment for large-scale video domain adaptation
Chen, M.H., Kira, Z., AlRegib, G., Yoo, J., Chen, R., Zheng, J.: · 2019
Later among the works it cites.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Wray, M., Larlus, D., Csurka, G., Damen, D.: · 2019
Later among the works it cites.
HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A., Zhukov, D., Alayrac, J.B., Tapaswi, M., Laptev, I., Sivic, J.: · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., Chintala, S.: · 2019
Later among the works it cites.
Understanding human hands in contact at internet scale
Shan, D., Geng, J., Shu, M., Fouhey, D.F.: · 2020
Closest in time.
Detection and Retrieval of Out-of-Distribution Objects in Semantic Segmentation
Oberdiek, P., Rottmann, M., Fink, G.A.: · 2020
Closest in time.
Progressive domain adaptation for object detection
Hsu, H.K., Yao, C.H., Tsai, Y.H., Hung, W.C., Tseng, H.Y., Singh, M., Yang, M.H.: · 2020
Closest in time.
Open Compound Domain Adaptation
Liu, Z., Miao, Z., Zhan, X., Lin, D., Yu, S.X., Icsi, U.C.B.: · 2020
Closest in time.
Moments in Time dataset: One million videos for event understanding
Monfort, M., Vondrick, C., Oliva, A., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S.A., Yan, T., Brown, L., Fan, Q., Gutfreund, D.: · 2020
Closest in time.
Rolling-unrolling LSTMs for action anticipation from first-person video
Furnari, A., Farinella, G.M.: · 2020
Closest in time.
Adversarial cross-domain action recognition with co-attention
Pan, B., Cao, Z., Adeli, E., Niebles, J.C.: · 2020
Closest in time.
Multi-modal domain adaptation for fine-grained action recognition
Munro, J., Damen, D.: · 2020
Closest in time.
Cross-domain first person audio-visual action recognition through relative norm alignment
Planamente, M., Plizzari, C., Alberti, E., Caputo, B.: · 2021
Closest in time.
Yang, L., Huang, Y., Sugano, Y., Sato, Y.: · 2021
Closest in time.
Plizzari, C., Planamente, M., Alberti, E., Caputo, B.: · 2021
Closest in time.