Fetching the paper…
Reading the bibliography…
The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification.
Origins of knowledge
Spelke, E.S., Breinlinger, K., Macomber, J., Jacobson, K.: · 1992
Earlier work this paper cites.
The scientist in the crib: What early learning tells us about the mind
Gopnik, A., Meltzoff, A.N., Kuhl, P.K.: · 2000
Earlier work this paper cites.
Visual recognition: As soon as you know it is there, you know what it is
Grill-Spector, K., Kanwisher, N.: · 2005
Earlier work this paper cites.
Objects in action: An approach for combining action understanding and object perception
Gupta, A., Davis, L.S.: · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
How to grow a mind: Statistics, structure, and abstraction
Tenenbaum, J.B., Kemp, C., Griffiths, T.L., Goodman, N.D.: · 2011
Earlier work this paper cites.
Wsabie: Scaling up to large vocabulary image annotation
Weston, J., Bengio, S., Usunier, N.: · 2011
Earlier work this paper cites.
Mid-level features improve recognition of interactive activities
Saenko, K., Packer, B., Chen, C., Bandla, S., Lee, Y., Jia, Y., Niebles, J., Koller, D., Fei-Fei, L., Grauman, K., et al.: · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., Mikolov, T.: · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., Sun, J.: · 2015
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Everingham, M., Eslami, S., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: · 2015
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Battaglia, P., Pascanu, R., Lai, M., Jimenez Rezende, D., et al.: · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Feichtenhofer, C., Pinz, A., Zisserman, A.: · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., Sorkine-Hornung, A.: · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J., Karpathy, A., Fei-Fei, L.: · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z.: · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., Weinberger, K.Q.: · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: · 2017
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., Girshick, R.: · 2017
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Aytar, Y., Vondrick, C., Torralba, A.: · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al.: · 2017
Earlier work this paper cites.
Actor-centric relation network
Sun, C., Shrivastava, A., Vondrick, C., Murphy, K., Sukthankar, R., Schmid, C.: · 2018
Earlier work this paper cites.
Videos as space-time region graphs
Wang, X., Gupta, A.: · 2018
Earlier work this paper cites.
Investigating human priors for playing video games
Dubey, R., Agrawal, P., Pathak, D., Griffiths, T.L., Efros, A.A.: · 2018
Earlier work this paper cites.
The developing infant creates a curriculum for statistical learning
Smith, L.B., Jayaraman, S., Clerkin, E., Yu, C.: · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K.: · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.: · 2018
Earlier work this paper cites.
Compositional learning for human object interaction
Kato, K., Li, Y., Gupta, A.: · 2018
Earlier work this paper cites.
Detecting and recognizing human-object interactions
Gkioxari, G., Girshick, R., Dollár, P., He, K.: · 2018
Cited alongside, same era.
Object level visual reasoning in videos
Baradel, F., Neverova, N., Wolf, C., Mille, J., Mori, G.: · 2018
Cited alongside, same era.
Attend and interact: Higher-order object interactions for video understanding
Ma, C.Y., Kadav, A., Melvin, I., Kira, Z., AlRegib, G., Graf, H.P.: · 2018
Cited alongside, same era.
Objects that sound
Arandjelovic, R., Zisserman, A.: · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Owens, A., Efros, A.A.: · 2018
Cited alongside, same era.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., He, K.: · 2018
Cited alongside, same era.
Unsupervised state representation learning in atari
Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M.A., Hjelm, R.D.: · 2019
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I., Hutter, F.: · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
Wu, C.Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., Girshick, R.: · 2019
Later among the works it cites.
Randaugment: Practical data augmentation with no separate search
Cubuk, E.D., Zoph, B., Shlens, J., Le, Q.V.: · 2019
Later among the works it cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., Kipf, T.: · 2020
Later among the works it cites.
Temporal pyramid network for action recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learnable pins: Cross-modal embeddings for person identity
Nagrani, A., Albanie, S., Zisserman, A.: · 2018
Cited alongside, same era.
Youtube-vos: A large-scale video object segmentation benchmark
Xu, N., Yang, L., Fan, Y., Yue, D., Liang, Y., Yang, J., Huang, T.: · 2018
Cited alongside, same era.
Referring relationships
Krishna, R., Chami, I., Bernstein, M., Fei-Fei, L.: · 2018
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Battaglia, P., Hamrick, J.B.C., Bapst, V., Sanchez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G.E., Vaswani, A., Allen, K., Nash, C., Langston, V.J., Dyer, C., Heess, N., Wierstra, D., Kohli, P., Botvinick, M., Vinyals, O., Li, Y., Pascanu, R.: · 2018
Cited alongside, same era.
Image generation from scene graphs
Johnson, J., Gupta, A., Fei-Fei, L.: · 2018
Cited alongside, same era.
Mapping images to scene graphs with permutation-invariant structured prediction
Herzig, R., Raboh, M., Chechik, G., Berant, J., Globerson, A.: · 2018
Cited alongside, same era.
Yang, C., Xu, Y., Shi, J., Dai, B., Zhou, B.: · 2020
Later among the works it cites.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Materzynska, J., Xiao, T., Herzig, R., Xu, H., Wang, X., Darrell, T.: · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Ji, J., Krishna, R., Fei-Fei, L., Niebles, J.C.: · 2020
Later among the works it cites.
Drg: Dual relation graph for human-object interaction detection
Gao, C., Xu, J., Zou, Y., Huang, J.B.: · 2020
Later among the works it cites.
Interactive fusion of multi-level features for compositional activity recognition
Yan, R., Xie, L., Shu, X., Tang, J.: · 2020
Later among the works it cites.
Self-supervised multimodal versatile networks
Alayrac, J.B., Recasens, A., Schneider, R., Arandjelović, R., Ramapuram, J., De Fauw, J., Smaira, L., Dieleman, S., Zisserman, A.: · 2020
Later among the works it cites.
Tao: A large-scale benchmark for tracking any object
Dave, A., Khurana, T., Tokmakov, P., Schmid, C., Ramanan, D.: · 2020
Later among the works it cites.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Girdhar, R., Ramanan, D.: · 2020
Later among the works it cites.
Spatio-temporal action detection with multi-object interaction
Xu, H., Yang, L., Sclaroff, S., Saenko, K., Darrell, T.: · 2020
Later among the works it cites.
Future video synthesis with object motion prediction
Wu, Y., Gao, R., Park, J., Chen, Q.: · 2020
Later among the works it cites.
Unsupervised object-centric video generation and decomposition in 3d
Henderson, P., Lampert, C.H.: · 2020
Later among the works it cites.
Learning canonical representations for scene graph to image generation
Herzig, R., Bar, A., Xu, H., Chechik, G., Darrell, T., Globerson, A.: · 2020
Later among the works it cites.
Object-centric forward modeling for model predictive control
Ye, Y., Gandhi, D., Gupta, A., Tulsiani, S.: · 2020
Later among the works it cites.
Understanding human hands in contact at internet scale
Shan, D., Geng, J., Shu, M., Fouhey, D.F.: · 2020
Later among the works it cites.
Object-region video transformers
Herzig, R., Ben-Avraham, E., Mangalam, K., Bar, A., Chechik, G., Rohrbach, A., Darrell, T., Globerson, A.: · 2021
Later among the works it cites.
Revisiting spatio-temporal layouts for compositional action recognition
Radevski, G., Moens, M.F., Tuytelaars, T.: · 2021
Later among the works it cites.
Unified graph structured models for video understanding
Arnab, A., Sun, C., Schmid, C.: · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Kolesnikov, A., Dosovitskiy, A., Weissenborn, D., Heigold, G., Uszkoreit, J., Beyer, L., Minderer, M., Dehghani, M., Houlsby, N., Gelly, S., Unterthiner, T., Zhai, X.: · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Patrick, M., Campbell, D., Asano, Y., Misra, I., Metze, F., Feichtenhofer, C., Vedaldi, A., Henriques, J.F.: · 2021
Later among the works it cites.
The ninth visual object tracking vot2021 challenge results
Kristan, M., Matas, J., Leonardis, A., Felsberg, M., Pflugfelder, R., Kämäräinen, J.K., Chang, H.J., Danelljan, M., Cehovin, L., Lukežič, A., et al.: · 2021
Later among the works it cites.
Motchallenge: A benchmark for single-camera multiple target tracking
Dendorfer, P., Osep, A., Milan, A., Schindler, K., Cremers, D., Reid, I., Roth, S., Leal-Taixé, L.: · 2021
Later among the works it cites.
Unidentified video objects: A benchmark for dense, open-world segmentation
Wang, W., Feiszli, M., Wang, H., Tran, D.: · 2021
Later among the works it cites.
Self-supervised video object segmentation by motion grouping
Yang, C., Lamdouar, H., Lu, E., Zisserman, A., Xie, W.: · 2021
Later among the works it cites.
Learning object-compositional neural radiance field for editable scene rendering
Yang, B., Zhang, Y., Xu, Y., Li, Y., Zhou, H., Bao, H., Zhang, G., Cui, Z.: · 2021
Later among the works it cites.
Bytetrack: Multi-object tracking by associating every detection box
Zhang, Y., Sun, P., Jiang, Y., Yu, D., Yuan, Z., Luo, P., Liu, W., Wang, X.: · 2021
Later among the works it cites.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., Schmid, C.: · 2021
Later among the works it cites.