Fetching the paper…
Reading the bibliography…
We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate future actions.
La naissance de l’intelligence chez l’enfant
Jean Piaget · 1935
Earlier work this paper cites.
A non-local algorithm for image denoising
Antoni Buades, Bartomeu Coll, and J-M Morel · 2005
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Hilde Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Activity forecasting
Kris M Kitani, Brian D Ziebart, James Andrew Bagnell, and Martial Hebert · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Sebastian Stein and Stephen J McKenna · 2013
Earlier work this paper cites.
Action-reaction: Forecasting the dynamics of human interaction
De-An Huang and Kris M Kitani · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Learning image representations tied to ego-motion
Dinesh Jayaraman and Kristen Grauman · 2015
Earlier work this paper cites.
Anticipating human activities using object affordances for reactive robotic response
Hema S Koppula and Ashutosh Saxena · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Recurrent neural networks for driver activity anticipation via sensory-fusion architecture
Ashesh Jain, Avi Singh, Hema S Koppula, Shane Soh, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Slow and steady feature analysis: higher order temporal coherence in video
Dinesh Jayaraman and Kristen Grauman · 2016
Earlier work this paper cites.
Learning activity progression in lstms for activity detection and early detection
Shugao Ma, Leonid Sigal, and Stan Sclaroff · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
Red: Reinforced encoder-decoder networks for action anticipation
Jiyang Gao, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Attentional pooling for action recognition
Rohit Girdhar and Deva Ramanan · 2017
Earlier work this paper cites.
ActionVLAD: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fründ, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Low-rank bilinear pooling for fine-grained classification
Shu Kong and Charless Fowlkes · 2017
Earlier work this paper cites.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Earlier work this paper cites.
First-person activity forecasting with online inverse reinforcement learning
Nicholas Rhinehart and Kris M Kitani · 2017
Earlier work this paper cites.
Weakly supervised action learning with rnn based fine-to-coarse modeling
Alexander Richard, Hilde Kuehne, and Juergen Gall · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
When will you do what?-anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Cited alongside, same era.
Modeling temporal structure with lstm for online action detection
Roeland De Geest and Tinne Tuytelaars · 2018
Cited alongside, same era.
Leveraging uncertainty to rethink loss functions and evaluation measures for egocentric action anticipation
Antonino Furnari, Sebastiano Battiato, and Giovanni Maria Farinella · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Cited alongside, same era.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Later among the works it cites.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema · 2020
Later among the works it cites.
Classifying, segmenting, and tracking object instances in video with mask propagation
Gedas Bertasius and Lorenzo Torresani · 2020
Later among the works it cites.
Cobe: Contextualized object embeddings from narrated instructional video
Gedas Bertasius and Lorenzo Torresani · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wen Liu, Weixin Luo, Dongze Lian, and Shenghua Gao · 2018
Cited alongside, same era.
Attention clusters: Purely attention based local feature integration for video classification
Xiang Long, Chuang Gan, Gerard de Melo, Jiajun Wu, Xiao Liu, and Shilei Wen · 2018
Cited alongside, same era.
Predicting future instance segmentation by forecasting convolutional features
Pauline Luc, Camille Couprie, Yann Lecun, and Jakob Verbeek · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Action anticipation by predicting future dynamic images
Cristian Rodriguez, Basura Fernando, and Hongdong Li · 2018
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2020
Later among the works it cites.
Rolling-unrolling lstms for action anticipation from first-person video
Antonino Furnari and Giovanni Maria Farinella · 2020
Later among the works it cites.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Rohit Girdhar and Deva Ramanan · 2020
Later among the works it cites.
Memory-augmented dense predictive coding for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
Directional temporal modeling for action recognition
Xinyu Li, Bing Shuai, and Joseph Tighe · 2020
Later among the works it cites.
Forecasting human object interaction: Joint prediction of motor attention and actions in first person video
Miao Liu, Siyu Tang, Yin Li, and James Rehg · 2020
Later among the works it cites.
EGO-TOPO: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Later among the works it cites.
Temporal aggregate representations for long-range video understanding
Fadime Sener, Dipika Singhania, and Angela Yao · 2020
Later among the works it cites.
Video representation learning with visual tempo consistency
Ceyuan Yang, Yinghao Xu, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.
ViVit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Forecasting action through contact representations from first person video
Eadom Dessalene, Michael Maynord, Chinmaya Devaraj, Cornelia Fermuller, and Yiannis Aloimonos · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Closest in time.
Anticipative Video Transformer @ EPIC-Kitchens Action Anticipation Challenge 2021
Rohit Girdhar and Kristen Grauman · 2021
Closest in time.
Transaction: Icl-sjtu submission to epic-kitchens action anticipation challenge 2021
Xiao Gu, Jianing Qiu, Yao Guo, Benny Lo, and Guang-Zhong Yang · 2021
Closest in time.
VidTr: Video transformer without convolutions
Xinyu Li, Yanyi Zhang, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Closest in time.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
Technical report: Temporal aggregate representations
Fadime Sener, Dibyadip Chatterjee, and Angela Yao · 2021
Closest in time.
Generic event boundary detection: A benchmark for event segmentation
Mike Zheng Shou, Deepti Ghadiyaram, Weiyao Wang, and Matt Feiszli · 2021
Closest in time.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Closest in time.
Learning to anticipate egocentric actions by imagination
Yu Wu, Linchao Zhu, Xiaohan Wang, Yi Yang, and Fei Wu · 2021
Closest in time.
Submission to epic-kitchens action anticipation challenge 2021
Yutaro Yamamuro, Kazuki Hanazawa, Masahiro Shida, Tsuyoshi Kodake, Shinji Takenaka, Yuji Sato, and Takeshi Fujimatsu · 2021
Closest in time.
Multi-modal temporal convolutional network for anticipating actions in egocentric videos
Olga Zatsarynna, Yazan Abu Farha, and Juergen Gall · 2021
Closest in time.
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun · 2021
Closest in time.