Fetching the paper…
Reading the bibliography…
Can we better anticipate an actor's future actions (e.g.
Videograph: Recognizing minutes-long human activities in videos
Noureldien Hussein, Efstratios Gavves, and Arnold W. M. Smeulders · 1905
Earlier work this paper cites.
Task-action grammars: A model of the mental representation of task languages
Stephen J Payne and Thomas RG Green · 1986
Earlier work this paper cites.
The minimalist grammar of action
Katerina Pastra and Yiannis Aloimonos · 2012
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Actionvlad: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Earlier work this paper cites.
When will you do what?-anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall · 2018
Earlier work this paper cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Earlier work this paper cites.
Time-conditioned action anticipation in one shot
Qiuhong Ke, Mario Fritz, and Bernt Schiele · 2019
Earlier work this paper cites.
Zero-shot anticipation for instructional activities
Fadime Sener and Angela Yao · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Earlier work this paper cites.
Long-term anticipation of activities with cycle consistency
Yazan Abu Farha, Qiuhong Ke, Bernt Schiele, and Juergen Gall · 2020
Earlier work this paper cites.
Vectornet: Encoding hd maps and agent dynamics from vectorized representation
Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid · 2020
Earlier work this paper cites.
Habitual behavior is goal-driven
Arie W Kruglanski and Ewa Szumowska · 2020
Earlier work this paper cites.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Cited alongside, same era.
Temporal aggregate representations for long-range video understanding
Fadime Sener, Dipika Singhania, and Angela Yao · 2020
Cited alongside, same era.
Procedure planning in instructional videos via contextual modeling and model-based policy learning
Jing Bi, Jiebo Luo, and Chenliang Xu · 2021
Cited alongside, same era.
Knowledge distillation for action anticipation via label smoothing
Guglielmo Camporese, Pasquale Coscia, Antonino Furnari, Giovanni Maria Farinella, and Lamberto Ballan · 2021
Cited alongside, same era.
Learning temporal dynamics from cycles in narrated video
Dave Epstein, Jiajun Wu, Cordelia Schmid, and Chen Sun · 2021
Cited alongside, same era.
Anticipative video transformer
Rohit Girdhar and Kristen Grauman · 2021
Plate: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg · 2022
Later among the works it cites.
Bevt: Bert pretraining of video transformers
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence · 2022
Later among the works it cites.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Goal-driven long-term trajectory prediction
Hung Tran, Vuong Le, and Truyen Tran · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al · 2022
Cited alongside, same era.
My view is the best view: Procedure learning from egocentric videos
Siddhant Bansal, Chetan Arora, and CV Jawahar · 2022
Cited alongside, same era.
Video+ clip baseline for ego4d long-term action anticipation
Srijan Das and Michael S Ryoo · 2022
Cited alongside, same era.
He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2022
Later among the works it cites.
Hiervl: Learning hierarchical video-language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman · 2023
Closest in time.
Videollm: Modeling video sequence with large language models
Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, et al · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
Text-derived knowledge helps vision: A simple cross-modal distillation for video-based action anticipation
Sayontan Ghosh, Tanvi Aggarwal, Minh Hoai, and Niranjan Balasubramanian · 2023
Closest in time.
Latency matters: Real-time action forecasting transformer
Harshayu Girase, Nakul Agarwal, Chiho Choi, and Karttikeya Mangalam · 2023
Closest in time.
Palm: Predicting actions through language models@ ego4d long-term action anticipation challenge 2023
Daoji Huang, Otmar Hilliges, Luc Van Gool, and Xi Wang · 2023
Closest in time.
Technical report for ego4d long term action anticipation challenge 2023
Tatsuya Ishibashi, Kosuke Ono, Noriyuki Kugo, and Yuji Sato · 2023
Closest in time.
Neuro-symbolic procedural planning with commonsense prompting
Yujie Lu, Weixi Feng, Wanrong Zhu, Wenda Xu, Xin Eric Wang, Miguel Eckstein, and William Yang Wang · 2023
Closest in time.
Linearly mapping from image to text space
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick · 2023
Closest in time.
Large language models as general pattern machines
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng · 2023
Closest in time.
Learning and verification of task structure in instructional videos
Medhini Narasimhan, Licheng Yu, Sean Bell, Ning Zhang, and Trevor Darrell · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Closest in time.