Fetching the paper…
Reading the bibliography…
A key challenge with procedure planning in instructional videos lies in how to handle a large decision space consisting of a multitude of action types that belong to various tasks.
Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
Nair, S.; and Finn, C. 2019 · 1909
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Unsupervised Learning from Narrated Instruction Videos
Alayrac, J.-B.; Bojanowski, P.; Agrawal, N.; Sivic, J.; Laptev, I.; and Lacoste-Julien, S. 2016 · 2016
Earlier work this paper cites.
On Calibration of Modern Neural Networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Earlier work this paper cites.
Learning Plannable Representations with Causal InfoGAN
Kurutach, T.; Tamar, A.; Yang, G.; Russell, S. J.; and Abbeel, P. 2018 · 2018
Earlier work this paper cites.
Universal Planning Networks: Learning Generalizable Representations for Visuomotor Control
Srinivas, A.; Jabri, A.; Abbeel, P.; Levine, S.; and Finn, C. 2018 · 2018
Earlier work this paper cites.
Towards Automatic Learning of Procedures from Web Instructional Videos
Zhou, L.; Xu, C.; and Corso, J. 2018 · 2018
Earlier work this paper cites.
Uncertainty-Aware Anticipation of Activities
Abu Farha, Y.; and Gall, J. 2019 · 2019
Earlier work this paper cites.
Neural Task Graphs: Generalizing to Unseen Tasks From a Single Video Demonstration
Huang, D.-A.; Nair, S.; Xu, D.; Zhu, Y.; Garg, A.; Fei-Fei, L.; Savarese, S.; and Niebles, J. C. 2019 · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Earlier work this paper cites.
COIN: A Large-Scale Dataset for Comprehensive Instructional Video Analysis
Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; and Zhou, J. 2019 · 2019
Earlier work this paper cites.
Cross-task Weakly Supervised Learning from Instructional Videos
Zhukov, D.; Alayrac, J.-B.; Cinbis, R. G.; Fouhey, D.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
Procedure planning in instructional videos
Chang, C.-Y.; Huang, D.-A.; Xu, D.; Adeli, E.; Fei-Fei, L.; and Niebles, J. C. 2020 · 2020
Cited alongside, same era.
Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
Damen, D.; Doughty, H.; Farinella, G. M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M. 2020 · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
Long-Horizon Visual Planning with Goal-Conditioned Hierarchical Predictors
Pertsch, K.; Rybkin, O.; Ebert, F.; Zhou, S.; Jayaraman, D.; Finn, C.; and Levine, S. 2020 · 2020
Cited alongside, same era.
Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy Learning
P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak Supervision
Zhao, H.; Hadji, I.; Dvornik, N.; Derpanis, K. G.; Wildes, R. P.; and Jepson, A. D. 2022 · 2022
Later among the works it cites.
HierVL: Learning Hierarchical Video-Language Embeddings
Ashutosh, K.; Girdhar, R.; Torresani, L.; and Grauman, K. 2023 · 2023
Closest in time.
Masked Diffusion Transformer is a Strong Image Synthesizer
Gao, S.; Zhou, P.; Cheng, M.-M.; and Yan, S. 2023 · 2023
Closest in time.
Chain of Thought Prompt Tuning in Vision Language Models
Ge, J.; Luo, H.; Qian, S.; Gan, Y.; Fu, J.; and Zhang, S. 2023 · 2023
Closest in time.
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bi, J.; Luo, J.; and Xu, C. 2021 · 2021
Cited alongside, same era.
PlaTe: Visually-Grounded Planning With Transformers in Procedural Tasks
Sun, J.; Huang, D.-A.; Lu, B.; Liu, Y.; Zhou, B.; and Garg, A. 2021 · 2021
Cited alongside, same era.
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Xu, H.; Ghosh, G.; Huang, P.-Y. B.; Okhonko, D.; Aghajanyan, A.; and Feichtenhofer, F. M. L. Z. C. 2021 · 2021
Cited alongside, same era.
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Grauman, K.; Westbury, A.; Byrne, E.; and et al. 2022 · 2022
Cited alongside, same era.
Pure Transformers are Powerful Graph Learners
Kim, J.; Nguyen, D.; Min, S.; Cho, S.; Lee, M.; Lee, H.; and Hong, S. 2022 · 2022
Cited alongside, same era.
Learning To Recognize Procedural Activities with Distant Supervision
Lin, X.; Petroni, F.; Bertasius, G.; Rohrbach, M.; Chang, S.-F.; and Torresani, L. 2022 · 2022
Cited alongside, same era.
Learning Parameterized Task Structure for Generalization to Unseen Entities
Liu, A.; Sohn, S.; Qazwini, M.; and Lee, H. 2022 · 2022
Cited alongside, same era.
Action Dynamics Task Graphs for Learning Plannable Representations of Procedural Tasks
Mao, W.; Desai, R.; Iuzzolino, M. L.; and Kamra, N. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Pretrained Language Models as Visual Planners for Human Assistance
Patel, D.; Eghbalzadeh, H.; Kamra, N.; Iuzzolino, M. L.; Jain, U.; and Desai, R. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Closest in time.
Learning Video Representations from Large Language Models
Zhao, Y.; Misra, I.; Krähenbühl, P.; and Girdhar, R. 2023 · 2023
Closest in time.
Fast Training of Diffusion Models with Masked Transformers
Zheng, H.; Nie, W.; Vahdat, A.; and Anandkumar, A. 2023 · 2023
Closest in time.
Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations
Zhong, Y.; Yu, L.; Bai, Y.; Li, S.; Yan, X.; and Li, Y. 2023 · 2023
Closest in time.
Procedure-Aware Pretraining for Instructional Video Understanding
Zhou, H.; Martín-Martín, R.; Kapadia, M.; Savarese, S.; and Niebles, J. C. 2023 · 2023
Closest in time.