Fetching the paper…
Reading the bibliography…
Given video demonstrations and paired narrations of an at-home procedural task such as changing a tire, we present an approach to extract the underlying task structure -- relevant actions and their temporal dependencies -- via action-centric task graphs.
A benchmark for structured procedural knowledge extraction from cooking videos
Xu, F. F.; Ji, L.; Shi, B.; Du, J.; Neubig, G.; Bisk, Y.; and Duan, N. 2020 · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Autonomously constructing hierarchical task networks for planning and human-robot collaboration
Hayes, B.; and Scassellati, B. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Jointly learning grounded task structures from language instruction and visual demonstration
Liu, C.; Yang, S.; Saba-Sadiya, S.; Shukla, N.; He, Y.; Zhu, S.-C.; and Chai, J. 2016 · 2016
Earlier work this paper cites.
Actions˜ transformations
Wang, X.; Farhadi, A.; and Gupta, A. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Cited alongside, same era.
Learning plannable representations with causal InfoGAN
Kurutach, T.; Tamar, A.; Yang, G.; Russell, S. J.; and Abbeel, P. 2018 · 2018
Cited alongside, same era.
Universal planning networks: Learning generalizable representations for visuomotor control
Srinivas, A.; Jabri, A.; Abbeel, P.; Levine, S.; and Finn, C. 2018 · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; and Davidson, J. 2019 · 2019
Cited alongside, same era.
Neural task graphs: Generalizing to unseen tasks from a single video demonstration
Huang, D.-A.; Nair, S.; Xu, D.; Zhu, Y.; Garg, A.; Fei-Fei, L.; Savarese, S.; and Niebles, J. C. 2019 · 2019
Cited alongside, same era.
Procedure planning in instructional videos
Chang, C.-Y.; Huang, D.-A.; Xu, D.; Adeli, E.; Fei-Fei, L.; and Niebles, J. C. 2020 · 2020
Later among the works it cites.
Dynamics Learning with Cascaded Variational Inference for Multi-Step Manipulation
Fang, K.; Zhu, Y.; Garg, A.; Savarese, S.; and Fei-Fei, L. 2020 · 2020
Later among the works it cites.
Multi-modal cooking workflow construction for food recipes
Pan, L.-M.; Chen, J.; Wu, J.; Liu, S.; Ngo, C.-W.; Kan, M.-Y.; Jiang, Y.; and Chua, T.-S. 2020 · 2020
Later among the works it cites.
Procedure planning in instructional videos via contextual modeling and model-based policy learning
Bi, J.; Luo, J.; and Xu, C. 2021 · 2021
Later among the works it cites.
PlaTe: Visually-grounded planning with transformers in procedural tasks
Sun, J.; Huang, D.-A.; Lu, B.; Liu, Y.-H.; Zhou, B.; and Garg, A. 2022 · 2022
Later among the works it cites.
P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak Supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhukov, D.; Alayrac, J.-B.; Cinbis, R. G.; Fouhey, D.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
Zhao, H.; Hadji, I.; Dvornik, N.; Derpanis, K. G.; Wildes, R. P.; and Jepson, A. D. 2022 · 2022
Later among the works it cites.