Fetching the paper…
Reading the bibliography…
Video procedure planning, i.e., planning a sequence of action steps given the video frames of start and goal states, is an essential ability for embodied AI.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Mysore, S.; Jensen, Z.; Kim, E.; Huang, K.; Chang, H.-S.; Strubell, E.; Flanigan, J.; McCallum, A.; and Olivetti, E. 2019 · 1905
Earlier work this paper cites.
A language-first approach for procedure planning
Liu, J.; Li, S.; Wang, Z.; Li, M.; and Ji, H. 2023b · 1954
Earlier work this paper cites.
Jansen, P. A. 2020 · 2009
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Tellex, S.; Kollar, T.; Dickerson, S.; Walter, M.; Banerjee, A.; Teller, S.; and Roy, N. 2011 · 2011
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Alayrac, J.-B.; Bojanowski, P.; Agrawal, N.; Sivic, J.; Laptev, I.; and Lacoste-Julien, S. 2016 · 2016
Earlier work this paper cites.
Who let the dogs out? modeling dog behavior from visual data
Ehsani, K.; Bagherinezhad, H.; Redmon, J.; Mottaghi, R.; and Farhadi, A. 2018 · 2018
Earlier work this paper cites.
Mishra, B. D.; Huang, L.; Tandon, N.; Yih, W.-t.; and Clark, P. 2018 · 2018
Earlier work this paper cites.
Universal planning networks: Learning generalizable representations for visuomotor control
Srinivas, A.; Jabri, A.; Abbeel, P.; Levine, S.; and Finn, C. 2018 · 2018
Earlier work this paper cites.
Uncertainty-aware anticipation of activities
Abu Farha, Y.; and Gall, J. 2019 · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; and Zhou, J. 2019 · 2019
Earlier work this paper cites.
Cross-task weakly supervised learning from instructional videos
Zhukov, D.; Alayrac, J.-B.; Cinbis, R. G.; Fouhey, D.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
Procedure planning in instructional videos
Chang, C.-Y.; Huang, D.-A.; Xu, D.; Adeli, E.; Fei-Fei, L.; and Niebles, J. C. 2020 · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
Miech, A.; Alayrac, J.-B.; Smaira, L.; Laptev, I.; Sivic, J.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
Procedure planning in instructional videos via contextual modeling and model-based policy learning
Bi, J.; Luo, J.; and Xu, C. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2023 · 2023
Later among the works it cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H.; Wu, J.; Washington, C.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2023 · 2023
Later among the works it cites.
Translating natural language to planning goals with large-language models
Xie, Y.; Yu, C.; Zhu, T.; Bai, J.; Gong, Z.; and Soh, H. 2023 · 2023
Later among the works it cites.
Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization
Zheng, M.; Gong, S.; Jin, H.; Peng, Y.; and Liu, Y. 2023 · 2023
Later among the works it cites.
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection
Lei, T.; Yin, S.; Peng, Y.; and Liu, Y. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Fu, C.; Gopalakrishnan, K.; Hausman, K.; et al. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Cited alongside, same era.
PlaTe: Visually-grounded planning with transformers in procedural tasks
Sun, J.; Huang, D.-A.; Lu, B.; Liu, Y.-H.; Zhou, B.; and Garg, A. 2022 · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2022 · 2022
Cited alongside, same era.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
Zhao, H.; Hadji, I.; Dvornik, N.; Derpanis, K. G.; Wildes, R. P.; and Jepson, A. D. 2022 · 2022
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023a
Cited in the paper.
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
Nagasinghe, K. R. Y.; Zhou, H.; Gunawardhana, M.; Min, M. R.; Harari, D.; and Khan, M. H. 2024 · 2024
Closest in time.
SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
Niu, Y.; Guo, W.; Chen, L.; Lin, X.; and Chang, S.-F. 2024 · 2024
Closest in time.
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
Ting Lei, S. Y.; and Liu, Y. 2024 · 2024
Closest in time.
Active Object Detection with Knowledge Aggregation and Distillation from Large Models
Yang, D.; and Liu, Y. 2024 · 2024
Closest in time.
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
Ye, Q.; Yu, Z.; Shao, R.; Xie, X.; Torr, P.; and Cao, X. 2024 · 2024
Closest in time.
Training Free Video Temporal Grounding using Large-scale Pre-trained Models
Zheng, M.; Cai, X.; Chen, Q.; Peng, Y.; and Liu, Y. 2024 · 2024
Closest in time.