Fetching the paper…
Reading the bibliography…
We present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Mpnet: Masked and permuted pre-training for language understanding, 2020
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, Mar. 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2021
Earlier work this paper cites.
Video + clip baseline for ego4d long-term action anticipation, 2022
Srijan Das and Michael S. Ryoo · 2022
Cited alongside, same era.
Egocentric video-language pretraining
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Cited alongside, same era.
vit-gpt2-image-captioning (revision 0e334c7), 2022
NLP Connect · 2022
Cited alongside, same era.
Complementary explanations for effective in-context learning
Xi Ye, Srini Iyer, Asli Celikyilmaz, Ves Stoyanov, Greg Durrett, and Ramakanth Pasunuru · 2022
Later among the works it cites.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Intention-conditioned long-term human egocentric action anticipation
Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…