Fetching the paper…
Reading the bibliography…
In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
Lotter, W.; Kreiman, G.; and Cox, D. 2017 · 2017
Earlier work this paper cites.
Deep future gaze: Gaze anticipation on egocentric videos using adversarial networks
Zhang, M.; Teck Ma, K.; Hwee Lim, J.; Zhao, Q.; and Feng, J. 2017 · 2017
Earlier work this paper cites.
Ha, D.; and Schmidhuber, J. 2018 · 2018
Earlier work this paper cites.
Stochastic Adversarial Video Prediction
Lee, A.; Strobel, M.; and Finn, C. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018 · 2018
Earlier work this paper cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Kazakos, E.; Nagrani, A.; Zisserman, A.; and Damen, D. 2019 · 2019
Earlier work this paper cites.
Lsta: Long short-term attention for egocentric action recognition
Sudhakaran, S.; Escalera, S.; and Lanz, O. 2019 · 2019
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2020 · 2020
Earlier work this paper cites.
Mutual context network for jointly estimating egocentric gaze and action
Huang, Y.; Cai, M.; Li, Z.; Lu, F.; and Sato, Y. 2020 · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
James, S.; Ma, Z.; Arrojo, D. R.; and Davison, A. J. 2020 · 2020
Earlier work this paper cites.
Selfpose: 3d egocentric pose estimation from a headset mounted camera
Tome, D.; Alldieck, T.; Peluse, P.; Pons-Moll, G.; Agapito, L.; Badino, H.; and De la Torre, F. 2020 · 2020
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Earlier work this paper cites.
Human hands as probes for interactive object understanding
Goyal, M.; Modi, S.; Goyal, R.; and Gupta, S. 2022 · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. 2022 · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2022 · 2022
Earlier work this paper cites.
Generative adversarial network for future hand segmentation from egocentric video
Jia, W.; Liu, M.; and Rehg, J. M. 2022 · 2022
Earlier work this paper cites.
Multiviz: Towards visualizing and understanding multimodal models
Liang, P. P.; Lyu, Y.; Chhablani, G.; Jain, N.; Deng, Z.; Wang, X.; Morency, L.-P.; and Salakhutdinov, R. 2022 · 2022
Cited alongside, same era.
Egocentric video-language pretraining
Lin, K. Q.; Wang, J.; Soldan, M.; Wray, M.; Yan, R.; XU, E. Z.; Gao, D.; Tu, R.-C.; Zhao, W.; Kong, W.; et al. 2022 · 2022
Cited alongside, same era.
Joint hand motion and interaction hotspots prediction from egocentric videos
Liu, S.; Tripathi, S.; Majumdar, S.; and Wang, X. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. 2023 · 2023
Taca: Upgrading your visual foundation model with task-agnostic compatible adapter
Zhang, B.; Ge, Y.; Xu, X.; Shan, Y.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Openvla: An open-source vision-language-action model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. 2024 · 2024
Later among the works it cites.
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Li, F.; Zhang, R.; Zhang, H.; Zhang, Y.; Li, B.; Li, W.; Ma, Z.; and Li, C. 2024 · 2024
Later among the works it cites.
World model on million-length video and language with ringattention
Liu, H.; Yan, W.; Zaharia, M.; and Abbeel, P. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing
Chen, W.-G.; Spiridonova, I.; Yang, J.; Gao, J.; and Li, C. 2023 · 2023
Cited alongside, same era.
Foundation models in robotics: Applications, challenges, and the future
Firoozi, R.; Tucker, J.; Tian, S.; Majumdar, A.; Sun, J.; Liu, W.; Zhu, Y.; Song, S.; Kapoor, A.; Hausman, K.; et al. 2023 · 2023
Cited alongside, same era.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
Ghosal, D.; Majumder, N.; Mehrish, A.; and Poria, S. 2023 · 2023
Cited alongside, same era.
Gaia-1: A generative world model for autonomous driving
Hu, A.; Russell, L.; Yeo, H.; Murez, Z.; Fedoseev, G.; Kendall, A.; Shotton, J.; and Corrado, G. 2023 · 2023
Cited alongside, same era.
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
Lai, B.; Ryan, F.; Jia, W.; Liu, M.; and Rehg, J. M. 2023 · 2023
Cited alongside, same era.
Ego-Body Pose Estimation via Ego-Head Pose Estimation
Li, J.; Liu, K.; and Wu, J. 2023 · 2023
Cited alongside, same era.
StillFast: An End-to-End Approach for Short-Term Object Interaction Anticipation
Ragusa, F.; Farinella, G. M.; and Furnari, A. 2023 · 2023
Cited alongside, same era.
Nymeria: A massive collection of multimodal egocentric daily motion in the wild
Ma, L.; Ye, Y.; Hong, F.; Guzov, V.; Jiang, Y.; Postyeni, R.; Pesqueira, L.; Gamino, A.; Baiyya, V.; Kim, H. J.; et al. 2024 · 2024
Later among the works it cites.
Llarva: Vision-action instruction tuning enhances robot learning
Niu, D.; Sharma, Y.; Biamby, G.; Quenum, J.; Bai, Y.; Shi, B.; Darrell, T.; and Herzig, R. 2024 · 2024
Later among the works it cites.
Genhowto: Learning to generate actions and state transformations from instructional videos
Souček, T.; Damen, D.; Wray, M.; Laptev, I.; and Sivic, J. 2024 · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Team, O. M.; Ghosh, D.; Walke, H.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Kreiman, T.; Xu, C.; et al. 2024 · 2024
Later among the works it cites.
Drivevlm: The convergence of autonomous driving and large vision-language models
Tian, X.; Gu, J.; Li, B.; Liu, Y.; Wang, Y.; Zhao, Z.; Zhan, K.; Jia, P.; Lang, X.; and Zhao, H. 2024 · 2024
Later among the works it cites.
Draganything: Motion control for anything using entity representation
Wu, W.; Li, Z.; Gu, Y.; Zhao, R.; He, Y.; Zhang, D. J.; Shou, M. Z.; Li, Y.; Gao, T.; and Zhang, D. 2024 · 2024
Later among the works it cites.
Whole-Body Conditioned Egocentric Video Prediction
Bai, Y.; Tran, D.; Bar, A.; LeCun, Y.; Darrell, T.; and Malik, J. 2025 · 2025
Closest in time.
Adaworld: Learning adaptable world models with latent actions
Gao, S.; Zhou, S.; Du, Y.; Zhang, J.; and Gan, C. 2025 · 2025
Closest in time.
Streamingt2v: Consistent, dynamic, and extendable long video generation from text
Henschel, R.; Khachatryan, L.; Poghosyan, H.; Hayrapetyan, D.; Tadevosyan, V.; Wang, Z.; Navasardyan, S.; and Shi, H. 2025 · 2025
Closest in time.
Lego: Learning egocentric action frame generation via visual instruction tuning
Lai, B.; Dai, X.; Chen, L.; Pang, G.; Rehg, J. M.; and Liu, M. 2025 · 2025
Closest in time.
Manigaussian: Dynamic gaussian splatting for multi-task robotic manipulation
Lu, G.; Zhang, S.; Wang, Z.; Liu, C.; Lu, J.; and Tang, Y. 2025 · 2025
Closest in time.
Tora: Trajectory-oriented diffusion transformer for video generation
Zhang, Z.; Liao, J.; Li, M.; Dai, Z.; Qiu, B.; Zhu, S.; Qin, L.; and Wang, W. 2025 · 2025
Closest in time.
Occworld: Learning a 3d occupancy world model for autonomous driving
Zheng, W.; Chen, W.; Huang, Y.; Zhang, B.; Duan, Y.; and Lu, J. 2025 · 2025
Closest in time.