Fetching the paper…
Reading the bibliography…
Ego-to-exo video generation refers to generating the corresponding exocentric video according to the egocentric video, providing valuable applications in AR/VR and embodied AI.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Schönberger, J. L. and Frahm, J.-M · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Identifying first-person camera wearers in third-person videos
Fan, C., Lee, J., Xu, M., Kumar Singh, K., Jae Lee, Y., Crandall, D. J., and Ryoo, M. S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Summarizing first-person videos from third persons’ points of view
Ho, H.-I., Chiu, W.-C., and Wang, Y.-C. F · 2018
Earlier work this paper cites.
Actor and observer: Joint modeling of first and third-person videos
Sigurdsson, G. A., Gupta, A., Schmid, C., Farhadi, A., and Alahari, K · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Lemma: A multiview dataset for learning multi-agent multi-view activities
Jia, B., Chen, Y., Huang, S., Zhu, Y., and Zhu, S.-C · 2020
Cited alongside, same era.
Imaginator: Conditional spatio-temporal gan for video generation
Wang, Y., Bilinski, P., Bremond, F., and Dantcheva, A · 2020
Cited alongside, same era.
Ego-exo: Transferring visual representations from third-person to first-person videos
Li, Y., Nagarajan, T., Xiong, B., and Grauman, K · 2021
Cited alongside, same era.
Cross-view exocentric to egocentric video synthesis
Liu, G., Tang, H., Latapie, H. M., Corso, J. J., and Yan, Y · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Home action genome: Cooperative compositional action understanding
3d human pose perception from egocentric stereo videos
Akada, H., Wang, J., Golyanik, V., and Theobalt, C · 2023
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Grauman, K., Westbury, A., Torresani, L., Kitani, K., Malik, J., Afouras, T., Ashutosh, K., Baiyya, V., Bansal, S., Boote, B., et al · 2023
Later among the works it cites.
Hu, Z. and Xu, D · 2023
Later among the works it cites.
Ego-body pose estimation via ego-head pose estimation
Li, J., Liu, K., and Wu, J · 2023
Later among the works it cites.
Trailblazer: Trajectory control for diffusion-based video generation, 2023
Ma, W.-D. K., Lewis, J. P., and Kleijn, W. B · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rai, N., Chen, H., Ji, J., Desai, R., Kozuka, K., Ishizaka, S., Adeli, E., and Niebles, J. C · 2021
Cited alongside, same era.
Motion representations for articulated animation
Siarohin, A., Woodford, O. J., Ren, J., Chai, M., and Tulyakov, S · 2021
Cited alongside, same era.
Seeing the unseen: Predicting the first-person camera wearer’s location and pose in third-person scenes
Wen, Y., Singh, K. K., Anderson, M., Jan, W.-P., and Lee, Y. J · 2021
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Cited alongside, same era.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2022
Cited alongside, same era.
Later among the works it cites.
Conditional image-to-video generation with latent flow diffusion models
Ni, H., Shi, C., Li, K., Huang, S. X., and Min, M. R · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Later among the works it cites.
Cross-view action recognition understanding from exocentric to egocentric perspective
Truong, T.-D. and Luu, K · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Later among the works it cites.
Xue, Z. and Grauman, K · 2023
Later among the works it cites.
Diffusion models: A comprehensive survey of methods and applications
Yang, L., Zhang, Z., Song, Y., Hong, S., Xu, R., Zhao, Y., Zhang, W., Cui, B., and Yang, M.-H · 2023
Later among the works it cites.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Yin, S., Wu, C., Liang, J., Shi, J., Li, H., Ming, G., and Duan, N · 2023
Later among the works it cites.