Fetching the paper…
Reading the bibliography…
We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective.
Image quality metrics: Psnr vs. ssim
Alain Hore and Djemel Ziou · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
The ycb object and model set: Towards common benchmarks for manipulation research
Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M Dollar · 2015
Earlier work this paper cites.
Action recognition in the presence of one egocentric and multiple static cameras
Bilge Soran, Ali Farhadi, and Linda Shapiro · 2015
Earlier work this paper cites.
Ego2top: Matching viewers in egocentric and top-view videos
Shervin Ardeshir and Ali Borji · 2016
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Identifying first-person camera wearers in third-person videos
Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar Singh, Yong Jae Lee, David J Crandall, and Michael S Ryoo · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Recycle-gan: Unsupervised video retargeting
Aayush Bansal, Shugao Ma, Deva Ramanan, and Yaser Sheikh · 2018
Earlier work this paper cites.
From third person to first person: Dataset and baselines for synthesis and retrieval
Mohamed Elfeki, Krishna Regmi, Shervin Ardeshir, and Ali Borji · 2018
Earlier work this paper cites.
Summarizing first-person videos from third persons’ points of view
Hsuan-I Ho, Wei-Chen Chiu, and Yu-Chiang Frank Wang · 2018
Earlier work this paper cites.
Cross-view image synthesis using conditional gans
Krishna Regmi and Ali Borji · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, and Sergey Levine · 2018
Earlier work this paper cites.
Joint person segmentation and identification in synchronized first-and third-person videos
Mingze Xu, Chenyou Fan, Yuchen Wang, Michael S Ryoo, and David J Crandall · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Cross-view image synthesis using geometry-guided conditional gans
Krishna Regmi and Ali Borji · 2019
Earlier work this paper cites.
Scene representation networks: Continuous 3d-structure-aware neural scene representations
Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein · 2019
Earlier work this paper cites.
Multi-channel attention selection gan with cascaded semantic guidance for cross-view image translation
Hao Tang, Dan Xu, Nicu Sebe, Yanzhi Wang, Jason J Corso, and Yan Yan · 2019
Cited alongside, same era.
What i see is what you see: Joint attention learning for first and third person video co-analysis
Huangyue Yu, Minjie Cai, Yunfei Liu, and Feng Lu · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Seeing the unseen: Predicting the first-person camera wearer’s location and pose in third-person scenes
Yangming Wen, Krishna Kumar Singh, Markham Anderson, Wei-Pang Jan, and Yong Jae Lee · 2021
Later among the works it cites.
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa · 2021
Later among the works it cites.
Human-to-robot imitation in the wild
Shikhar Bahl, Abhinav Gupta, and Deepak Pathak · 2022
Later among the works it cites.
Cascaded diffusion models for high fidelity image generation
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans · 2022
Later among the works it cites.
Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation
Rishabh Jangir, Nicklas Hansen, Sambaran Ghosal, Mohit Jain, and Xiaolong Wang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exocentric to egocentric image generation via parallel generative adversarial network
Gaowen Liu, Hao Tang, Hugo Latapie, and Yan Yan · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Cited alongside, same era.
pytorch-fid: FID Score for PyTorch
Maximilian Seitzer · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
Dandan Shan, Jiaqi Geng, Michelle Shu, and David F. Fouhey · 2020
Cited alongside, same era.
Synsin: End-to-end view synthesis from a single image
Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson · 2020
Cited alongside, same era.
First-and third-person video co-analysis by learning spatial-temporal joint attention
Huangyue Yu, Minjie Cai, Yunfei Liu, and Feng Lu · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Aria pilot dataset
Zhaoyang Lv, Edward Miller, Jeff Meissner, Luis Pesqueira, Chris Sweeney, Jing Dong, Lingni Ma, Pratik Patel, Pierre Moulon, Kiran Somasundaram, Omkar Parkhi, Yuyang Zou, Nikhil Raina, Steve Saarinen, Yusuf M Mansour, Po-Kang Huang, Zijian Wang, Anton Troynikov, Raul Mur Artal, Daniel DeTone, Daniel Barnes, Elizabeth Argall, Andrey Lobanovskiy, David Jaeyun Kim, Philippe Bouttefroy, Julian Straub, Jakob Julian Engel, Prince Gupta, Mingfei Yan, Renzo De Nardi, and Richard Newcombe · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Later among the works it cites.
Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan · 2022
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2022
Later among the works it cites.
Look outside the room: Synthesizing a consistent long-term 3d scene video from a single image
Xuanchi Ren and Xiaolong Wang · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao · 2022
Later among the works it cites.
Estimating egocentric 3d human pose in the wild with external weak supervision
Jian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Diogo Luvizon, and Christian Theobalt · 2022
Later among the works it cites.
Novel view synthesis with diffusion models
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi · 2022
Later among the works it cites.
Zero-shot robot manipulation from passive human videos
Homanga Bharadhwaj, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar · 2023
Later among the works it cites.
Generative novel view synthesis with 3d-aware diffusion models
Eric R Chan, Koki Nagano, Matthew A Chan, Alexander W Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein · 2023
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2023
Later among the works it cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Later among the works it cites.
Where are we in the search for an artificial visual cortex for embodied intelligence?
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Pieter Abbeel, Jitendra Malik, et al · 2023
Later among the works it cites.
Project aria: A new tool for egocentric multi-modal ai research
Kiran Somasundaram, Jing Dong, Huixuan Tang, Julian Straub, Mingfei Yan, Michael Goesele, Jakob Julian Engel, Renzo De Nardi, and Richard Newcombe · 2023
Later among the works it cites.
Consistent view synthesis with pose-guided diffusion models
Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf · 2023
Later among the works it cites.
Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment
Zihui Xue and Kristen Grauman · 2023
Later among the works it cites.
Affordance diffusion: Synthesizing hand-object interactions
Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu · 2023
Later among the works it cites.