Fetching the paper…
Reading the bibliography…
In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role.
The hidden dimension
Edward Twitchell Hall · 1966
Earlier work this paper cites.
Environment and the spatial arrangement of conversational encounters
T Matthew Ciolek and Adam Kendon · 1980
Earlier work this paper cites.
The strength of weak ties: A network theory revisited
Mark Granovetter · 1983
Earlier work this paper cites.
The child’s theory of mind
Henry M Wellman · 1992
Earlier work this paper cites.
Gaze cueing of attention: visual attention, social cognition, and individual differences
Alexandra Frischen, Andrew P Bayliss, and Steven P Tipper · 2007
Earlier work this paper cites.
Learning to set waypoints for audio-visual navigation
Changan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao, Santhosh Kumar Ramakrishnan, and Kristen Grauman · 2008
Earlier work this paper cites.
Auditory and visual orienting responses in listeners with and without hearing-impairment
W Owen Brimijoin, David McShefferty, and Michael A Akeroyd · 2010
Earlier work this paper cites.
Social interaction discovery by statistical analysis of f-formations
Marco Cristani, Loris Bazzani, Giulia Paggetti, Andrea Fossati, Diego Tosato, Alessio Del Bue, Gloria Menegaz, and Vittorio Murino · 2011
Earlier work this paper cites.
Detecting f-formations as dominant sets
Hayley Hung and Ben Kröse · 2011
Earlier work this paper cites.
The anatomy of the facebook social graph
Johan Ugander, Brian Karrer, Lars Backstrom, and Cameron Marlow · 2011
Earlier work this paper cites.
Social interactions: A first-person perspective
Alircza Fathi, Jessica K Hodgins, and James M Rehg · 2012
Earlier work this paper cites.
Learning to discover social circles in ego networks
Jure Leskovec and Julian Mcauley · 2012
Earlier work this paper cites.
First-person activity recognition: What are they doing to me?
Michael S Ryoo and Larry Matthies · 2013
Earlier work this paper cites.
Uncovering interactions and interactors: Joint estimation of head, body orientation and f-formations from surveillance videos
Elisa Ricci, Jagannadan Varadarajan, Ramanathan Subramanian, Samuel Rota Bulo, Narendra Ahuja, and Oswald Lanz · 2015
Earlier work this paper cites.
Social saliency prediction
Hyun Soo Park and Jianbo Shi · 2015
Earlier work this paper cites.
Detecting bids for eye contact using a wearable camera
Zhefan Ye, Yin Li, Yun Liu, Chanel Bridges, Agata Rozga, and James M Rehg · 2015
Earlier work this paper cites.
Egocentric future localization
Hyun Soo Park, Jyh-Jing Hwang, Yedong Niu, and Jianbo Shi · 2016
Earlier work this paper cites.
Detecting conversational groups in images and sequences: A robust game-theoretic approach
Sebastiano Vascon, Eyasu Z Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, and Vittorio Murino · 2016
Cited alongside, same era.
Recognizing micro-actions and reactions from paired egocentric videos
Ryo Yonetani, Kris M Kitani, and Yoichi Sato · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Cited alongside, same era.
Egocentric network analysis: Foundations, methods, and models
Brea L Perry, Bernice A Pescosolido, and Stephen P Borgatti · 2018
Cited alongside, same era.
4d human body capture from egocentric video via 3d scene grounding
Miao Liu, Dexin Yang, Yan Zhang, Zhaopeng Cui, James M Rehg, and Siyu Tang · 2021
Later among the works it cites.
Cross-domain first person audio-visual action recognition through relative norm alignment
Mirco Planamente, Chiara Plizzari, Emanuele Alberti, and Barbara Caputo · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Generative adversarial network for future hand segmentation from egocentric video
Wenqi Jia, Miao Liu, and James M Rehg · 2022
Later among the works it cites.
Egocentric deep multi-channel audio-visual active speaker localization
Hao Jiang, Calvin Murdock, and Vamsi Krishna Ithapu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Future person localization in first-person videos
Takuma Yagi, Karttikeya Mangalam, Ryo Yonetani, and Yoichi Sato · 2018
Cited alongside, same era.
Recognizing f-formations in the open world
Hooman Hedayati, Daniel Szafir, and Sean Andrist · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Cited alongside, same era.
Why are conversations limited to about four people? a theoretical exploration of the conversation size constraint
Jaimie Arona Krems and Jason Wilkes · 2019
Cited alongside, same era.
Soundspaces: Audio-visual navigation in 3d environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman · 2020
Cited alongside, same era.
Visualechoes: Spatial image representation learning through echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah, Carl Schissler, and Kristen Grauman · 2020
Cited alongside, same era.
Egocom: A multi-person multi-modal egocentric communications dataset
Curtis Northcutt, Shengxin Zha, Steven Lovegrove, and Richard Newcombe · 2020
Cited alongside, same era.
In the eye of transformer: Global-local correlation for egocentric gaze estimation
Bolin Lai, Miao Liu, Fiona Ryan, and James Rehg · 2022
Later among the works it cites.
Holistic-guided disentangled learning with cross-video semantics mining for concurrent first-person and third-person activity recognition
Tianshan Liu, Rui Zhao, Wenqi Jia, Kin-Man Lam, and Jun Kong · 2022
Later among the works it cites.
Sound source selection based on head movements in natural group conversation
Hao Lu and W Owen Brimijoin · 2022
Later among the works it cites.
Learning state-aware visual representations from audible interactions
Himangi Mittal, Pedro Morgado, Unnat Jain, and Abhinav Gupta · 2022
Later among the works it cites.
Conversation group detection with spatio-temporal context
Stephanie Tan, David MJ Tax, and Hayley Hung · 2022
Later among the works it cites.
Sonicverse: A multisensory simulation platform for training household agents that see and hear
Ruohan Gao, Hao Li, Gokul Dharan, Zhuzhu Wang, Chengshu Li, Fei Xia, Silvio Savarese, Li Fei-Fei, and Jiajun Wu · 2023
Closest in time.
Egocentric audio-visual object localization
Chao Huang, Yapeng Tian, Anurag Kumar, and Chenliang Xu · 2023
Closest in time.
Quavf: Quality-aware audio-visual fusion for ego4d talking to me challenge
Hsi-Che Lin, Chien-Yi Wang, Min-Hung Chen, Szu-Wei Fu, and Yu-Chiang Frank Wang · 2023
Closest in time.
Chat2map: Efficient scene mapping from multi-ego conversations
Sagnik Majumder, Hao Jiang, Pierre Moulon, Ethan Henderson, Paul Calamia, Kristen Grauman, and Vamsi Krishna Ithapu · 2023
Closest in time.
Egocentric auditory attention localization in conversations
Fiona Ryan, Hao Jiang, Abhinav Shukla, James M Rehg, and Vamsi Krishna Ithapu · 2023
Closest in time.
Egocentric video task translation@ ego4d challenge 2022
Zihui Xue, Yale Song, Kristen Grauman, and Lorenzo Torresani · 2023
Closest in time.