Fetching the paper…
Reading the bibliography…
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite.
A technique for computer detection and correction of spelling errors
Fred J Damerau · 1964
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al · 1966
Earlier work this paper cites.
Episodic and semantic memory
E. Tulving · 1972
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Church and Patrick Hanks · 1990
Earlier work this paper cites.
Action phases and mind-sets, Handbook of motivation and cognition: Foundations of social behavior
P. Gollwitzer · 1990
Earlier work this paper cites.
Minimum audible angle thresholds for sources varying in both elevation and azimuth
David R Perrott and Kourosh Saberi · 1990
Earlier work this paper cites.
https://www.nist.gov/sites/default/files/documents/2017/09/26/spk-2000-plan-v1.0.htm_.pdf
NIST SRE 2000 Evaluation Plan · 2000
Earlier work this paper cites.
Testing the correlation of word error rate and perplexity
Dietrich Klakow and Jochen Peters · 2002
Earlier work this paper cites.
Bidirectional lstm networks for improved phoneme classification and recognition
Alex Graves, Santiago Fernández, and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
The AMI meeting corpus
Iain McCowan, Jean Carletta, Wessel Kraaij, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Wilfried Post, Dennis Reidsma, and Pierre Wellner · 2005
Earlier work this paper cites.
Robust speaker diarization for meetings
Xavier Anguera Miró · 2006
Earlier work this paper cites.
Multiple object tracking performance metrics and evaluation in a smart room environment
Keni Bernardin, Alexander Elbs, and Rainer Stiefelhagen · 2006
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al · 2006
Earlier work this paper cites.
Audio-visual speech recognition using lip information extracted from side-face images
Koji Iwano, Tomoaki Yoshinaga, Satoshi Tamura, and Sadaoki Furui · 2007
Earlier work this paper cites.
The cocktail party problem: what is it? how can it be solved? and why should animal behaviorists study it?
Mark A Bee and Christophe Micheyl · 2008
Earlier work this paper cites.
Evaluating multiple object tracking performance: the clear mot metrics
Keni Bernardin and Rainer Stiefelhagen · 2008
Earlier work this paper cites.
An evaluation of psychophysical models of auditory change perception
Christophe Micheyl, Christian Kaernbach, and Laurent Demany · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Guide to the carnegie mellon university multimodal activity (cmu-mmac) database
F. De la Torre, J. Hodgins, J. Montano, S. Valcarcel, R. Forcada, and J. Macey · 2009
Earlier work this paper cites.
Audio/video fusion for objects recognition
Loic Lacheze, Yan Guo, Ryad Benosman, Bruno Gas, and Charlie Couverture · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Here’s looking at you, kid
Manuel J Marín-Jiménez, Andrew Zisserman, and Vittorio Ferrari · 2011
Earlier work this paper cites.
The Kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Speaker diarization: A review of recent research
Xavier Anguera, Simon Bozonnet, Nicholas Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals · 2012
Earlier work this paper cites.
Social interactions: A first-person perspective
Alireza Fathi, Jessica K. Hodgins, and James M. Rehg · 2012
Earlier work this paper cites.
Social interactions: A first-person perspective
A. Fathi, J. K. Hodgins, and J. M. Rehg · 2012
Earlier work this paper cites.
Activity forecasting
Kris M. Kitani, Brian Ziebart, James D. Bagnell, and Martial Hebert · 2012
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Y. J. Lee, J. Ghosh, and K. Grauman · 2012
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Y. J. Lee, J. Ghosh, and K. Grauman · 2012
Earlier work this paper cites.
Connecting meeting behavior with extraversion-a systematic study
Bruno Lepri, Ramanathan Subramanian, Kyriaki Kalimeri, Jacopo Staiano, Fabio Pianesi, and Nicu Sebe · 2012
Earlier work this paper cites.
3D social saliency from head-mounted cameras
Hyun Soo Park, Eakta Jain, and Yaser Sheikh · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
Hamed Pirsiavash and Deva Ramanan · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
H. Pirsiavash and D. Ramanan · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Modeling actions through state changes
A. Fathi and J. Rehg · 2013
Earlier work this paper cites.
Modeling actions through state changes
Alireza Fathi and James M Rehg · 2013
Earlier work this paper cites.
Ikeabot: An autonomous multi-robot coordinated furniture assembly system
Ross A Knepper, Todd Layton, John Romanishin, and Daniela Rus · 2013
Earlier work this paper cites.
Model recommendation with virtual probes for ego-centric hand detection
Cheng Li and Kris Kitani · 2013
Earlier work this paper cites.
Learning to predict gaze in egocentric video
Yin Li, Alireza Fathi, and James M. Rehg · 2013
Earlier work this paper cites.
Story-driven summarization for egocentric video
Zheng Lu and Kristen Grauman · 2013
Earlier work this paper cites.
First-person activity recognition: What are they doing to me?
M. S. Ryoo and L. Matthies · 2013
Earlier work this paper cites.
Learning 6d object pose estimation using 3d object coordinates
Eric Brachmann, Alexander Krull, Frank Michel, Stefan Gumhold, Jamie Shotton, and Carsten Rother · 2014
Earlier work this paper cites.
You-Do, I-Learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video
Dima Damen, Teesid Leelasawassuk, Osian Haines, Andrew Calway, and Walterio Mayol-Cuevas · 2014
Earlier work this paper cites.
Nonverbal Communication in Human Interaction
Mark L. Knapp, Judith A. Hall, and Terrence G. Horgan · 2014
Earlier work this paper cites.
Unsupervised feature learning for 3d scene labeling
Kevin Lai, Liefeng Bo, and Dieter Fox · 2014
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Detecting people looking at each other in videos
Manuel Jesús Marín-Jiménez, Andrew Zisserman, Marcin Eichner, and Vittorio Ferrari · 2014
Earlier work this paper cites.
Direction of arrival based spatial covariance model for blind sound source separation
Joonas Nikunen and Tuomas Virtanen · 2014
Earlier work this paper cites.
Behavioral Imaging and Autism
James M. Rehg, Agata Rozga, Gregory D. Abowd, and Matthew S. Goodwin · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions
Sven Bambach, Stefan Lee, David J. Crandall, and Chen Yu · 2015
Earlier work this paper cites.
The yale human grasping dataset: Grasp, object, and task data in household and machine shop environments
Ian M Bullock, Thomas Feix, and Aaron M Dollar · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J. Lim, and Edward H. Adelson · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J Lim, and Edward H Adelson · 2015
Earlier work this paper cites.
Predicting important objects for egocentric video summarization
Yong Jae Lee and Kristen Grauman · 2015
Earlier work this paper cites.
Personal object discovery in first-person videos
Cewu Lu, Renjie Liao, and Jiaya Jia · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
Where are they looking?
Adria Recasens, Aditya Khosla, Carl Vondrick, and Antonio Torralba · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
SHI Xingjian, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Temporal perception and prediction in ego-centric video
Yipin Zhou and Tamara L Berg · 2015
Earlier work this paper cites.
Understanding hand-object manipulation with grasp types and object attributes
Minjie Cai, Kris M Kitani, and Yoichi Sato · 2016
Earlier work this paper cites.
You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance
Dima Damen, Teesid Leelasawassuk, and Walterio Mayol-Cuevas · 2016
Earlier work this paper cites.
Summarization of egocentric videos: A comprehensive survey
Ana Garcia Del Molino, Cheston Tan, Joo-Hwee Lim, and Ah-Hwee Tan · 2016
Earlier work this paper cites.
Multiview rgb-d dataset for object instance detection
Georgios Georgakis, Md Alimoor Reza, Arsalan Mousavian, Phi-Hung Le, and Jana Košecká · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Auditory distance perception in humans: a review of cues, development, neuronal bases, and effects of sensory loss
Andrew J Kolarik, Brian CJ Moore, Pavel Zahorik, Silvia Cirstea, and Shahina Pardhan · 2016
Earlier work this paper cites.
Anticipating human activities using object affordances for reactive robotic response
Hema S. Koppula and Ashutosh Saxena · 2016
Earlier work this paper cites.
Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks
Alexei A. Efros Krishna Kumar Singh, Kayvon Fatahalian · 2016
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2016
Earlier work this paper cites.
Egocentric future localization
H. S. Park, J.-J. Hwang, Y. Niu, and J. Shi · 2016
Earlier work this paper cites.
Egocentric future localization
H. S. Park, J.-J. Hwang, Y. Niu, and J. Shi · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Cited alongside, same era.
Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks
Krishna Kumar Singh, Kayvon Fatahalian, and Alexei A Efros · 2016
Cited alongside, same era.
Detecting engagement in egocentric video
Yu-Chuan Su and Kristen Grauman · 2016
Cited alongside, same era.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
LaSOT: A High-Quality Benchmark for Large-Scale Single Object Tracking
Heng Fan, Haibin Ling, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, and Chunyuan Liao · 2019
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
Datasets for face and object detection in fisheye images
Jianglin Fu, Ivan V Bajić, and Rodney G Vaughan · 2019
Later among the works it cites.
What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention
Antonino Furnari and Giovanni Maria Farinella · 2019
Later among the works it cites.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Recognizing micro-actions and reactions from paired egocentric videos
Ryo Yonetani, Kris M. Kitani, and Yoichi Sato · 2016
Cited alongside, same era.
Visual motif discovery via first-person vision
Ryo Yonetani, Kris M Kitani, and Yoichi Sato · 2016
Cited alongside, same era.
Learning temporal transformations from time-lapse videos
Y. Zhou and T. Berg · 2016
Cited alongside, same era.
Learning temporal transformations from time-lapse videos
Yipin Zhou and Tamara L Berg · 2016
Cited alongside, same era.
Joint discovery of object states and manipulation actions
Jean-Baptiste Alayrac, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2017
Cited alongside, same era.
First-person action-object detection with egonet
Gedas Bertasius, Hyun Soo Park, Stella X. Yu, and Jianbo Shi · 2017
Cited alongside, same era.
Later among the works it cites.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Later among the works it cites.
Timeception for complex action recognition
Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders · 2019
Later among the works it cites.
Seeing through sounds: Predicting visual semantic segmentation results from multichannel audio signals
Go Irie, Mirela Ostrek, Haochen Wang, Hirokazu Kameoka, Akisato Kimura, Takahito Kawanishi, and Kunio Kashino · 2019
Later among the works it cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Later among the works it cites.
Gaze360: Physically Unconstrained Gaze Estimation in the Wild
Petr Kellnhofer, Simon Stent, Wojciech Matusik, and Antonio Torralba · 2019
Later among the works it cites.
Bmn: Boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Later among the works it cites.
Laeo-net: revisiting people looking at each other in videos
Manuel J Marin-Jimenez, Vicky Kalogeiton, Pablo Medina-Suarez, and Andrew Zisserman · 2019
Later among the works it cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Grounded human-object interaction hotspots from video
Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman · 2019
Later among the works it cites.
Future event prediction: If and when
Lukas Neumann, Andrew Zisserman, and Andrea Vedaldi · 2019
Later among the works it cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Later among the works it cites.
Task-driven modular networks for zero-shot compositional learning
Senthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, and Marc’Aurelio Ranzato · 2019
Later among the works it cites.
Egocentric visitors localization in cultural sites
F. Ragusa, A. Furnari, S. Battiato, G. Signorello, and G. M. Farinella · 2019
Later among the works it cites.
Ava-activespeaker: An audio-visual dataset for active speaker detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al · 2019
Later among the works it cites.
Learning to localize sound sources in visual scenes: Analysis and applications
A. Senocak, T.-H. Oh, J. Kim, M. Yang, and I. S. Kweon · 2019
Later among the works it cites.
Learning semantic embedding spaces for slicing vegetables
Mohit Sharma, Kevin Zhang, and Oliver Kroemer · 2019
Later among the works it cites.
The replica dataset: A digital replica of indoor spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al · 2019
Later among the works it cites.
Learning a generative model for multi-step human-object interactions from videos
He Wang, Sören Pirk, Ersin Yumer, Vladimir G Kim, Ozan Sener, Srinath Sridhar, and Leonidas J Guibas · 2019
Later among the works it cites.
Fast online object tracking and segmentation: A unifying approach, 2019
Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H. S. Torr · 2019
Later among the works it cites.
Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl · 2019
Later among the works it cites.
Self-supervised Learning of Audio-Visual Objects from Video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
Know Your Surroundings: Exploiting Scene Information for Object Tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte · 2020
Later among the works it cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Soundspaces: Audio-visual navigation in 3d environments
C. Chen, U. Jain, C. Schissler, S. V. Amengual Gari, Z. Al-Halah, V. Ithapu, P. Robinson, and K. Grauman · 2020
Later among the works it cites.
Detection of eye contact with deep neural networks is as accurate as human experts
Eunji Chong, Elysha Clark-Whitney, Audrey Southerland, Elizabeth Stubbs, Chanel Miller, Eliana L Ajodan, Melanie R Silverman, Catherine Lord, Agata Rozga, Rebecca M Jones, and James M Rehg · 2020
Later among the works it cites.
Detecting Attended Visual Targets in Video
Eunji Chong, Yongxin Wang, Nataniel Ruiz, and James M. Rehg · 2020
Later among the works it cites.
In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, and Icksang Han · 2020
Later among the works it cites.
Spot the conversation: speaker diarisation in the wild
Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, and Andrew Zisserman · 2020
Later among the works it cites.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Later among the works it cites.
CN-CELEB: a challenging Chinese speaker recognition dataset
Yue Fan, JW Kang, LT Li, KC Li, HL Chen, ST Cheng, PY Zhang, ZY Zhou, YQ Cai, and Dong Wang · 2020
Later among the works it cites.
Rolling-unrolling lstms for action anticipation from first-person video
Antonino Furnari and Giovanni Farinella · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Later among the works it cites.
spaCy: Industrial-strength Natural Language Processing in Python, 2020
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Later among the works it cites.
A multi-view dataset for learning multi-agent multi-task activities
Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu, and Song-Chun Zhu · 2020
Later among the works it cites.
The eighth visual object tracking VOT2020 challenge results, 2020
Matej Kristan, Ales Leonardis, Jiri Matas, Michael Felsberg, Roman Pflugfelder, Joni-Kristian Kamarainen, Luka Čehovin Zajc, Martin Danelljan, Alan Lukezic, Ondrej Drbohlav, Linbo He, Yushan Zhang, Song Yan, Jinyu Yang, Gustavo Fernandez, and et al · 2020
Later among the works it cites.
The open images dataset v4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Later among the works it cites.
Forecasting human-object interaction: joint prediction of motor attention and actions in first person video
Miao Liu, Siyu Tang, Yin Li, and James M Rehg · 2020
Later among the works it cites.
You2me: Inferring body pose in egocentric video via first and second person interactions
Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman · 2020
Later among the works it cites.
Egocom: A multi-person multi-modal egocentric communications dataset
C. Northcutt, S. Zha, S. Lovegrove, and R. Newcombe · 2020
Later among the works it cites.
Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, and Caroline Pantofaru · 2020
Later among the works it cites.
Superglue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich · 2020
Later among the works it cites.
Understanding human hands in contact at internet scale
Dandan Shan, Jiaqi Geng, Michelle Shu, and David Fouhey · 2020
Later among the works it cites.
Understanding human hands in contact at internet scale
Dandan Shan, Jiaqi Geng, Michelle Shu, and David Fouhey · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Later among the works it cites.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem · 2020
Later among the works it cites.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Later among the works it cites.
Learning 2d temporal adjacent networks formoment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Later among the works it cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al · 2021
Closest in time.
Rescaling egocentric vision
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2021
Closest in time.
Jacob Donley, Vladimir Tourbabin, Jung-Suk Lee, Mark Broyles, Hao Jiang, Jie Shen, Maja Pantic, Vamsi Krishna Ithapu, and Ravish Mehra · 2021
Closest in time.
Boosting image-based mutual gaze detection using pseudo 3d gaze
Bardia Doosti, Ching-Hui Chen, Raviteja Vemulapalli, Xuhui Jia, Yukun Zhu, and Bradley Green · 2021
Closest in time.
Is first person vision challenging for object tracking?
Matteo Dunnhofer, Antonino Furnari, Giovanni Maria Farinella, and Christian Micheloni · 2021
Closest in time.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Closest in time.
Dual Attention Guided Gaze Target Detection in the Wild
Yi Fang, Jiapeng Tang, Wang Shen, Wei Shen, Xiao Gu, Li Song, and Guangtao Zhai · 2021
Closest in time.
VisualVoice: Audio-visual speech separation with cross-modal consistency
R. Gao and K. Grauman · 2021
Closest in time.
Anticipative video transformer
Rohit Girdhar and Kristen Grauman · 2021
Closest in time.
GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
Lianghua Huang, Xin Zhao, and Kaiqi Huang · 2021
Closest in time.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Closest in time.
In the Eye of the Beholder: Gaze and Actions in First Person Video
Yin Li, Miao Liu, and Jame Rehg · 2021
Closest in time.
Ego-exo: Transferring visual representations from third-person to first-person videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman · 2021
Closest in time.
Deep template-based object instance detection
Jean-Philippe Mercier, Mathieu Garon, Philippe Giguere, and Jean-Francois Lalonde · 2021
Closest in time.
A review of speaker diarization: Recent advances with deep learning
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu J Han, Shinji Watanabe, and Shrikanth Narayanan · 2021
Closest in time.
The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain
Francesco Ragusa, Antonino Furnari, Salvatore Livatino, and Giovanni Maria Farinella · 2021
Closest in time.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Closest in time.
Predicting the future from first person (egocentric) vision: A survey
Ivan Rodin, Antonino Furnari, Dimitrios Mavroedis, and Giovanni Maria Farinella · 2021
Closest in time.
Silero vad: Pre-trained enterprise-grade voice activity detector (VAD), number detector and language classifier
Silero Team · 2021
Closest in time.
Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li · 2021
Closest in time.
Video self-stitching graph network for temporal action localization
Chen Zhao, Ali K Thabet, and Bernard Ghanem · 2021
Closest in time.
Embracing uncertainty: Decoupling and de-bias for robust temporal grounding
Hao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen, and Chuanping Hu · 2021
Closest in time.
Deep audio-visual learning: A survey
Hao Zhu, Man-Di Luo, Rui Wang, Ai-Hua Zheng, and Ran He · 2021
Closest in time.
Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implementation and analysis on standard tasks
Federico Landini, Ján Profant, Mireia Diez, and Lukáš Burget · 2022
Closest in time.