Fetching the paper…
Reading the bibliography…
We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
METEOR: an automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Associating semantically structured cooking videos with their preparation steps
K. Miura, M. Takano, R. Hamada, I. Ide, S. Sakai, and H. Tanaka · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
L. V. D. Maaten and G. Hinton · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
E. Tzeng, J. Hoffman, N. Zhang, K. Saenko, and T. Darrell · 2014
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Y. Ganin and V. Lempitsky · 2015
Earlier work this paper cites.
Delving into egocentric actions
Y. Li, Z. Ye, and J. M. Rehg · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Simultaneous deep transfer across domains and tasks
E. Tzeng, J. Hoffman, T. Darrell, and K. Saenko · 2015
Earlier work this paper cites.
CIDEr: consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2016
Earlier work this paper cites.
Simple online and realtime tracking
A. Bewley, Z. Ge, L. Ott, F. T. Ramos, and B. Upcroft · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Earlier work this paper cites.
Unified deep supervised domain adaptation and generalization
S. Motiian, M. Piccirilli, D. A. Adjeroh, and G. Doretto · 2017
Earlier work this paper cites.
Partial adversarial domain adaptation
Z. Cao, L. Ma, M. Long, and J. Wang · 2018
Earlier work this paper cites.
Actor and observer: joint modeling of first and third-person videos
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari · 2018
Earlier work this paper cites.
Bidirectional attentive fusion with context gating for dense video captioning
J. Wang, W. Jiang, L. Ma, W. Liu, and Y. Xu · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong · 2018
Earlier work this paper cites.
Temporal attentive alignment for large-scale video domain adaptation
M.-H. Chen, Z. Kira, G. Alregib, J. Yoo, R. Chen, and J. Zheng · 2019
Earlier work this paper cites.
Taking a closer look at domain shift: Category-level adversaries for semantics consistent domain adaptation
Y. Luo, L. Zheng, T. Guan, J. Yu, and Y. Yang · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J. B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Earlier work this paper cites.
COIN: A large-scale dataset for comprehensive instructional video analysis
Y. Tang, D. Ding, Y. Rao, Y. Zheng, D. Zhang, L. Zhao, J. Lu, and J. Zhou · 2019
Earlier work this paper cites.
H+O: unified egocentric recognition of 3D hand-object poses and interactions
B. Tekin, F. Bogo, and M. Pollefeys · 2019
Earlier work this paper cites.
Temporal segment networks for action recognition in videos
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2019
Cited alongside, same era.
Youmakeup: A large-scale domain-specific multimodal dataset for fine-grained semantic comprehension
W. Wang, Y. Wang, S. Chen, and Q. Jin · 2019
Cited alongside, same era.
Unsupervised and semi-supervised domain adaptation for action recognition from drones
J. Choi, G. Sharma, M. Chandraker, and J.-B. Huang · 2020
Cited alongside, same era.
SODA: story oriented dense video captioning evaluation framework
S. Fujita, T. Hirao, H. Kamigaito, M. Okumura, and M. Nagata · 2020
Cited alongside, same era.
Understanding self-training for gradual domain adaptation
A. Kumar, T. Ma, and P. Liang · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
MovieCuts: a new dataset and benchmark for cut type recognition
A. Pardo, F. C. Heilbron, J. L. Alcázar, A. K. Thabet, and B. Ghanem · 2022
Later among the works it cites.
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
F. Sener, D. Chatterjee, D. Shelepov, K. He, D. Singhania, R. Wang, and A. Yao · 2022
Later among the works it cites.
Look for the change: Learning object states and state-modifying actions from untrimmed web videos
T. Soucek, J.-B. Alayrac, A. Miech, I. Laptev, and J. Sivic · 2022
Later among the works it cites.
Background mixup data augmentation for hand and object-in-contact detection
K. Tango, T. Ohkawa, R. Furuta, and Y. Sato · 2022
Later among the works it cites.
Understanding gradual domain adaptation: Improved analysis, optimal path and beyond
H. Wang, B. Li, and H. Zhao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman · 2020
Cited alongside, same era.
Multi-modal domain adaptation for fine-grained action recognition
J. Munro and D. Damen · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. Fouhey · 2020
Cited alongside, same era.
Shot contrastive self-supervised learning for scene boundary detection
S. Chen, X. Nie, D. Fan, D. Zhang, V. Bhat, and R. Hamid · 2021
Cited alongside, same era.
Learning cross-modal contrastive features for video domain adaptation
D. Kim, Y.-H. Tsai, B. Zhuang, X. Yu, S. Sclaroff, K. Saenko, and M. Chandraker · 2021
Cited alongside, same era.
H2O: two hands manipulating objects for first person interaction recognition
T. Kwon, B. Tekin, J. Stühmer, F. Bogo, and M. Pollefeys · 2021
Cited alongside, same era.
Ego-exo: Transferring visual representations from third-person to first-person videos
Y. Li, T. Nagarajan, B. Xiong, and K. Grauman · 2021
Cited alongside, same era.
L. Zhang, S. Zhou, S. Stent, and J. Shi · 2022
Later among the works it cites.
Video captioning based on both egocentric and exocentric views of robot vision for human-robot interaction
S.-H. Kang and J.-H. Han · 2023
Closest in time.
Ego-humans: An ego-centric 3d multi-human benchmark
R. Khirodkar, A. Bansal, L. Ma, R. Newcombe, M. Vo, and K. Kitani · 2023
Closest in time.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollar, and R. Girshick · 2023
Closest in time.
Cross-domain video action recognition via adaptive gradual learning
D. Liu, Z. Bao, J. Mi, Y. Gan, M. Ye, and J. Zhang · 2023
Closest in time.
AssemblyHands: Towards egocentric activity understanding via 3D hand pose estimation
T. Ohkawa, K. He, F. Sener, T. Hodan, L. Tran, and C. Keskin · 2023
Closest in time.
Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
S. Pramanick, Y. Song, S. Nag, K. Q. Lin, H. Shah, M. Z. Shou, R. Chellappa, and P. Zhang · 2023
Closest in time.
A new dataset and approach for timestamp supervised action segmentation using human object interaction
S. I. Sayed, R. Ghoddoosian, B. Trivedi, and V. Athitsos · 2023
Closest in time.
Cross-view action recognition understanding from exocentric to egocentric perspective
T.-D. Truong and K. Luu · 2023
Closest in time.
Learning from semantic alignment between unpaired multiviews for egocentric video recognition
Q. Wang, L. Zhao, L. Yuan, T. Liu, and X. Peng · 2023
Closest in time.
NewsNet: a novel dataset for hierarchical temporal segmentation
H. Wu, K. Chen, H. Liu, M. Zhuge, B. Li, R. Qiao, X. Shu, B. Gan, L. Xu, B. Ren, M. Xu, W. Zhang, R. Ramachandra, C.-W. Lin, and B. Ghanem · 2023
Closest in time.
Z. Xue and K. Grauman · 2023
Closest in time.
VidChapters-7M: video chapters at scale
A. Yang, A. Nagrani, I. Laptev, J. Sivic, and C. Schmid · 2023
Closest in time.
Vid2Seq: large-scale pretraining of a visual language model for dense video captioning
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid · 2023
Closest in time.
Beyond Instructions: a taxonomy of information types in how-to videos
S. Yang, S. Kwak, J. Lee, and J. Kim · 2023
Closest in time.
Fine-grained affordance annotation for egocentric hand-object interaction videos
Z. Yu, Y. Huang, R. Furuta, T. Yagi, Y. Goutsu, and Y. Sato · 2023
Closest in time.
Hierarchical video-moment retrieval and step-captioning
A. Zala, J. Cho, S. Kottur, X. Chen, B. Oguz, Y. Mehdad, and M. Bansal · 2023
Closest in time.
Benchmarks and challenges in pose estimation for egocentric hand interactions with objects
Z. Fan, T. Ohkawa, L. Yang, N. Lin, Z. Zhou, S. Zhou, J. Liang, Z. Gao, X. Zhang, X. Zhang, F. Li, L. Zheng, F. Lu, K. A. Zeid, B. Leibe, J. On, S. Baek, A. Prakash, S. Gupta, K. He, Y. Sato, O. Hilliges, H. J. Chang, and A. Yao · 2024
Closest in time.
Generative hierarchical temporal transformer for hand action recognition and motion prediction
Y. Wen, H. Pan, T. Ohkawa, L. Yang, J. Pan, Y. Sato, T. Komura, and W. Wang · 2024
Closest in time.
Retrieval-augmented egocentric video captioning
J. Xu, Y. Huang, J. Hou, G. Chen, Y. Zhang, R. Feng, and W. Xie · 2024
Closest in time.