Fetching the paper…
Reading the bibliography…
With the exponential growth of video data, there is an urgent need for automated technology to analyze and comprehend video content.
TDN: Temporal Difference Networks for Efficient Action Recognition
Wang, L.; Tong, Z.; Ji, B.; and Wu, G. 2021a · 1904
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Span-based localizing network for natural language video localization
Zhang, H.; Sun, A.; Jing, W.; and Zhou, J. T. 2020 · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Kuehne, H.; Arslan, A.; and Serre, T. 2014 · 2014
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
Heilbron, F. C.; Escorcia, V.; Ghanem, B.; and Niebles, J. C. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Online Action Detection
Geest, R. D.; Gavves, E.; Ghodrati, A.; Li, Z.; Snoek, C.; and Tuytelaars, T. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Gool, L. V. 2016 · 2016
Earlier work this paper cites.
CUHK & ETHZ & SIAT Submission to ActivityNet Challenge 2016
Xiong, Y.; Wang, L.; Wang, Z.; Zhang, B.; Song, H.; Li, W.; Lin, D.; Qiao, Y.; Gool, L. V.; and Tang, X. 2016 · 2016
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
The Kinetics Human Action Video Dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; Suleyman, M.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Aggregated Residual Transformations for Deep Neural Networks
Xie, S.; Girshick, R. B.; Dollár, P.; Tu, Z.; and He, K. 2017 · 2017
Earlier work this paper cites.
BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
Lin, T.; Zhao, X.; Su, H.; Wang, C.; and Yang, M. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
A Closer Look at Spatiotemporal Convolutions for Action Recognition
Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; and Paluri, M. 2018 · 2018
Earlier work this paper cites.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q. V.; and Salakhutdinov, R. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Ms-tcn: Multi-stage temporal convolutional network for action segmentation
Farha, Y. A.; and Gall, J. 2019 · 2019
Earlier work this paper cites.
SlowFast Networks for Video Recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Earlier work this paper cites.
What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention
Furnari, A.; and Farinella, G. M. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Video Classification With Channel-Separated Convolutional Networks
Tran, D.; Wang, H.; Feiszli, M.; and Torresani, L. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2020
Earlier work this paper cites.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Joshi, M.; Chen, D.; Liu, Y.; Weld, D. S.; Zettlemoyer, L.; and Levy, O. 2020 · 2020
Earlier work this paper cites.
Human action recognition using fusion of multiview and deep features: an application to video surveillance
Khan, M. A.; Javed, K.; Khan, S. A.; Saba, T.; Habib, U.; Khan, J. A.; and Abbasi, A. A. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
Temporal Aggregate Representations for Long-Range Video Understanding
Sener, F.; Singhania, D.; and Yao, A. 2020 · 2020
Cited alongside, same era.
G-TAD: Sub-Graph Localization for Temporal Action Detection
Xu, M.; Zhao, C.; Rojas, D. S.; Thabet, A. K.; and Ghanem, B. 2020 · 2020
Cited alongside, same era.
Xcit: Cross-covariance image transformers
Ali, A.; Touvron, H.; Caron, M.; Bojanowski, P.; Douze, M.; Joulin, A.; Laptev, I.; Neverova, N.; Synnaeve, G.; Verbeek, J.; et al. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Later among the works it cites.
BEiT: BERT Pre-Training of Image Transformers
Bao, H.; Dong, L.; Piao, S.; and Wei, F. 2022 · 2022
Later among the works it cites.
PaLM: Scaling Language Modeling with Pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; Schuh, P.; Shi, K.; Tsvyashchenko, S.; Maynez, J.; Rao, A.; Barnes, P.; Tay, Y.; Shazeer, N.; Prabhakaran, V.; Reif, E.; Du, N.; Hutchinson, B.; Pope, R.; Bradbury, J.; Austin, J.; Isard, M.; Gur-Ari, G.; Yin, P.; Duke, T.; Levskaya, A.; Ghemawat, S.; Dev, S.; Michalewski, H.; Garcia, X.; Misra, V.; Robinson, K.; Fedus, L.; Zhou, D.; Ippolito, D.; Luan, D.; Lim, H.; Zoph, B.; Spiridonov, A.; Sepassi, R.; Dohan, D.; Agrawal, S.; Omernick, M.; Dai, A. M.; Pillai, T. S.; Pellat, M.; Lewkowycz, A.; Moreira, E.; Child, R.; Polozov, O.; Lee, K.; Zhou, Z.; Wang, X.; Saeta, B.; Diaz, M.; Firat, O.; Catasta, M.; Wei, J.; Meier-Hellstern, K.; Eck, D.; Dean, J.; Petrov, S.; and Fiedel, N. 2022 · 2022
Later among the works it cites.
Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100
Damen, D.; Doughty, H.; Farinella, G. M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lucic, M.; and Schmid, C. 2021 · 2021
Cited alongside, same era.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Bain, M.; Nagrani, A.; Varol, G.; and Zisserman, A. 2021 · 2021
Cited alongside, same era.
Is Space-Time Attention All You Need for Video Understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Cited alongside, same era.
DCAN: Improving Temporal Action Detection via Dual Context Aggregation
Chen, G.; Zheng, Y.; Wang, L.; and Lu, T. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Multiscale Vision Transformers
Fan, H.; Xiong, B.; Mangalam, K.; Li, Y.; Yan, Z.; Malik, J.; and Feichtenhofer, C. 2021 · 2021
Cited alongside, same era.
Anticipative Video Transformer
Girdhar, R.; and Grauman, K. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2022 · 2022
Later among the works it cites.
Masked Autoencoders As Spatiotemporal Learners
Feichtenhofer, C.; Fan, H.; Li, Y.; and He, K. 2022 · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. 2022 · 2022
Later among the works it cites.
Training Compute-Optimal Large Language Models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L. A.; Welbl, J.; Clark, A.; Hennigan, T.; Noland, E.; Millican, K.; van den Driessche, G.; Damoc, B.; Guy, A.; Osindero, S.; Simonyan, K.; Elsen, E.; Rae, J. W.; Vinyals, O.; and Sifre, L. 2022 · 2022
Later among the works it cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Later among the works it cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. C. H. 2022a · 2022
Later among the works it cites.
Egocentric video-language pretraining
Lin, K. Q.; Wang, J.; Soldan, M.; Wray, M.; Yan, R.; XU, E. Z.; Gao, D.; Tu, R.-C.; Zhao, W.; Kong, W.; et al. 2022 · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue
OpenAI, T. 2022 · 2022
Later among the works it cites.
Smith, S.; Patwary, M.; Norick, B.; LeGresley, P.; Rajbhandari, S.; Casper, J.; Liu, Z.; Prabhumoye, S.; Zerveas, G.; Korthikanti, V.; Zheng, E.; Child, R.; Aminabadi, R. Y.; Bernauer, J.; Song, X.; Shoeybi, M.; He, Y.; Houston, M.; Tiwary, S.; and Catanzaro, B. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R.; De Freitas, D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.-T.; Jin, A.; Bos, T.; Baker, L.; Du, Y.; et al. 2022 · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z.; Song, Y.; Wang, J.; and Wang, L. 2022 · 2022
Later among the works it cites.
Finetuned Language Models are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Later among the works it cites.
Focal modulation networks
Yang, J.; Li, C.; Dai, X.; and Gao, J. 2022 · 2022
Later among the works it cites.
ActionFormer: Localizing Moments of Actions with Transformers
Zhang, C.; Wu, J.; and Li, Y. 2022 · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M. T.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Later among the works it cites.
Real-Time Online Video Detection with Temporal Smoothing Transformers
Zhao, Y.; and Krähenbühl, P. 2022 · 2022
Later among the works it cites.
Learning Video Representations from Large Language Models
Zhao, Y.; Misra, I.; Krähenbühl, P.; and Girdhar, R. 2022 · 2022
Later among the works it cites.
InternChat: Solving Vision-Centric Tasks by Interacting with Chatbots Beyond Language
Liu, Z.; He, Y.; Wang, W.; Wang, W.; Wang, Y.; Chen, S.; Zhang, Q.; Yang, Y.; Li, Q.; Yu, J.; et al. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory
Ramakrishnan, S. K.; Al-Halah, Z.; and Grauman, K. 2023 · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Closest in time.
Boundary-Denoising for Video Activity Localization
Xu, M.; Soldan, M.; Gao, J.; Liu, S.; Pérez-Rúa, J.-M.; and Ghanem, B. 2023 · 2023
Closest in time.