Fetching the paper…
Reading the bibliography…
Large pretrained (e.g., "foundation") models exhibit distinct capabilities depending on the domain of data they are trained on.
Hierarchical mixtures of experts and the em algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Lifelong learning algorithms
S. Thrun · 1998
Earlier work this paper cites.
Video summarization: methods and landscape
M. Barbieri, L. Agnihotri, and N. Dimitrova · 2003
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2006
Earlier work this paper cites.
Self-taught learning: transfer learning from unlabeled data
R. Raina, A. Battle, H. Lee, B. Packer, and A. Y. Ng · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Temporal segmentation and activity classification from first-person sensing
E. H. Spriggs, F. De La Torre, and M. Hebert · 2009
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Kinectfusion: real-time 3d reconstruction and interaction using a moving depth camera
S. Izadi, D. Kim, O. Hilliges, D. Molyneaux, R. Newcombe, P. Kohli, J. Shotton, S. Hodges, D. Freeman, A. Davison, et al · 2011
Earlier work this paper cites.
Unsupervised testing strategies for asr
B. Strope, D. Beeferman, A. Gruenstein, and X. Lei · 2011
Earlier work this paper cites.
Fast unsupervised ego-action learning for first-person sports videos
K. M. Kitani, T. Okabe, Y. Sato, and A. Sugimoto · 2011
Earlier work this paper cites.
Understanding egocentric activities
A. Fathi, A. Farhadi, and J. M. Rehg · 2011
Earlier work this paper cites.
Human activity prediction: Early recognition of ongoing activities from streaming videos
M. S. Ryoo · 2011
Earlier work this paper cites.
Unsupervised and transfer learning challenge: a deep learning approach
G. Mesnil, Y. Dauphin, X. Glorot, S. Rifai, Y. Bengio, I. Goodfellow, E. Lavoie, X. Muller, G. Desjardins, D. Warde-Farley, et al · 2012
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Y. J. Lee, J. Ghosh, and K. Grauman · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
H. Pirsiavash and D. Ramanan · 2012
Earlier work this paper cites.
Activity forecasting
K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert · 2012
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
First-person activity recognition: What are they doing to me?
M. S. Ryoo and L. Matthies · 2013
Earlier work this paper cites.
Pixel-level hand detection in ego-centric videos
C. Li and K. M. Kitani · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Mixture of experts: a literature survey
S. Masoudnia and R. Ebrahimpour · 2014
Earlier work this paper cites.
Max-margin early event detectors
M. Hoai and F. De la Torre · 2014
Earlier work this paper cites.
Asymmetric LSH (ALSH) for sublinear time maximum inner product search (MIPS)
A. Shrivastava and P. Li · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
A. M. Dai and Q. V. Le · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Predicting important objects for egocentric video summarization
Y. J. Lee and K. Grauman · 2015
Earlier work this paper cites.
Pooled motion features for first-person videos
M. S. Ryoo, B. Rothrock, and L. Matthies · 2015
Earlier work this paper cites.
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions
S. Bambach, S. Lee, D. J. Crandall, and C. Yu · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Unsupervised pretraining for sequence to sequence learning
P. Ramachandran, P. J. Liu, and Q. V. Le · 2016
Earlier work this paper cites.
A review on automatic speech recognition architecture and approaches
S. Karpagavalli and E. Chandra · 2016
Earlier work this paper cites.
Places: An image database for deep scene understanding
B. Zhou, A. Khosla, A. Lapedriza, A. Torralba, and A. Oliva · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Going deeper into first-person activity recognition
M. Ma, H. Fan, and K. M. Kitani · 2016
Earlier work this paper cites.
Summarization of egocentric videos: A comprehensive survey
A. G. Del Molino, C. Tan, J.-H. Lim, and A.-H. Tan · 2016
Cited alongside, same era.
Anticipating visual representations from unlabeled video
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Cited alongside, same era.
First-person activity forecasting with online inverse reinforcement learning
N. Rhinehart and K. M. Kitani · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
P. Rajpurkar, R. Jia, and P. Liang · 2018
Cited alongside, same era.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Later among the works it cites.
Transformer is all you need: Multimodal multitask learning with a unified transformer
R. Hu and A. Singh · 2021
Later among the works it cites.
Concadia: Tackling image accessibility with context
E. Kreiss, N. D. Goodman, and C. Potts · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Cited alongside, same era.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
N. C. Mithun, J. Li, F. Metze, and A. K. Roy-Chowdhury · 2018
Cited alongside, same era.
A joint sequence fusion model for video question answering and retrieval
Y. Yu, J. Kim, and G. Kim · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Charades-ego: A large-scale dataset of paired third and first person videos
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari · 2018
Cited alongside, same era.
Predicting visual features from text for image and video caption retrieval
J. Dong, X. Li, and C. G. Snoek · 2018
Cited alongside, same era.
First-person hand action benchmark with rgb-d videos and 3d hand pose annotations
G. Garcia-Hernando, S. Yuan, S. Baek, and T.-K. Kim · 2018
Cited alongside, same era.
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill · 2021
Later among the works it cites.
Clipcap: Clip prefix for image captioning
R. Mokady, A. Hertz, and A. H. Bermano · 2021
Later among the works it cites.
Clip2tv: An empirical study on transformer-based methods for video-text retrieval
Z. Gao, J. Liu, S. Chen, D. Chang, H. Zhang, and J. Yuan · 2021
Later among the works it cites.
Robust fine-tuning of zero-shot models
M. Wortsman, G. Ilharco, M. Li, J. W. Kim, H. Hajishirzi, A. Farhadi, H. Namkoong, and L. Schmidt · 2021
Later among the works it cites.
Zero-shot image-to-text generation for visual-semantic arithmetic
Y. Tewel, Y. Shalev, I. Schwartz, and L. Wolf · 2021
Later among the works it cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Later among the works it cites.
A straightforward framework for video retrieval using clip
J. A. Portillo-Quintero, J. C. Ortiz-Bayliss, and H. Terashima-Marín · 2021
Later among the works it cites.
Clip2video: Mastering video-text retrieval via image clip
H. Fang, P. Xiong, L. Xu, and Y. Chen · 2021
Later among the works it cites.
Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss
X. Cheng, H. Lin, X. Wu, F. Yang, and D. Shen · 2021
Later among the works it cites.
Ego-exo: Transferring visual representations from third-person to first-person videos
Y. Li, T. Nagarajan, B. Xiong, and K. Grauman · 2021
Later among the works it cites.
Recent advances in video question answering: A review of datasets and methods
D. Patel, R. Parikh, and Y. Shastri · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Later among the works it cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Z. Yang, Z. Gan, J. Wang, X. Hu, Y. Lu, Z. Liu, and L. Wang · 2021
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer · 2021
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Later among the works it cites.
Audio retrieval with natural language queries
A.-M. Oncescu, A. Koepke, J. F. Henriques, Z. Akata, and S. Albanie · 2021
Later among the works it cites.
Video summarization using deep neural networks: A survey
E. Apostolidis, E. Adamantidou, A. I. Metsai, V. Mezaris, and I. Patras · 2021
Later among the works it cites.
Rethinking attention with performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlós, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
Later among the works it cites.
Hopfield networks is all you need
H. Ramsauer, B. Schäfl, J. Lehner, P. Seidl, M. Widrich, L. Gruber, M. Holzleitner, T. Adler, D. P. Kreil, M. K. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2021
Later among the works it cites.
K. Choromanski, H. Chen, H. Lin, Y. Ma, A. Sehanobish, D. Jain, M. S. Ryoo, J. Varley, A. Zeng, V. Likhosherstov, D. Kalashnikov, V. Sindhwani, and A. Weller · 2021
Later among the works it cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Closest in time.
Do as i can and not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, and M. Yan · 2022
Closest in time.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Closest in time.
Xirl: Cross-embodiment inverse reinforcement learning
K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi · 2022
Closest in time.
Clip models are few-shot learners: Empirical studies on vqa and visual entailment
H. Song, L. Dong, W.-N. Zhang, T. Liu, and F. Wei · 2022
Closest in time.
Merlot reserve: Neural script knowledge through vision and language and sound
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi · 2022
Closest in time.
Large pretrained models on multimodal sentiment analysis
Y. Song, X. Fan, Y. Yang, G. Ren, and W. Pan · 2022
Closest in time.
mslam: Massively multilingual joint pre-training for speech and text
A. Bapna, C. Cherry, Y. Zhang, Y. Jia, M. Johnson, Y. Cheng, S. Khanuja, J. Riesa, and A. Conneau · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Closest in time.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Closest in time.
Language models can see: Plugging visual controls in text generation
Y. Su, T. Lan, Y. Liu, F. Liu, D. Yogatama, Y. Wang, L. Kong, and N. Collier · 2022
Closest in time.
https://cloud.google.com/speech-to-text
Speech-to-text: Automatic speech recognition | google cloud · 2022
Closest in time.
Block-nerf: Scalable large scene neural view synthesis
M. Tancik, V. Casser, X. Yan, S. Pradhan, B. Mildenhall, P. P. Srinivasan, J. T. Barron, and H. Kretzschmar · 2022
Closest in time.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Closest in time.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Closest in time.
Disentangled representation learning for text-video retrieval
Q. Wang, Y. Zhang, Y. Zheng, P. Pan, and X.-S. Hua · 2022
Closest in time.