Fetching the paper…
Reading the bibliography…
A vast amount of audio-visual data is available on the Internet thanks to video streaming services, to which users upload their content.
Referit game: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. L · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Manning, C. D., Surdeanu, M., Bauer, J., Finkel, J., Bethard, S. J., and McClosky, D · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Yu, S.-I., Jiang, L., and Hauptmann, A · 2014
Earlier work this paper cites.
Searching persuasively: Joint event detection and evidence recounting with limited supervision
Chang, X., Yu, Y.-L., Yang, Y., and Hauptmann, A. G · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, B. G. and Niebles, J. C · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Malmaud, J., Huang, J., Rathod, V., Johnston, N., Rabinovich, A., and Murphy, K · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
Sener, O., Zamir, A. R., Savarese, S., and Saxena, A · 2015
Cited alongside, same era.
You only look once: Unified, real-time object detection
Redmon, J., Divvala, S., Girshick, R., and Farhadi, A · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Rohrback, A., Rohrbach, M., Hu, R., Darrell, T., and Schiele, B · 2016
Cited alongside, same era.
See, hear, and read: Deep aligned representations
Aytar, Y., Vondrick, C., and Torralba, A · 2017
Cited alongside, same era.
Mask r-cnn
He, K., Gkioxari, G., Dollar, P., and Girshick, R · 2017
Cited alongside, same era.
Visually grounded learning of keyword prediction from untranscribed speech
Kamper, H., Settle, S., Shakhnarovich, G., and Livescu, K · 2017
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., and Rubinstein, M · 2018
Later among the works it cites.
Curriculumnet: Weakly supervised learning from large-scale web images
Guo, S., Huang, W., Zhang, H., Zhuang, C., Dong, D., Scott, M. R., and Huang, D · 2018
Later among the works it cites.
Finding “it”: Weakly-supervised, reference-aware visual grounding in instructional videos
Huang, D. A., Buch, S., Dery, L., Garg, A., Fei-Fei, L., and Niebles, J. C · 2018
Later among the works it cites.
How2: a large-scale dataset for multimodal language understanding
Sanabria, R., Caglayan, O., Palaskar, S., Elliott, D., Barrault, L., Specia, L., and Metze, F · 2018
Later among the works it cites.
The sound of pixels
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., and Torralba, A · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2017
Cited alongside, same era.
Yolo9000: Better, faster, stronger
Redmon, J. and Farhadi, A · 2017
Cited alongside, same era.
Attend in groups: a weakly-supervised deep learning framework for learning from web data
Zhuang, B., Liu, L., Li, Y., Shen, C., and Reid, I · 2017
Cited alongside, same era.
Objects that sound
Arandjelovic, R. and Zisserman, A · 2018
Cited alongside, same era.
Vision as an interlingua: Learning multilingual semantic embeddings of untranscribed speech
Harwath, D., Chuang, G., and Glass, J
Cited in the paper.
Jointly discovering visual objects and spoken words from raw sensory input
Harwath, D., Recasens, A., Surís, D., Chuang, G., Torralba, A., and Glass, J
Cited in the paper.
Later among the works it cites.
Toward self-supervised object detection in unlabeled videos
Amrani, E., Ben-Ari, R., Hakim, T., and Bronstein, A · 2019
Closest in time.
maskrcnn-benchmark: Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch
Massa, F. and Girshick, R · 2019
Closest in time.
Not all frames are equal: Weakly-supervised video grounding with contextual similarity and visual clustering losses
Shi, J., Xu, J., Gong, B., and Xu, C · 2019
Closest in time.
Grounded video description
Zhou, L., Kalantidis, Y., Chen, X., Corso, J. J., and Rohrbach, M · 2019
Closest in time.