Fetching the paper…
Reading the bibliography…
This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Multimedia Tools and Applications 25
Snoek, C.G., Worring, M.: Multimodal video indexing: A review of the state-of-the-art · 2005
Earlier work this paper cites.
In: Proceedings of the IEEE Conference on Computer Vision & Pattern Recognition (CVPR), pp. 248–255. IEEE, Miami, FL, USA (2009)
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: A large-scale hierarchical image database · 2009
Earlier work this paper cites.
In: Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pp. 91–99. Association for Computational Linguistics, Los Angeles, CA, USA (2010)
Feng, Y., Lapata, M.: Visual information in semantic representation · 2010
Earlier work this paper cites.
In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 64–71. Association for Computational Linguistics, Vancouver, Canada (2017)
Gella, S., Keller, F.: An analysis of action recognition datasets for language and vision tasks · 2011
Earlier work this paper cites.
Artificial Intelligence 193
Navigli, R., Ponzetto, S.P.: BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network · 2012
Earlier work this paper cites.
In: Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp. 644–648. Association for Computational Linguistics, Atlanta, Georgia (2013)
Dyer, C., Chahuneau, V., Smith, N.A.: A simple, fast, and effective reparameterization of IBM Model 2 · 2013
Earlier work this paper cites.
In: C.J.C. Burges, L. Bottou, M. Welling, Z. Ghahramani, K.Q. Weinberger (eds.) Advances in Neural Information Processing Systems 26, pp. 3111–3119. Curran Associates, Inc. (2013)
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed representations of words and phrases and their compositionality · 2013
Earlier work this paper cites.
In: D. Fleet, T. Pajdla, B. Schiele, T. Tuytelaars (eds.) Proceedings of the European Conference on Computer Vision (ECCV), pp. 740–755. Springer International Publishing, Zurich, Switzerland (2014)
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common objects in context · 2014
Earlier work this paper cites.
Transactions of the Association for Computational Linguistics 2
Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions · 2014
Earlier work this paper cites.
In: Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 2425–2433. IEEE, Santiago, Chile (2015)
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: VQA: Visual question answering · 2015
Earlier work this paper cites.
Chen, X., Fang, H., Lin, T.Y., Vedantam, R., Gupta, S., Dollár, P., Zitnick, C.L.: Microsoft COCO captions: Data collection and evaluation server · 2015
Earlier work this paper cites.
In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 148–158. Association for Computational Linguistics, Lisbon, Portugal (2015)
Kiela, D., Vulić, I., Clark, S.: Visual bilingual lexicon induction with transferred ConvNet features · 2015
Earlier work this paper cites.
International Journal of Computer Vision 115
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet large scale visual recognition challenge · 2015
Cited alongside, same era.
In: Proceedings of the IEEE Conference on Computer Vision & Pattern Recognition (CVPR), pp. 3156–3164. IEEE, Boston, MA, USA (2015)
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator · 2015
Cited alongside, same era.
Language and Linguistics Compass 10
Baroni, M.: Grounding distributional semantics in the visual world · 2016
Cited alongside, same era.
Journal of Artificial Intelligence Research 55
Bernardi, R., Cakici, R., Elliott, D., Erdem, A., Erdem, E., Ikizler-Cinbis, N., Keller, F., Muscat, A., Plank, B.: Automatic description generation from images: A survey of models, datasets, and evaluation measures · 2016
Cited alongside, same era.
In: Proceedings of the 5th Workshop on Vision and Language, pp. 70–74. Association for Computational Linguistics, Berlin, Germany (2016)
Elliott, D., Frank, S., Sima’an, K., Specia, L.: Multi30K: Multilingual English-German image descriptions · 2016
In: Proceedings of the Second Conference on Machine Translation, pp. 215–233. Association for Computational Linguistics, Copenhagen, Denmark (2017)
Elliott, D., Frank, S., Barrault, L., Bougares, F., Specia, L.: Findings of the second shared task on multimodal machine translation and multilingual image description · 2017
Later among the works it cites.
IEEE Transactions on Pattern Analysis and Machine Intelligence 40
Ramisa, A., Yan, F., Moreno-Noguer, F., Mikolajczyk, K.: BreakingNews: Article annotation by image and text processing · 2017
Later among the works it cites.
Aafaq, N., Gilani, S.Z., Liu, W., Mian, A.: Video description: A survey of methods, datasets and evaluation metrics · 2018
Later among the works it cites.
IEEE Transactions on Pattern Analysis and Machine Intelligence 41
Baltrušaitis, T., Ahuja, C., Morency, L.P.: Multimodal machine learning: A survey and taxonomy · 2018
Later among the works it cites.
In: Proceedings of the 27th International Conference on Computational Linguistics, pp. 2325–2339. Association for Computational Linguistics, Santa Fe, NM, USA (2018)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 182–192. Association for Computational Linguistics, San Diego, California (2016)
Gella, S., Lapata, M., Keller, F.: Unsupervised visual sense disambiguation for verbs using multimodal embeddings · 2016
Cited alongside, same era.
In: Proceedings of the IEEE Conference on Computer Vision & Pattern Recognition (CVPR), pp. 770–778. IEEE, Las Vegas, NV, USA (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Cited alongside, same era.
In: N.C.C. Chair), K. Choukri, T. Declerck, S. Goggi, M. Grobelnik, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, S. Piperidis (eds.) Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016). European Language Resources Association (ELRA), Paris, France (2016)
Hollink, L., Bedjeti, A., van Harmelen, M., Elliott, D.: A corpus of images and text in online news · 2016
Cited alongside, same era.
In: N.C.C. Chair), K. Choukri, T. Declerck, S. Goggi, M. Grobelnik, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, S. Piperidis (eds.) Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), pp. 923–929. European Language Resources Association (ELRA), Portorož, Slovenia (2016)
Lison, P., Tiedemann, J.: OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles · 2016
Cited alongside, same era.
In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1780–1790. Association for Computational Linguistics, Berlin, Germany (2016)
Miyazaki, T., Shimizu, N.: Cross-lingual image caption generation · 2016
Cited alongside, same era.
In: Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp. 435–440. Association for Computational Linguistics, San Diego, California (2016)
Paetzold, G., Specia, L.: Inferring psycholinguistic properties of words · 2016
Cited alongside, same era.
In: Proceedings of the IEEE Conference on Computer Vision & Pattern Recognition (CVPR), pp. 1080–1089. IEEE, Honolulu, HI, USA (2017)
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J.M.F., Parikh, D., Batra, D.: Visual dialog · 2017
Cited alongside, same era.
Beinborn, L., Botschen, T., Gurevych, I.: Multimodal grounding for language processing · 2018
Later among the works it cites.
In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2566–2576. Association for Computational Linguistics, Melbourne, Australia (2018)
Hewitt, J., Ippolito, D., Callahan, B., Kriz, R., Wijaya, D.T., Callison-Burch, C.: Learning translations via images with a massively multilingual image dataset · 2018
Later among the works it cites.
https://www.imdb.com/interfaces/ (2019)
IMDb: IMDb · 2018
Later among the works it cites.
In: N.C.C. chair), K. Choukri, C. Cieri, T. Declerck, S. Goggi, K. Hasida, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, S. Piperidis, T. Tokunaga (eds.) Proceedings of the Language Resources and Evaluation Conference, pp. 3810–3817. European Language Resources Association (ELRA), Miyazaki, Japan (2018)
Lala, C., Specia, L.: Multimodal lexical translation · 2018
Later among the works it cites.
Marín, J., Biswas, A., Ofli, F., Hynes, N., Salvador, A., Aytar, Y., Weber, I., Torralba, A.: Recipe1m: A dataset for learning cross-modal embeddings for cooking recipes and food images · 2018
Later among the works it cites.
http://www.opensubtitles.org/ (2019)
OpenSubtitles: Subtitles - download movie and TV Series subtitles · 2018
Later among the works it cites.
In: Proceedings of the 13th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Papers), pp. 40–153. Boston, MA, USA (2018)
Schamoni, S., Hitschler, J., Riezler, S.: A dataset and reranking method for multimodal MT of user-generated image captions · 2018
Later among the works it cites.
In: Workshop on Shortcomings in Vision and Language (SiVL) (2019)
Lala, C., Madhyastha, P., Specia, L.: Grounded word sense translation · 2019
Later among the works it cites.