Fetching the paper…
Reading the bibliography…
Reliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning.
VisualBERT: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 1908
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Climbing the tower of babel: Unsupervised multilingual learning
Snyder, B. and Barzilay, R · 2010
Earlier work this paper cites.
On achieving and evaluating language-independence in nlp
Bender, E. M · 2011
Earlier work this paper cites.
WALS Online
Dryer, M. S. and Haspelmath, M. (eds.) · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Visual bilingual lexicon induction with transferred ConvNet features
Kiela, D., Vulić, I., and Clark, S · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Bernardi, R., Cakici, R., Elliott, D., Erdem, A., Erdem, E., Ikizler-Cinbis, N., Keller, F., Muscat, A., and Plank, B · 2016
Earlier work this paper cites.
Multi30K: Multilingual English-German image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
TasvirEt: Görüntülerden otomatik türkçe açıklama olusturma Için bir denektaçı veri kümesi (TasvirEt: A benchmark dataset for automatic Turkish description generation from images)
Unal, M. E., Citamak, B., Yagcioglu, S., Erdem, A., Erdem, E., Cinbis, N. I., and Cakici, R · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
Findings of the second shared task on multimodal machine translation and multilingual image description
Elliott, D., Frank, S., Barrault, L., Bougares, F., and Specia, L · 2017
Earlier work this paper cites.
Image pivoting for learning multilingual multimodal representations
Gella, S., Sennrich, R., Keller, F., and Lapata, M · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., Bernstein, M. S., and Fei-Fei, L · 2017
Earlier work this paper cites.
Fluency-guided cross-lingual image captioning
Lan, W., Li, X., and Dong, J · 2017
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Littell, P., Mortensen, D. R., Lin, K., Kairis, K., Turner, C., and Levin, L · 2017
Earlier work this paper cites.
Cross-linguistic differences and similarities in image descriptions
van Miltenburg, E., Elliott, D., and Vossen, P · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollar, P., Tu, Z., and He, K · 2017
Earlier work this paper cites.
STAIR captions: Constructing a large-scale Japanese image caption dataset
Yoshikawa, Y., Shigeto, Y., and Takeuchi, A · 2017
Earlier work this paper cites.
Baselines and test data for cross-lingual inference
Agić, Ž. and Schluter, N · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L · 2018
Earlier work this paper cites.
Findings of the third shared task on multimodal machine translation
Barrault, L., Bougares, F., Specia, L., Lala, C., Elliott, D., and Frank, S · 2018
Earlier work this paper cites.
Assessing multilingual multimodal image description: Studies of native speaker preferences and translator choices
Frank, S., Elliott, D., and Specia, L · 2018
Earlier work this paper cites.
Neural task representations as weak supervision for model agnostic cross-lingual transfer
Jauhar, S. K., Gamon, M., and Pantel, P · 2018
Earlier work this paper cites.
Learning visually grounded sentence representations
Kiela, D., Conneau, A., Jabri, A., and Nickel, M · 2018
Earlier work this paper cites.
Bridging languages through images with deep partial canonical correlation analysis
Rotman, G., Vulić, I., and Reichart, R · 2018
Cited alongside, same era.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Cited alongside, same era.
DIDEC: The Dutch image description and eye-tracking corpus
van Miltenburg, E., Kádár, Á., Koolen, R., and Krahmer, E · 2018
Cited alongside, same era.
Grounded textual entailment
Vu, H. T., Greco, C., Erofeeva, A., Jafaritazehjan, S., Linders, G., Tanti, M., Testoni, A., Bernardi, R., and Gatt, A · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
XCOPA: A multilingual dataset for causal commonsense reasoning
Ponti, E. M., Glavaš, G., Majewska, O., Liu, Q., Vulić, I., and Korhonen, A · 2020
Later among the works it cites.
RussianSuperGLUE: A Russian language understanding evaluation benchmark
Shavrina, T., Fenogenova, A., Anton, E., Shevelev, D., Artemova, E., Malykh, V., Mikhailov, V., Tikhonova, M., Chertok, A., and Evlampiev, A · 2020
Later among the works it cites.
VL-BERT: Pre-training of generic visual-linguistic representations
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., and Dai, J · 2020
Later among the works it cites.
IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., and Purwarianti, A · 2020
Later among the works it cites.
Multimodal transformer for multimodal machine translation
Yao, S. and Wan, X · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Baltrusaitis, T., Ahuja, C., and Morency, L.-P · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Conneau, A. and Lample, G · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Multi-head attention with diversity for learning grounded multilingual multimodal representations
Huang, P.-Y., Chang, X., and Hauptmann, A · 2019
Cited alongside, same era.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Towards zero-shot cross-lingual image retrieval and tagging
Aggarwal, P., Tambi, R., and Kale, A · 2021
Later among the works it cites.
Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and-language BERTs
Bugliarello, E., Cotterell, R., Okazaki, N., and Elliott, D · 2021
Later among the works it cites.
Cross-lingual visual pre-training for multimodal machine translation
Caglayan, O., Kuyu, M., Amac, M. S., Madhyastha, P., Erdem, E., Erdem, A., and Specia, L · 2021
Later among the works it cites.
Liro: Benchmark and leaderboard for romanian language tasks
Dumitrescu, S. D., Rebeja, P., Lorincz, B., Gaman, M., Avram, A., Ilie, M., Pruteanu, A., Stan, A., Rosia, L., Iacobescu, C., Morogan, L., Dima, G., Marchidan, G., Rebedea, T., Chitez, M., Yogatama, D., Ruder, S., Ionescu, R. T., Pascanu, R., and Patraucean, V · 2021
Later among the works it cites.
Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models
Huang, P.-Y., Patrick, M., Hu, J., Neubig, G., Metze, F., and Hauptmann, A · 2021
Later among the works it cites.
MURAL: Multimodal, multitask representations across languages
Jain, A., Guo, M., Srinivasan, K., Chen, T., Kudugunta, S., Jia, C., Yang, Y., and Baldridge, J · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Later among the works it cites.
mTVR: Multilingual moment retrieval in videos
Lei, J., Berg, T., and Bansal, M · 2021
Later among the works it cites.
VALUE: A multi-task benchmark for video-and-language understanding evaluation
Li, L., Lei, J., Gan, Z., Yu, L., Chen, Y.-C., Pillai, R., Cheng, Y., Zhou, L., Wang, X. E., Wang, W. Y., Berg, T. L., Bansal, M., Liu, J., Wang, L., and Liu, Z · 2021
Later among the works it cites.
Visually grounded reasoning across languages and cultures
Liu, F., Bugliarello, E., Ponti, E. M., Reddy, S., Collier, N., and Elliott, D · 2021
Later among the works it cites.
A Hindi image caption generation framework using deep learning
Mishra, S. K., Dhir, R., Saha, S., and Bhattacharyya, P · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Ni, M., Huang, H., Su, L., Cui, E., Bharti, T., Wang, L., Zhang, D., and Duan, N · 2021
Later among the works it cites.
KLUE: Korean language understanding evaluation
Park, S., Moon, J., Kim, S., Cho, W. I., Han, J. Y., Park, J., Song, C., Kim, J., Song, Y., Oh, T., Lee, J., Oh, J., Lyu, S., Jeong, Y., Lee, I., Seo, S., Lee, D., Kim, H., Lee, M., Jang, S., Do, S., Kim, S., Lim, K., Lee, J., Park, K., Shin, J., Kim, S., Park, L., Oh, A., Ha, J.-W., and Cho, K · 2021
Later among the works it cites.
Minimax and neyman–Pearson meta-learning for outlier languages
Ponti, E. M., Aralikatte, R., Shrivastava, D., Reddy, S., and Søgaard, A · 2021
Later among the works it cites.
XTREME-R: Towards more challenging and nuanced multilingual evaluation
Ruder, S., Constant, N., Botha, J., Siddhant, A., Firat, O., Fu, J., Liu, P., Hu, J., Garrette, D., Neubig, G., and Johnson, M · 2021
Later among the works it cites.
It’s not just size that matters: Small language models are also few-shot learners
Schick, T. and Schütze, H · 2021
Later among the works it cites.
WIT: Wikipedia-Based Image Text Dataset for Multimodal Multilingual Machine Learning , pp. 2443–2449
Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M · 2021
Later among the works it cites.
GEM: A general evaluation benchmark for multimodal tasks
Su, L., Duan, N., Cui, E., Ji, L., Wu, C., Luo, H., Liu, Y., Zhong, M., Bharti, T., and Sacheti, A · 2021
Later among the works it cites.
MultiSubs: A large-scale multimodal and multilingual dataset
Wang, J., Madhyastha, P., Figueiredo, J., Lala, C., and Specia, L · 2021
Later among the works it cites.
Superb: Speech processing universal performance benchmark
Yang, S.-w., Chi, P.-H., Chuang, Y.-S., Lai, C.-I., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G.-T., hsien Huang, T., Tseng, W.-C., tik Lee, K., Liu, D.-R., Huang, Z., Dong, S., Li, S.-W., Watanabe, S., rahman Mohamed, A., and yi Lee, H · 2021
Later among the works it cites.
Broaden the vision: Geo-diverse visual commonsense reasoning
Yin, D., Li, L. H., Hu, Z., Peng, N., and Chang, K.-W · 2021
Later among the works it cites.
A closer look at few-shot crosslingual transfer: The choice of shots matters
Zhao, M., Zhu, Y., Shareghi, E., Vulić, I., Reichart, R., Korhonen, A., and Schütze, H · 2021
Later among the works it cites.
Uc2: Universal cross-lingual cross-modal vision-and-language pre-training
Zhou, M., Zhou, L., Wang, S., Cheng, Y., Li, L., Yu, Z., and Liu, J · 2021
Later among the works it cites.
Composable sparse fine-tuning for cross-lingual transfer
Ansell, A., Ponti, E., Korhonen, A., and Vulić, I · 2022
Closest in time.
Systematic inequalities in language technology performance across the world’s languages
Blasi, D., Anastasopoulos, A., and Neubig, G · 2022
Closest in time.
Retrieve fast, rerank smart: Cooperative and joint approaches for improved cross-modal retrieval
Geigle, G., Pfeiffer, J., Reimers, N., Vulić, I., and Gurevych, I · 2022
Closest in time.
xGQA: Cross-lingual visual question answering
Pfeiffer, J., Geigle, G., Kamath, A., Steitz, J.-M., Roth, S., Vulić, I., and Gurevych, I · 2022
Closest in time.