Fetching the paper…
Reading the bibliography…
We address the challenging problem of Natural Language Comprehension beyond plain-text documents by introducing the TILT neural network architecture which simultaneously learns layout information, visual features, and textual semantics.
Harley, A.W., Ufkes, A., Derpanis, K.G.: Evaluation of deep convolutional nets for document image classification and retrieval. In: ICDAR (2015)
2015
Earlier work this paper cites.
Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional Networks for Biomedical Image Segmentation. In: MICCAI (2015)
2015
Earlier work this paper cites.
Dai, J., Li, Y., He, K., Sun, J.: R-FCN: Object detection via region-based fully convolutional networks. In: NeurIPS (2016)
2016
Earlier work this paper cites.
Hewlett, D., Lacoste, A., Jones, L., Polosukhin, I., Fandrianto, A., Han, J., Kelcey, M., Berthelot, D.: WikiReading: A novel large-scale language understanding task over Wikipedia. In: ACL (2016)
2016
Earlier work this paper cites.
Kumar, A., Irsoy, O., Ondruska, P., Iyyer, M., Bradbury, J., Gulrajani, I., Zhong, V., Paulus, R., Socher, R.: Ask me anything: Dynamic memory networks for natural language processing. In: ICML (2016)
2016
Earlier work this paper cites.
Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P.: SQuAD: 100,000+ questions for machine comprehension of text. In: EMNLP (2016)
2016
Earlier work this paper cites.
Sennrich, R., Haddow, B., Birch, A.: Neural machine translation of rare words with subword units. In: ACL (2016)
2016
Earlier work this paper cites.
Lai, G., Xie, Q., Liu, H., Yang, Y., Hovy, E.: RACE: Large-scale ReAding comprehension dataset from examinations. In: EMNLP (2017)
2017
Earlier work this paper cites.
Palm, R.B., Winther, O., Laws, F.: CloudScan - a configuration-free invoice analysis system using recurrent neural networks. In: ICDAR (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Cho, M., Amplayo, R., Hwang, S.w., Park, J.: Adversarial TableQA: Attention supervision for question answering on tables. In: PMLR (2018)
2018
Earlier work this paper cites.
Choi, E., He, H., Iyyer, M., Yatskar, M., tau Yih, W., Choi, Y., Liang, P., Zettlemoyer, L.: QuAC: Question answering in context. In: EMNLP (2018)
2018
Earlier work this paper cites.
Kafle, K., Price, B.L., Cohen, S., Kanan, C.: DVQA: understanding data visualizations via question answering. In: CVPR (2018)
2018
Earlier work this paper cites.
Kahou, S.E., Michalski, V., Atkinson, A., Kádár, Á., Trischler, A., Bengio, Y.: FigureQA: An annotated figure dataset for visual reasoning. In: ICLR (2018)
2018
Earlier work this paper cites.
Kudo, T.: Subword regularization: Improving neural network translation models with multiple subword candidates. In: ACL (2018)
2018
Earlier work this paper cites.
Lee, K.H., Chen, X., Hua, G., Hu, H., He, X.: Stacked cross attention for image-text matching. In: ECCV (2018)
2018
Earlier work this paper cites.
McCann, B., Keskar, N.S., Xiong, C., Socher, R.: The natural language decathlon: Multitask learning as question answering (2018), arXiv preprint
2018
Earlier work this paper cites.
Denk, T.I., Reisswig, C.: BERTgrid: Contextualized embedding for 2d document representation and understanding (2019), arXiv preprint
2019
Earlier work this paper cites.
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., Gardner, M.: DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In: NAACL-HLT (2019)
2019
Earlier work this paper cites.
Ethayarajh, K.: How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In: EMNLP-IJCNLP (2019)
2019
Cited alongside, same era.
Ho, J., Kalchbrenner, N., Weissenborn, D., Salimans, T.: Axial attention in multidimensional transformers (2019), arXiv preprint
2019
Cited alongside, same era.
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., Jawahar, C.: ICDAR2019 competition on scanned receipt OCR and information extraction. In: ICDAR (2019)
2019
Cited alongside, same era.
Jaume, G., Ekenel, H.K., Thiran, J.P.: FUNSD: A dataset for form understanding in noisy scanned documents. In: ICDAR-OST (2019)
2019
Cited alongside, same era.
Keskar, N., McCann, B., Xiong, C., Socher, R.: Unifying question answering and text classification via span extraction (2019), arXiv preprint
2019
Herzig, J., Nowak, P.K., Müller, T., Piccinno, F., Eisenschlos, J.: TaPas: Weakly supervised table parsing via pre-training. In: ACL (2020)
2020
Later among the works it cites.
Hwang, W., Yim, J., Park, S., Yang, S., Seo, M.: Spatial dependency parsing for semi-structured document information extraction (2020), arXiv preprint
2020
Later among the works it cites.
Kasai, J., Pappas, N., Peng, H., Cross, J., Smith, N.A.: Deep encoder, shallow decoder: Reevaluating the speed-quality tradeoff in machine translation (2020), arXiv preprint
2020
Later among the works it cites.
Khashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., Hajishirzi, H.: UnifiedQA: Crossing format boundaries with a single QA system. In: EMNLP-Findings (2020)
2020
Later among the works it cites.
Khot, T., Clark, P., Guerquin, M., Jansen, P., Sabharwal, A.: QASC: A dataset for question answering via sentence composition. In: AAAI (2020)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.W., Dai, A.M., Uszkoreit, J., Le, Q., Petrov, S.: Natural questions: A benchmark for question answering research. TACL (2019)
2019
Cited alongside, same era.
Le, H., Sahoo, D., Chen, N., Hoi, S.: Multimodal transformer networks for end-to-end video-grounded dialogue systems. In: ACL (2019)
2019
Cited alongside, same era.
Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., Chang, K.W.: VisualBERT: A simple and performant baseline for vision and language (2019), arXiv preprint
2019
Cited alongside, same era.
Liu, X., Gao, F., Zhang, Q., Zhao, H.: Graph convolution for multimodal information extraction from visually rich documents. In: NAACL-HLT (2019)
2019
Cited alongside, same era.
Ma, J., Qin, S., Su, L., Li, X., Xiao, L.: Fusion of image-text attention for transformer-based multimodal machine translation. In: IALP (2019)
2019
Cited alongside, same era.
Park, S., Shin, S., Lee, B., Lee, J., Surh, J., Seo, M., Lee, H.: CORD: A consolidated receipt dataset for post-ocr parsing. In: Document Intelligence Workshop at NeurIPS (2019)
2019
Cited alongside, same era.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners (2019), technical report
2019
Cited alongside, same era.
2020
Later among the works it cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., Zettlemoyer, L.: BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In: ACL (2020)
2020
Later among the works it cites.
Powalski, R., Stanislawek, T.: UniCase – rethinking casing in language models (2020), arXiv prepint
2020
Later among the works it cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. JMRL (2020)
2020
Later among the works it cites.
Ren, Y., Liu, J., Tan, X., Zhao, Z., Zhao, S., Liu, T.Y.: A study of non-autoregressive model for sequence generation. In: ACL (2020)
2020
Later among the works it cites.
Sidorov, O., Hu, R., Rohrbach, M., Singh, A.: TextCaps: A dataset for image captioning with reading comprehension. In: ECCV (2020)
2020
Later among the works it cites.
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., Dai, J.: VL-BERT: pre-training of generic visual-linguistic representations. In: ICLR (2020)
2020
Later among the works it cites.
Xu, Y., Xu, Y., Lv, T., Cui, L., Wei, F., Wang, G., Lu, Y., Florencio, D., Zhang, C., Che, W., Zhang, M., Zhou, L.: LayoutLMv2: Multi-modal pre-training for visually-rich document understanding (2020), arXiv preprint
2020
Later among the works it cites.
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., Zhou, M.: LayoutLM: Pre-training of text and layout for document image understanding. In: KDD (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Łukasz Garncarek, Powalski, R., Stanisławek, T., Topolski, B., Halama, P., Turski, M., Graliński, F.: LAMBERT: Layout-aware (language) modeling using bert for information extraction (2021), accepted to ICDAR 2021
2021
Closest in time.
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., Yang, Z., Zhang, Y., Tao, D.: A survey on visual transformer (2021), arXiv preprint
2021
Closest in time.
Hong, T., Kim, D., Ji, M., Hwang, W., Nam, D., Park, S.: BROS: A pre-trained language model for understanding texts in document (2021), openreview.net preprint
2021
Closest in time.
Mathew, M., Karatzas, D., Jawahar, C.: DocVQA: A dataset for VQA on document images. In: WACV (2021)
2021
Closest in time.
Stanisławek, T., Graliński, F., Wróblewska, A., Lipiński, D., Kaliska, A., Rosalska, P., Topolski, B., Biecek, P.: Kleister: Key information extraction datasets involving long documents with complex layouts (2021), accepted to ICDAR 2021
2021
Closest in time.