Fetching the paper…
Reading the bibliography…
In the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Bert-based multi-head selection for joint entity-relation extraction
Weipeng Huang, Xingyi Cheng, Taifeng Wang, and Wei Chu. 2019 · 1908
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Zilong Wang, Mingjie Zhan, Xuebo Liu, and Ding Liang. 2020 · 2010
Earlier work this paper cites.
A probabilistic approach to printed document understanding
Eric Medvet, Alberto Bartoli, and Giorgio Davanzo. 2011 · 2011
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jean-Philippe Thiran Guillaume Jaume, Hazim Kemal Ekenel. 2019 · 2019
Earlier work this paper cites.
End-to-end neural relation extraction using deep biaffine attention
Dat Quoc Nguyen and Karin Verspoor. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Low-resource response generation with template prior
Ze Yang, Wei Wu, Jian Yang, Can Xu, and Zhoujun Li. 2019 · 2019
Earlier work this paper cites.
An improved privacy-preserving stochastic gradient descent algorithm
Xianfu Cheng, Yanqing Yao, and Ao Liu. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
A survey on deep learning for named entity recognition
Jing Li, Aixin Sun, Jianglei Han, and Chenliang Li. 2020 · 2020
Earlier work this paper cites.
Improving neural machine translation with soft template prediction
Jian Yang, Shuming Ma, Dongdong Zhang, Zhoujun Li, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Pick: processing key information extraction from documents using improved graph learning-convolutional networks
Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao. 2021 · 2020
Cited alongside, same era.
Trie: end-to-end text reading and information extraction for document understanding
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. 2020 · 2020
Cited alongside, same era.
Rethinking pre-training and self-training
Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk, and Quoc Le. 2020 · 2020
Cited alongside, same era.
Infoxlm: An information-theoretic framework for cross-lingual language model pre-training
Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, He-Yan Huang, and Ming Zhou. 2021 · 2021
Cited alongside, same era.
Chatgpt for shaping the future of dentistry: the potential of multi-modal large language model
Hanyao Huang, Ou Zheng, Dongdong Wang, Jiayi Yin, Zijin Wang, Shengxuan Ding, Heng Yin, Chuan Xu, Renjie Yang, Qian Zheng, et al. 2023 · 2023
Later among the works it cites.
Trocr: Transformer-based optical character recognition with pre-trained models
Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei. 2023 · 2023
Later among the works it cites.
Centre: A paragraph-level chinese dataset for relation extraction among enterprises
Peipei Liu, Hong Li, Zhiyu Wang, Yimo Ren, Jie Liu, Fei Lv, Hongsong Zhu, and Limin Sun. 2023 · 2023
Later among the works it cites.
Geolayoutlm: Geometric pre-training for visual information extraction
Chuwei Luo, Changxu Cheng, Qi Zheng, and Cong Yao. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lei Cui, Yiheng Xu, Tengchao Lv, and Furu Wei. 2021 · 2021
Cited alongside, same era.
Structurallm: Structural pre-training for form understanding
Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2021 · 2021
Cited alongside, same era.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. 2021 · 2021
Cited alongside, same era.
An improved stochastic gradient descent algorithm based on rényi differential privacy
XianFu Cheng, YanQing Yao, Liying Zhang, Ao Liu, and Zhoujun Li. 2022 · 2022
Cited alongside, same era.
Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents
Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2022 · 2022
Cited alongside, same era.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, et al. 2022 · 2022
Cited alongside, same era.
Deep purified feature mining model for joint named entity recognition and relation extraction
Youwei Wang, Ying Wang, Zhongchuan Sun, Yinghao Li, Shizhe Hu, and Yangdong Ye. 2023 · 2023
Later among the works it cites.
Multi-stage pre-training enhanced by chatgpt for multi-scenario multi-domain dialogue summarization
Weixiao Zhou, Gengyao Li, Xianfu Cheng, Xinnian Liang, Junnan Zhu, Feifei Zhai, and Zhoujun Li. 2023 · 2023
Later among the works it cites.
xcot: Cross-lingual instruction tuning for cross-lingual chain-of-thought reasoning
Linzheng Chai, Jian Yang, Tao Sun, Hongcheng Guo, Jiaheng Liu, Bing Wang, Xinnian Liang, Jiaqi Bai, Tongliang Li, Qiyao Peng, and Zhoujun Li. 2024 · 2024
Closest in time.
Sviptr: Fast and efficient scene text recognition with vision permutable extractor
Xianfu Cheng, Weixiao Zhou, Xiang Li, Jian Yang, Hang Zhang, Tao Sun, Wei Zhang, Yuying Mai, Tongliang Li, Xiaoming Chen, et al. 2024 · 2024
Closest in time.
Information extraction with differentiable beam search on graph rnns
Niama El Khbir, Nadi Tomeh, and Thierry Charnois. 2024 · 2024
Closest in time.
Layoutllm: Large language model instruction tuning for visually rich document understanding
Masato Fujitake. 2024 · 2024
Closest in time.
Gpt-4o: The cutting-edge advancement in multimodal llm
Raisa Islam and Owana Marzia Moushi. 2024 · 2024
Closest in time.
Span-based joint entity and relation extraction augmented with sequence tagging mechanism
Bin Ji, Shasha Li, Hao Xu, Jie Yu, Jun Ma, Huijun Liu, and Jing Yang. 2024 · 2024
Closest in time.
Entity-relation extraction as full shallow semantic dependency parsing
Shu Jiang, Zuchao Li, Hai Zhao, and Weiping Ding. 2024 · 2024
Closest in time.
Adaptive token selection and fusion network for multimodal sentiment analysis
Xiang Li, Ming Lu, Ziming Guo, and Xiaoming Zhang. 2024 · 2024
Closest in time.
m3p: Towards multimodal multilingual translation with multimodal prompt
Jian Yang, Hongcheng Guo, Yuwei Yin, Jiaqi Bai, Bing Wang, Jiaheng Liu, Xinnian Liang, Linzheng Chai, Liqun Yang, and Zhoujun Li. 2024 · 2024
Closest in time.