Fetching the paper…
Reading the bibliography…
Document visual question answering (DocVQA) pipelines that answer questions from documents have broad applications.
Video google: a text retrieval approach to object matching in videos
Sivic and Zisserman · 2003
Earlier work this paper cites.
Inverted files for text search engines
Justin Zobel and Alistair Moffat · 2006
Earlier work this paper cites.
An overview of the tesseract ocr engine
Ray Smith · 2007
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jégou, Matthijs Douze, and Cordelia Schmid · 2011
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
pdf2image, 2017
Edouard Belval · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chana, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Scene Text Visual Question Answering
Ali Furkan Biten, Andres Mafla, Lluis Gomez, Valveny C V Jawahar, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
pdfminer.six, 2019
pdfminer · 2019
Earlier work this paper cites.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, O Barlas, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Earlier work this paper cites.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab and Matei Zaharia · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen Tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr)
Jamshed Memon, Maira Sami, Rizwan Ahmed Khan, and Mueen Uddin · 2020
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 2020
Earlier work this paper cites.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
Gautier Izacard and Edouard Grave · 2021
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2021
Earlier work this paper cites.
Playwright for python, 2021
Microsoft · 2021
Earlier work this paper cites.
Multimodalqa: Complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant · 2021
Earlier work this paper cites.
Retrieving and reading: A comprehensive survey on open-domain question answering
Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Murag: Multimodal retrieval-augmented generator for open question answering over images and text
Wenhu Chen, Hexiang Hu, Xi Chen, Pat Verga, and William W Cohen · 2022
Cited alongside, same era.
The PyPDF2 library, version 2, 2022
Mathieu Fenniak and PyPDF2 Contributors · 2022
Cited alongside, same era.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Cited alongside, same era.
Ocr-free document understanding transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park · 2022
Cited alongside, same era.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, and C.V. Jawahar · 2022
Cited alongside, same era.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Mingshi Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, Qin Jin, Liang He, Xin Lin, and Feiyan Huang · 2023
Later among the works it cites.
ScreenAI: A Vision-Language Model for UI and Infographics Understanding, 2024
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, and Abhanshu Sharma · 2024
Closest in time.
LongAlign: A recipe for long context alignment of large language models
Yushi Bai, Xin Lv, Jiajie Zhang, Yuze He, Ji Qi, Lei Hou, Jie Tang, Yuxiao Dong, and Juanzi Li · 2024
Closest in time.
PaliGemma: A versatile 3B VLM for transfer, 2024
Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia · 2022
Cited alongside, same era.
A-okvqa: A benchmark for visual question answering using world knowledge, 2022
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi · 2022
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection, 2023
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Qwen-VL: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Pdfvqa: A new dataset for real-world vqa on pdf documents
Yihao Ding, Siwen Luo, Hyunsuk Chung, and Soyeon Caren Han · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang · 2023
Cited alongside, same era.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Gram: Global reasoning for multi-page vqa
Tsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts, Roy Ganz, Elad Ben Avraham, Aviad Aberdam, Shahar Tsiper, and Ron Litman · 2024
Closest in time.
Arctic-TILT. Business Document Understanding at Sub-Billion Scale, 2024
Łukasz Borchmann, Michał Pietruszka, Wojciech Jaśkowski, Dawid Jurkiewicz, Piotr Halama, Paweł Józiak, Łukasz Garncarek, Paweł Liskowski, Karolina Szyndler, Andrzej Gretkowski, Julita Ołtusek, Gabriela Nowakowska, Artur Zawłocki, Łukasz Duhr, Paweł Dyda, and Michał Turski · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al · 2024
Closest in time.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Closest in time.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, et al · 2024
Closest in time.
The faiss library, 2024
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Closest in time.
ColPali: Efficient Document Retrieval with Vision Language Models, 2024
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Gemini Team · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding, 2024
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou · 2024
Closest in time.
The llama 3 herd of models, 2024
Llama Team · 2024
Closest in time.
DeepSeek-VL: towards real-world vision-language understanding
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Yaofeng Sun, et al · 2024
Closest in time.
Mmlongbench-doc: Benchmarking long-context document understanding with visualizations
Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, et al · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
The chromium projects, 2024
The Chromium Project Authors · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin · 2024
Closest in time.
LLaVA-UHD: An LMM perceiving any aspect ratio and high-resolution images
Ruyi Xu, Yuan Yao, Zonghao Guo, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, and Gao Huang · 2024
Closest in time.
RLAIF-V: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Tianyu Yu, Haoye Zhang, Yuan Yao, Yunkai Dang, Da Chen, Xiaoman Lu, Ganqu Cui, Taiwen He, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2024
Closest in time.