Fetching the paper…
Reading the bibliography…
We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities.
FCOS: Fully Convolutional One-Stage Object Detection
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. 2019 · 1904
Earlier work this paper cites.
Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Minghui Liao, Pengyuan Lyu, Minghang He, Cong Yao, Wenhao Wu, and Xiang Bai. 2019 · 1908
Earlier work this paper cites.
RandAugment: Practical data augmentation with no separate search
Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2019 · 1909
Earlier work this paper cites.
VarifocalNet: An IoU-aware Dense Object Detector
Haoyang Zhang, Ying Wang, Feras Dayoub, and Niko Sünderhauf. 2021 · 2008
Earlier work this paper cites.
PP-OCR: A Practical Ultra Lightweight OCR System
Yuning Du, Chenxia Li, Ruoyu Guo, Xiaoting Yin, Weiwei Liu, Jun Zhou, Yifan Bai, Zilin Yu, Yehua Yang, Qingqing Dang, and Haoshuang Wang. 2020 · 2009
Earlier work this paper cites.
Word Spotting in the Wild. In Proceedings of the 11th European Conference on Computer Vision: Part I (Heraklion, Crete, Greece) (ECCV’10) . Springer-Verlag, Berlin, Heidelberg, 591–604
Kai Wang and Serge Belongie. 2010 · 2010
Earlier work this paper cites.
MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. 2021 · 2012
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Reading Text in the Wild with Convolutional Neural Networks
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
Baoguang Shi, Xiang Bai, and Cong Yao. 2017 · 2016
Earlier work this paper cites.
Robust Scene Text Recognition with Automatic Rectification. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 4168–4176
Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2016 · 2016
Earlier work this paper cites.
Rosetta: Large Scale System for Text Detection and Recognition in Images. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (London, United Kingdom) (KDD ’18) . Association for Computing Machinery, New York, NY, USA, 71–79
Fedor Borisyuk, Albert Gordo, and Viswanath Sivakumar. 2018 · 2018
Earlier work this paper cites.
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2018 · 2018
Cited alongside, same era.
Path Aggregation Network for Instance Segmentation
Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. 2018 · 2018
Cited alongside, same era.
Arbitrary-Oriented Scene Text Detection via Rotation Proposals
Jianqi Ma, Weiyuan Shao, Hao Ye, Li Wang, Hong Wang, Yingbin Zheng, and Xiangyang Xue. 2018 · 2018
Cited alongside, same era.
ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . 1516–1520
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar. 2019 · 2019
Cited alongside, same era.
FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions
Alvin Wan, Xiaoliang Dai, Peizhao Zhang, Zijian He, Yuandong Tian, Saining Xie, Bichen Wu, Matthew Yu, Tao Xu, Kan Chen, Péter Vajda, and Joseph Gonzalez. 2020 · 2020
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022 · 2022
Later among the works it cites.
Domain Prompts: Towards memory and compute efficient domain adaptation of ASR systems. In Proc. Interspeech 2022 . 684–688
Saket Dingliwal, Ashish Shenoy, Sravan Bodapati, Ankur Gandhe, Ravi Teja Gadde, and Katrin Kirchhoff. 2022 · 2022
Later among the works it cites.
Towards End-to-End Unified Scene Text Detection and Layout Analysis. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 1039–1049
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. 2022b · 2022
Later among the works it cites.
Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Remember the Context! ASR Slot Error Correction Through Memorization. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . 236–243
Dhanush Bekal, Ashish Shenoy, Monica Sunkara, Sravan Bodapati, and Katrin Kirchhoff. 2021 · 2021
Cited alongside, same era.
Efficient domain adaptation of language models in ASR systems using Prompt-tuning
Saket Dingliwal, Ashish Shenoy, Sravan Bodapati, Ankur Gandhe, Ravi Teja Gadde, and Katrin Kirchhoff. 2021 · 2021
Cited alongside, same era.
DocVQA: A Dataset for VQA on Document Images. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV) . 2199–2208
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021 · 2021
Cited alongside, same era.
STRIDE: Scene Text Recognition In-Device. In 2021 International Joint Conference on Neural Networks (IJCNN) . 1–8
Rachit S Munjal, Arun D Prabhu, Nikhil Arora, Sukumar Moharana, and Gopi Ramena. 2021 · 2021
Cited alongside, same era.
A White Paper on Neural Network Quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort. 2021 · 2021
Cited alongside, same era.
ASR Adaptation for E-commerce Chatbots using Cross-Utterance Context and Multi-Task Language Modeling. In Proceedings of the 4th Workshop on e-Commerce and NLP , Shervin Malmasi, Surya Kallumadi, Nicola Ueffing, Oleg Rokhlenko, Eugene Agichtein, and Ido Guy (Eds.). Association for Computational Linguistics, Online, 18–25
Ashish Shenoy, Sravan Bodapati, and Katrin Kirchhoff. 2021a · 2021
Cited alongside, same era.
Adapting Long Context NLM for ASR Rescoring in Conversational Agents. In Proc. Interspeech 2021 . 3246–3250
Ashish Shenoy, Sravan Bodapati, Monica Sunkara, Srikanth Ronanki, and Katrin Kirchhoff. 2021b · 2021
Cited alongside, same era.
Xue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming, Wentao Wang, Qi Tian, and Junchi Yan. 2022 · 2022
Later among the works it cites.
OpenAI (2023). 2023 · 2023
Later among the works it cites.
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. 2023 · 2023
Later among the works it cites.
Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang. 2023 · 2023
Later among the works it cites.
BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions
Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu. 2023 · 2023
Later among the works it cites.
Exploring OCR Capabilities of GPT-4V(ision) : A Quantitative and In-depth Evaluation
Yongxin Shi, Dezhi Peng, Wenhui Liao, Zening Lin, Xinhong Chen, Chongyu Liu, Yuyi Zhang, and Lianwen Jin. 2023 · 2023
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, and et al. 2023 · 2023
Later among the works it cites.
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, Qin Jin, Liang He, Xin Alex Lin, and Fei Huang. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.