Fetching the paper…
Reading the bibliography…
The paper introduces the UniMER dataset, marking the first study on Mathematical Expression Recognition (MER) targeting complex real-world scenarios.
Binary codes capable of correcting deletions, insertions, and reversals
Levenshtein, V. I.; et al. 1966 · 1966
Earlier work this paper cites.
Syntax-directed recognition of hand-printed two-dimensional mathematics
Anderson, R. H. 1967 · 1967
Earlier work this paper cites.
Ambiguity and constraint in mathematical expression recognition
Miller, E.; and Viola, P. 1998 · 1998
Earlier work this paper cites.
Error detection, error correction and performance evaluation in on-line mathematical expression recognition
Chan, K.; and Yeung, D. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Infty: an integrated ocr system for mathematical documents
Suzuki, M.; Tamari, F.; Fukuda, R.; Uchida, S.; and Kanahori, T. 2003 · 2003
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
ICFHR 2014 competition on recognition of on-line handwritten mathematical expressions (CROHME 2014)
Mouchere, H.; Viard-Gaudin, C.; Zanibbi, R.; and Garain, U. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
ICFHR2016 CROHME: Competition on recognition of online handwritten mathematical expressions
Mouchère, H.; Viard-Gaudin, C.; Zanibbi, R.; and Garain, U. 2016 · 2016
Earlier work this paper cites.
Image-to-markup generation with coarse-to-fine attention
Deng, Y.; Kanervisto, A.; Ling, J.; and Rush, A. M. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Zhang, J.; Du, J.; Zhang, S.; Liu, D.; Hu, Y.; Hu, J.; Wei, S.; and Dai, L. 2017 · 2017
Earlier work this paper cites.
Pattern generation strategies for improving recognition of handwritten mathematical expressions
Le, A. D.; Indurkhya, B.; and Nakagawa, M. 2019 · 2019
Cited alongside, same era.
ICDAR 2019 CROHME+ TFD: Competition on recognition of handwritten mathematical expressions and typeset formula detection
Mahdavi, M.; Zanibbi, R.; Mouchere, H.; Viard-Gaudin, C.; and Garain, U. 2019 · 2019
Cited alongside, same era.
Improving attention-based handwritten mathematical expression recognition with scale augmentation and drop attention
Li, Z.; Jin, L.; Lai, S.; and Zhu, Y. 2020 · 2020
Cited alongside, same era.
Handwritten mathematical expression recognition via paired adversarial learning
Wu, J.-W.; Yin, F.; Zhang, Y.-M.; Zhang, X.-Y.; and Liu, C.-L. 2020 · 2020
Cited alongside, same era.
A tree-structured decoder for image-to-markup generation
Zhang, J.; Du, J.; Yang, Y.; Song, Y.-Z.; Wei, S.; and Dai, L. 2020 · 2020
Cited alongside, same era.
Nougat: Neural optical understanding for academic documents
Blecher, L.; Cucurull, G.; Scialom, T.; and Stojnic, R. 2023 · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2023 · 2023
Later among the works it cites.
Vary: Scaling up the vision vocabulary for large vision-language models
Wei, H.; Kong, L.; Chen, J.; Zhao, L.; Ge, Z.; Yang, J.; Sun, J.; Han, C.; and Zhang, X. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
pix2tex - LaTeX OCR
Blecher, L. 2022 · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chu, X.; Tian, Z.; Zhang, B.; Wang, X.; and Shen, C. 2021 · 2021
Cited alongside, same era.
Convit: Improving vision transformers with soft convolutional inductive biases
d’Ascoli, S.; Touvron, H.; Leavitt, M. L.; Morcos, A. S.; Biroli, G.; and Sagun, L. 2021 · 2021
Cited alongside, same era.
Handwritten mathematical expression recognition with bidirectionally trained transformer
Zhao, W.; Gao, L.; Yan, Z.; Peng, S.; Du, L.; and Zhang, Z. 2021 · 2021
Cited alongside, same era.
Handwritten mathematical expression recognition via attention aggregation based bi-directional mutual learning
Bian, X.; Qin, B.; Xin, X.; Li, J.; Su, X.; and Wang, Y. 2022 · 2022
Cited alongside, same era.
Cmt: Convolutional neural networks meet vision transformers
Guo, J.; Han, K.; Wu, H.; Tang, Y.; Chen, X.; Wang, Y.; and Xu, C. 2022 · 2022
Cited alongside, same era.
Ocr-free document understanding transformer
Kim, G.; Hong, T.; Yim, M.; Nam, J.; Park, J.; Yim, J.; Hwang, W.; Yun, S.; Han, D.; and Park, S. 2022 · 2022
Cited alongside, same era.
When counting meets HMER: counting-aware network for handwritten mathematical expression recognition
Li, B.; Yuan, Y.; Liang, D.; Liu, X.; Ji, Z.; Bai, J.; Liu, W.; and Bai, X. 2022 · 2022
Cited alongside, same era.
Chen, Z.; Wang, W.; Tian, H.; Ye, S.; Gao, Z.; Cui, E.; Tong, W.; Hu, K.; Luo, J.; Ma, Z.; et al. 2024 · 2024
Closest in time.
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Wang, B.; Ouyang, L.; Wei, X.; Zhang, S.; Duan, H.; Cao, M.; et al. 2024 · 2024
Closest in time.
Opendatalab: Empowering general artificial intelligence with open datasets
He, C.; Li, W.; Jin, Z.; Xu, C.; Wang, B.; and Lin, D. 2024 · 2024
Closest in time.
MMSci: A Multimodal Multi-Discipline Dataset for PhD-Level Scientific Comprehension
Li, Z.; Yang, X.; Choi, K.; Zhu, W.; Hsieh, R.; Kim, H.; Lim, J. H.; Ji, S.; Lee, B.; Yan, X.; et al. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Vigc: Visual instruction generation and correction
Wang, B.; Wu, F.; Han, X.; Peng, J.; Zhong, H.; Zhang, P.; Dong, X.; Li, W.; Li, W.; Wang, J.; et al. 2024 · 2024
Closest in time.
Xia, R.; Mao, S.; Yan, X.; Zhou, H.; Zhang, B.; Peng, H.; Pi, J.; Fu, D.; Wu, W.; Ye, H.; et al. 2024 · 2024
Closest in time.
Zhang, P.; Dong, X.; Zang, Y.; Cao, Y.; Qian, R.; Chen, L.; Guo, Q.; Duan, H.; Wang, B.; Ouyang, L.; et al. 2024 · 2024
Closest in time.