Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs.
Musical scales and the generalized circle of fifths
John Clough and Gerald Myerson · 1986
Earlier work this paper cites.
Optical music recognition using projections
Ichiro Fujinaga · 1988
Earlier work this paper cites.
MusiXTEX. using TEX to write polyphonic or instrumental music
Daniel Taupin, Ross Mitchell, and Andreas Egler · 1993
Earlier work this paper cites.
Harmonic experience: Tonal harmony from its natural origins to its modern expression
William Allaudin Mathieu · 1997
Earlier work this paper cites.
The challenge of optical music recognition
David Bainbridge and Tim Bell · 2001
Earlier work this paper cites.
Music information processing using the humdrum toolkit: Concepts, examples, and lessons
David Huron · 2002
Earlier work this paper cites.
Reading sheet music facilitates sensorimotor mu-desynchronization in musicians
Lawrence Paul Behmer Jr and Kelly J Jantzen · 2011
Earlier work this paper cites.
Optical music recognition: state-of-the-art and open issues
Ana Rebelo, Ichiro Fujinaga, Filipe Paszkiewicz, Andre RS Marcal, Carlos Guedes, and Jaime S Cardoso · 2012
Earlier work this paper cites.
CVC-MUSCIMA: A ground-truth of handwritten music score images for writer identification and staff removal
Alicia Fornés, Anjan Dutta, Albert Gordo, and Josep Lladós · 2012
Earlier work this paper cites.
The basics of reading music
Kevin Meixner · 2015
Earlier work this paper cites.
Optical music recognition with convolutional sequence-to-sequence models
Eelco van der Wel and Karen Ullrich · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Deepscores-a dataset for segmentation, detection and classification of tiny objects
Lukas Tuggener, Ismail Elezi, Jurgen Schmidhuber, Marcello Pelillo, and Thilo Stadelmann · 2018
Earlier work this paper cites.
End-to-end neural optical music recognition of monophonic scores
Jorge Calvo-Zaragoza and David Rizo · 2018
Earlier work this paper cites.
Optical music recognition: State of the art and major challenges
Elona Shatri and György Fazekas · 2020
Earlier work this paper cites.
Understanding optical music recognition
Jorge Calvo-Zaragoza, Jan Hajič Jr, and Alexander Pacha · 2020
Earlier work this paper cites.
Pp-ocr: A practical ultra lightweight ocr system
Yuning Du, Chenxia Li, Ruoyu Guo, Xiaoting Yin, Weiwei Liu, Jun Zhou, Yifan Bai, Zilin Yu, Yehua Yang, Qingqing Dang, et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Cited alongside, same era.
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner · 2021
Cited alongside, same era.
Doremi: First glance at a universal omr dataset
Elona Shatri and György Fazekas · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Breezewhite/oemer: v0.1.7, October 2023
Yoyo, Christian Liebhardt, and Sayooj Samuel · 2023
Cited alongside, same era.
A unified representation framework for the evaluation of optical music recognition systems
Pau Torras, Sanket Biswas, and Alicia Fornés · 2024
Later among the works it cites.
Practical end-to-end optical music recognition for pianoform music
Jiří Mayer, Milan Straka, Jan Hajič, and Pavel Pecina · 2024
Later among the works it cites.
Optical music recognition in manuscripts from the ricordi archive
Federico Simonetta, Rishav Mondal, Luca Andrea Ludovico, and Stavros Ntalampiras · 2024
Later among the works it cites.
Sheet music transformer: End-to-end optical music recognition beyond monophonic transcription, 2024
Antonio Ríos-Vila, Jorge Calvo-Zaragoza, and Thierry Paquet · 2024
Later among the works it cites.
Sheet music transformer++: End-to-end full-page optical music recognition for pianoform sheet music
Antonio Rıos-Vila, Jorge Calvo-Zaragoza, David Rizo, and Thierry Paquet · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tromr:transformer-based polyphonic optical music recognition
Yixuan Li, Huaping Liu, Qiang Jin, Miaomiao Cai, and Peng Li · 2023
Cited alongside, same era.
Musicagent: An ai agent for music understanding and generation with large language models
Dingyao Yu, Kaitao Song, Peiling Lu, Tianyu He, Xu Tan, Wei Ye, Shikun Zhang, and Jiang Bian · 2023
Cited alongside, same era.
Layoutgpt: Compositional visual planning and generation with large language models
Weixi Feng, Wanrong Zhu, Tsu-jui Fu, Varun Jampani, Arjun Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang · 2023
Cited alongside, same era.
Acrobat AI Assistant, 2024
Adobe Inc · 2024
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Cited alongside, same era.
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, et al · 2024
Cited alongside, same era.
Natural language understanding and inference with mllm in visual question answering: A survey
Jiayi Kuang, Ying Shen, Jingyou Xie, Haohao Luo, Zhe Xu, Ronghao Li, Yinghui Li, Xianfeng Cheng, Xika Lin, and Yu Han · 2024
Cited alongside, same era.
Chatmusician: Understanding and generating music intrinsically with llm
Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, et al · 2024
Later among the works it cites.
Paligemma 2: A family of versatile vlms for transfer
Andreas Steiner, André Susano Pinto, Michael Tschannen, Daniel Keysers, Xiao Wang, Yonatan Bitton, Alexey Gritsenko, Matthias Minderer, Anthony Sherbondy, Shangbang Long, et al · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al · 2024
Later among the works it cites.
Trins: Towards multimodal language models that can read
Ruiyi Zhang, Yanzhe Zhang, Jian Chen, Yufan Zhou, Jiuxiang Gu, Changyou Chen, and Tong Sun · 2024
Later among the works it cites.
Llava-read: Enhancing reading ability of multimodal language models
Ruiyi Zhang, Yufan Zhou, Jian Chen, Jiuxiang Gu, Changyou Chen, and Tong Sun · 2024
Later among the works it cites.
MMR: Evaluating reading ability of large multimodal models
Jian Chen, Ruiyi Zhang, Yufan Zhou, Ryan Rossi, Jiuxiang Gu, and Changyou Chen · 2024
Later among the works it cites.
Textlap: Customizing language models for text-to-layout planning
Jian Chen, Ruiyi Zhang, Yufan Zhou, Jennifer Healey, Jiuxiang Gu, Zhiqiang Xu, and Changyou Chen · 2024
Later among the works it cites.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, et al · 2024
Later among the works it cites.
Ocrbench: on the hidden mystery of ocr in large multimodal models
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai · 2024
Later among the works it cites.