Fetching the paper…
Reading the bibliography…
The adoption of tablets with touchscreens and styluses is increasing, and a key feature is converting handwriting to text, enabling search, indexing, and AI assistance.
Hmm based online handwriting recognition
Hu, J., Brown, M., and Turin, W · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The state of the art in japanese online handwriting recognition compared to techniques in western handwriting recognition
Jaeger, S., Liu, C.-L., and Nakagawa, M · 2003
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
A novel connectionist system for unconstrained handwriting recognition
Graves, A., Liwicki, M., Fernández, S., Bertolami, R., Bunke, H., and Schmidhuber, J · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Hannun, A. Y., Case, C., Casper, J., Catanzaro, B., Diamos, G. F., Elsen, E., Prenger, R. J., Satheesh, S., Sengupta, S., Coates, A., and Ng, A · 2014
Earlier work this paper cites.
State-of-the-art speech recognition with sequence-to-sequence models
Chiu, C.-C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, K., Jaitly, N., Li, B., Chorowski, J., and Bacchiani, M · 2017
Earlier work this paper cites.
A comparison of sequence-to-sequence models for speech recognition
Prabhavalkar, R., Rao, K., Sainath, T. N., Li, B., Johnson, L. M., and Jaitly, N · 2017
Earlier work this paper cites.
As we may ink?: Learning from everyday analog pen use to improve digital ink experiences
Riche, Y., Henry Riche, N., Hinckley, K., Panabaker, S., Fuelling, S., and Williams, S · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Deepwriting: Making digital ink editable via deep generative modeling
Aksan, E., Pece, F., and Hilliges, O · 2018
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Icfhr 2018–competition on vietnamese online handwritten text recognition using hands-vnondb (vohtr2018)
Nguyen, H. T., Nguyen, C. T., and Nakagawa, M · 2018
Earlier work this paper cites.
Deep context: End-to-end contextual speech recognition
Pundak, G., Sainath, T. N., Prabhavalkar, R., Kannan, A., and Zhao, D · 2018
Earlier work this paper cites.
End to end recognition system for recognizing offline unconstrained vietnamese handwriting
Le, A. D., Nguyen, H. T., and Nakagawa, M · 2019
Cited alongside, same era.
Evaluating sequence-to-sequence models for handwritten text recognition, 2019
Michael, J., Labahn, R., Grüning, T., and Zöllner, J · 2019
Cited alongside, same era.
Multi-modal attention network for handwritten mathematical expression recognition
Wang, J., Du, J., Zhang, J., and Wang, Z.-R · 2019
Cited alongside, same era.
Unified vision-language pre-training for image captioning and vqa, 2019
Zhou, L., Palangi, H., Zhang, L., Hu, H., Corso, J. J., and Gao, J · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Palm 2 technical report, 2023
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Díaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y · 2023
Later among the works it cites.
Introducing our multimodal models, 2023
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., and Taşırlar, S · 2023
Later among the works it cites.
Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Dehghani, M., Mustafa, B., Djolonga, J., Heek, J., Minderer, M., Caron, M., Steiner, A., Puigcerver, J., Geirhos, R., Alabdulmohsin, I., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast multi-language lstm-based online handwriting recognition
Carbune, V., Gonnet, P., Deselaers, T., Rowley, H., Daryin, A., Calvo, M., Wang, L.-L., Keysers, D., Feuz, S., and Gervais, P · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Sketchformer: Transformer-based representation for sketched structure, 2020
Ribeiro, L. S. F., Bui, T., Collomosse, J., and Ponti, M · 2020
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer, 2021
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Cited alongside, same era.
Survey on handwritten recognition
Al Sayed, I., Mawlod, A., Alnajjar, A., and Gheni, H · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning, 2022
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Transformer-based models for arabic online handwriting recognition
Alwajih, F., Badr, E., and Abdou, S · 2022
Cited alongside, same era.
Later among the works it cites.
Msdoctr-lite: A lite transformer for full page multi-script handwriting recognition
Dhiaf, M., Rouhou, A. C., Kessentini, Y., and Salem, S. B · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model, 2023
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P · 2023
Later among the works it cites.
Detect handwriting in image, 2023
Google Cloud · 2023
Later among the works it cites.
Speech translation with large language models: An industrial practice, 2023
Huang, Z., Ye, R., Ko, T., Dong, Q., Cheng, S., Wang, M., and Li, H · 2023
Later among the works it cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Lee, K., Joshi, M., Turc, I. R., Hu, H., Liu, F., Eisenschlos, J. M., Khandelwal, U., Shaw, P., Chang, M.-W., and Toutanova, K · 2023
Later among the works it cites.
Visual instruction tuning, 2023
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Later among the works it cites.
Goat: Fine-tuned llama outperforms gpt-4 on arithmetic tasks, 2023
Liu, T. and Low, B. K. H · 2023
Later among the works it cites.
Painter: Teaching auto-regressive language models to draw sketches, 2023
Pourreza, R., Bhattacharyya, A., Panchal, S., Lee, M., Madan, P., and Memisevic, R · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen, 2023
Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., de Chaumont Quitry, F., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., Muckenhirn, H., Padfield, D., Qin, J., Rozenberg, D., Sainath, T., Schalkwyk, J., Sharifi, M., Ramanovich, M. T., Tagliasacchi, M., Tudor, A., Velimirović, M., Vincent, D., Yu, J., Wang, Y., Zayats, V., Zeghidour, N., Zhang, Y., Zhang, Z., Zilka, L., and Frank, C · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Icdar 2023 crohme: Competition on recognition of handwritten mathematical expressions
Xie, Y., Mouchère, H., Simistira Liwicki, F., Rakesh, S., Saini, R., Nakagawa, M., Nguyen, C. T., and Truong, T.-N · 2023
Later among the works it cites.
https://www.kaggle.com/competitions/quickdraw-doodle-recognition , 2023
Kaggle Quick, Draw! competition · 2024
Closest in time.
Mathwriting dataset, 2024
Gervais, P., Fadeeva, A., and Maksai, A · 2024
Closest in time.