Fetching the paper…
Reading the bibliography…
This paper introduces AnyTrans, an all-encompassing framework for the task-Translate AnyText in the Image (TATI), which includes multilingual text translation and text fusion within images.
Pp-ocr: A practical ultra lightweight ocr system
Yuning Du, Chenxia Li, Ruoyu Guo, Xiaoting Yin, Weiwei Liu, Jun Zhou, Yifan Bai, Zilin Yu, Yehua Yang, Qingqing Dang, et al. 2020 · 2009
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero. 2014 · 2014
Earlier work this paper cites.
Multi-task learning for multiple language translation
Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, and Haifeng Wang. 2015 · 2015
Earlier work this paper cites.
Multimodal attention for neural machine translation
Ozan Caglayan, Loïc Barrault, and Fethi Bougares. 2016 · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Attention-based multimodal neural machine translation
Po-Yao Huang, Frederick Liu, Sz-Rung Shiang, Jean Oh, and Chris Dyer. 2016 · 2016
Earlier work this paper cites.
Doubly-attentive decoder for multi-modal neural machine translation
Iacer Calixto, Qun Li, and Nick Campbell. 2017 · 2017
Earlier work this paper cites.
A teacher-student framework for zero-resource neural machine translation
Yun Chen, Yang Liu, Yong Cheng, and Victor O.K. Li. 2017 · 2017
Earlier work this paper cites.
Imagination improves multimodal translation
Desmond Elliott and Ákos Kádár. 2017 · 2017
Earlier work this paper cites.
Gan(generative adversarial nets)
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Attention strategies for multi-source sequence-to-sequence learning
Jindřich Libovický and Jindřich Helcl. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao. 2017 · 2017
Earlier work this paper cites.
East: An efficient and accurate scene text detector
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang. 2017 · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. 2017 · 2017
Earlier work this paper cites.
Multi-content gan for few-shot font style transfer
Samaneh Azadi, Matthew Fisher, Vladimir Kim, Zhaowen Wang, Eli Shechtman, and Trevor Darrell. 2018 · 2018
Earlier work this paper cites.
Input combination strategies for multi-source transformer decoder
Jindřich Libovický, Jindřich Helcl, and David Mareček. 2018 · 2018
Earlier work this paper cites.
Multi-oriented scene text detection via corner localization and region segmentation
Pengyuan Lyu, Cong Yao, Wenhao Wu, Shuicheng Yan, and Xiang Bai. 2018 · 2018
Earlier work this paper cites.
Arbitrary-oriented scene text detection via rotation proposals
Jianqi Ma, Weiyuan Shao, Hao Ye, Li Wang, Hong Wang, Yingbin Zheng, and Xiangyang Xue. 2018 · 2018
Earlier work this paper cites.
Rapid adaptation of neural machine translation to new languages
Graham Neubig and Junjie Hu. 2018 · 2018
Cited alongside, same era.
Joint Training for Pivot-Based Neural Machine Translation , page 41–54
Yong Cheng. 2019 · 2019
Cited alongside, same era.
Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition – rrc-mlt-2019
Nibal Nayef, Cheng-Lin Liu, Jean-Marc Ogier, Yash Patel, Michal Busta, ChowdhuryPinaki Nath, Karatzas Dimosthenis, Wafa Khlif, Jiri Matas, Umapada Pal, and Burie Jean-Christophe. 2019 · 2019
Cited alongside, same era.
Multimodal machine translation through visuals and speech
Umut Sulubacak, Ozan Caglayan, Stig-Arne Grönroos, Aku Rouhe, Desmond Elliott, Lucia Specia, and Jörg Tiedemann. 2019 · 2019
Cited alongside, same era.
Editing text in the wild
Liang Wu, Chengquan Zhang, Jiaming Liu, Junyu Han, Jingtuo Liu, Errui Ding, and Xiang Bai. 2019 · 2019
Cited alongside, same era.
Real-time scene text detection with differentiable binarization
Diffedit: Diffusion-based semantic image editing with mask guidance
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. 2022 · 2022
Later among the works it cites.
Improving end-to-end text image translation from the auxiliary text translation task
Cong Ma, Yaping Zhang, Mei Tu, Xu Han, Linghui Wu, Yang Zhao, and Yu Zhou. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. 2022 · 2022
Later among the works it cites.
Palette: Image-to-image diffusion models
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. 2022 · 2022
Later among the works it cites.
Prompting palm for translation: Assessing strategies and performance
David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, and George Foster. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen, and Xiang Bai. 2020 · 2020
Cited alongside, same era.
Towards end-to-end in-image neural machine translation
Elman Mansimov, Mitchell Stern, Mia Chen, Orhan Firat, Jakob Uszkoreit, and Puneet Jain. 2020 · 2020
Cited alongside, same era.
Comet: A neural framework for mt evaluation
Ricardo Rei, Craig Stewart, AnaC Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Cited alongside, same era.
Most: A multi-oriented scene text detector with localization refinement
Minghang He, Minghui Liao, Zhibo Yang, Humen Zhong, Jun Tang, Wenqing Cheng, Cong Yao, Yongpan Wang, and Xiang Bai. 2021 · 2021
Cited alongside, same era.
Trocr: Transformer-based optical character recognition with pre-trained models
Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Zhang Cha, Zhoujun Li, and Furu Wei. 2021 · 2021
Cited alongside, same era.
Multi-modal neural machine translation with deep semantic interactions
Jinsong Su, Jinchang Chen, Hui Jiang, Chulun Zhou, Huan Lin, Yubin Ge, Qingqiang Wu, and Yongxuan Lai. 2021 · 2021
Cited alongside, same era.
Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. 2022 · 2022
Later among the works it cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Textdiffuser-2: Unleashing the power of language models for text rendering
Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei. 2023 · 2023
Later among the works it cites.
Wordart designer: User-driven artistic typography synthesis using large language models
Jun-Yan He, Zhi-Qi Cheng, Chenyang Li, Jingdong Sun, Wangmeng Xiang, Xianhui Lin, Xiaoyang Kang, Zengke Jin, Yusen Hu, Bin Luo, et al. 2023 · 2023
Later among the works it cites.
Exploring better text image translation with multimodal codebook
Zhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang, Jian Luan, Bin Wang, Degen Huang, and Jinsong Su. 2023 · 2023
Later among the works it cites.
Strokenet: Stroke assisted and hierarchical graph reasoning networks
Lei Li, Kai Fan, and Chun Yuan. 2023 · 2023
Later among the works it cites.
On the hidden mystery of ocr in large multimodal models
Yuliang Liu, Zhang Li, Hongliang Li, Wenwen Yu, Mingxin Huang, Dezhi Peng, Mingyu Liu, Mingrui Chen, Chunyuan Li, Lianwen Jin, et al. 2023 · 2023
Later among the works it cites.
Glyphdraw: Learning to draw chinese characters in image synthesis models coherently
Jian Ma, Mingjun Zhao, Chen Chen, Ruichen Wang, Di Niu, Haonan Lu, and Xiaodong Lin. 2023 · 2023
Later among the works it cites.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. 2023 · 2023
Later among the works it cites.
Anytext: Multilingual visual text generation and editing
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, and Xuansong Xie. 2023 · 2023
Later among the works it cites.
Chinese text recognition with a pre-trained clip-like model through image-ids aligning
Haiyang Yu, Xiaocong Wang, Bin Li, and Xiangyang Xue. 2023 · 2023
Later among the works it cites.
Tim: Teaching large language models to translate with comparison
Jiali Zeng, Fandong Meng, Yongjing Yin, and Jie Zhou. 2023 · 2023
Later among the works it cites.
Textdiffuser: Diffusion models as text painters
Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei. 2024 · 2024
Closest in time.
Towards boosting many-to-many multilingual machine translation with large language models
Pengzhi Gao, Zhongjun He, Hua Wu, and Haifeng Wang. 2024 · 2024
Closest in time.