Fetching the paper…
Reading the bibliography…
Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity.
A spatial extension of cielab for digital color image reproduction
X. Zhang, B. A. Wandell, et al · 1996
Earlier work this paper cites.
The cma evolution strategy: a comparing review
N. Hansen · 2006
Earlier work this paper cites.
Learning photographic global tonal adjustment with a database of input / output image pairs
V. Bychkovsky, S. Paris, E. Chan, and F. Durand · 2011
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
P.-Y. Chen, H. Zhang, Y. Sharma, J. Yi, and C.-J. Hsieh · 2017
Earlier work this paper cites.
Exposure: A white-box photo post-processing framework
Y. Hu, H. He, C. Xu, B. Wang, and S. Lin · 2018
Earlier work this paper cites.
Automatic isp image quality tuning using nonlinear optimization
J. Nishimura, T. Gerasimow, R. Sushma, A. Sutic, C.-T. Wu, and G. Michael · 2018
Earlier work this paper cites.
Hyperparameter optimization in black-box image processing using differentiable proxies
E. Tseng, F. Yu, Y. Yang, F. Mannan, K. S. Arnaud, D. Nowrouzezahrai, J.-F. Lalonde, and F. Heide · 2019
Earlier work this paper cites.
Unpaired image enhancement featuring reinforcement-learning-controlled image editing software
S. Kosugi and T. Yamasaki · 2020
Earlier work this paper cites.
Hardware-in-the-loop end-to-end optimization of camera image processing pipelines
A. Mosleh, A. Sharma, E. Onzon, F. Mannan, N. Robidoux, and F. Heide · 2020
Earlier work this paper cites.
Learning image-adaptive 3d lookup tables for high performance photo enhancement in real-time
H. Zeng, J. Cai, L. Li, Z. Cao, and L. Zhang · 2020
Earlier work this paper cites.
Ppr10k: A large-scale portrait photo retouching dataset with human-region mask and group-level consistency
J. Liang, H. Zeng, M. Cui, X. Xie, and L. Zhang · 2021
Earlier work this paper cites.
Reconfigisp: Reconfigurable camera image processing pipeline
K. Yu, Z. Li, Y. Peng, C. C. Loy, and J. Gu · 2021
Earlier work this paper cites.
Harmonizer: Learning to perform white-box image and video harmonization
Z. Ke, C. Sun, L. Zhu, K. Xu, and R. W. Lau · 2022
Earlier work this paper cites.
Neural photo-finishing
E. Tseng, Y. Zhang, L. Jebe, X. Zhang, Z. Xia, Y. Fan, F. Heide, and J. Chen · 2022
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks, A. Holynski, and A. A. Efros · 2023
Earlier work this paper cites.
Metagpt: Meta programming for multi-agent collaborative framework
S. Hong, X. Zheng, J. Chen, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, et al · 2023
Earlier work this paper cites.
Viescore: Towards explainable metrics for conditional image synthesis evaluation
M. Ku, D. Jiang, C. Wei, X. Yue, and W. Chen · 2023
Earlier work this paper cites.
Langchain: Build context-aware reasoning applications
LangChain · 2023
Earlier work this paper cites.
Rsfnet: A white-box image retouching approach using region-specific color filters
W. Ouyang, Y. Dong, X. Kang, P. Ren, X. Xu, and X. Xie · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Earlier work this paper cites.
Openagents: An open platform for language agents in the wild
T. Xie, F. Zhou, Z. Cheng, P. Shi, L. Weng, Y. Liu, T. J. Hua, J. Zhao, Q. Liu, C. Liu, et al · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Earlier work this paper cites.
Magicbrush: A manually annotated dataset for instruction-guided image editing
K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su · 2023
Cited alongside, same era.
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu, et al · 2024
Cited alongside, same era.
Guiding instruction-based image editing via multimodal large language models
T.-J. Fu, W. Hu, X. Du, W. Y. Wang, Y. Yang, and Z. Gan · 2024
Cited alongside, same era.
Lightrag: Simple and fast retrieval-augmented generation
Z. Guo, L. Xia, Y. Yu, T. Ao, and C. Huang · 2024
Cited alongside, same era.
Audiogpt: Understanding and generating speech, music, sound, and talking head
R. Huang, M. Li, D. Yang, J. Shi, X. Chang, Z. Ye, Y. Wu, Z. Hong, J. Huang, J. Liu, et al · 2024
Cited alongside, same era.
Ultraedit: Instruction-based fine-grained image editing at scale
H. Zhao, X. S. Ma, L. Chen, S. Si, R. Wu, K. An, P. Yu, M. Zhang, Q. Li, and B. Chang · 2024
Later among the works it cites.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Y. Zheng, R. Zhang, J. Zhang, Y. Ye, Z. Luo, Z. Feng, and Y. Ma · 2024
Later among the works it cites.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.
R1-v: Reinforcing super generalization ability in vision-language models with less than $3
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci · 2025
Closest in time.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu, et al · 2024
Cited alongside, same era.
Hq-edit: A high-quality dataset for instruction-based image editing
M. Hui, S. Yang, B. Zhao, Y. Shi, H. Wang, P. Wang, Y. Zhou, and C. Xie · 2024
Cited alongside, same era.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Cited alongside, same era.
Preference optimization for reasoning with pseudo feedback
F. Jiao, G. Guo, X. Zhang, N. F. Chen, S. Joty, and F. Wei · 2024
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, et al · 2024
Cited alongside, same era.
Apigen: Automated pipeline for generating verifiable and diverse function-calling datasets
Z. Liu, T. Hoang, J. Zhang, M. Zhu, T. Lan, J. Tan, W. Yao, Z. Liu, Y. Feng, R. RN, et al · 2024
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Y. Hu, and S. Lin · 2025
Closest in time.
Search-r1: Training llms to reason and leverage search engines with reinforcement learning
B. Jin, H. Zeng, Z. Yue, D. Wang, H. Zamani, and J. Han · 2025
Closest in time.
Jarvisir: Elevating autonomous driving perception with intelligent image restoration
Y. Lin, Z. Lin, H. Chen, P. Pan, C. Li, S. Chen, W. Kairun, Y. Jin, W. Li, and X. Ding · 2025
Closest in time.
Step1x-edit: A practical framework for general image editing
S. Liu, Y. Han, P. Xing, F. Yin, R. Wang, W. Cheng, J. Liao, Y. Wang, H. Fu, C. Han, et al · 2025
Closest in time.
Visual-rft: Visual reinforcement fine-tuning
Z. Liu, Z. Sun, Y. Zang, X. Dong, Y. Cao, H. Duan, D. Lin, and J. Wang · 2025
Closest in time.
Ui-r1: Enhancing action prediction of gui agents by reinforcement learning
Z. Lu, Y. Chai, Y. Guo, X. Yin, L. Liu, H. Wang, G. Xiong, and H. Li · 2025
Closest in time.
Unitok: A unified tokenizer for visual generation and understanding
C. Ma, Y. Jiang, J. Wu, J. Yang, X. Yu, Z. Yuan, B. Peng, and X. Qi · 2025
Closest in time.
Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning
F. Meng, L. Du, Z. Liu, Z. Zhou, Q. Lu, D. Fu, B. Shi, W. Wang, J. He, K. Zhang, et al · 2025
Closest in time.
Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse
Z. Pan and H. Liu · 2025
Closest in time.
Toolrl: Reward is all tool learning needs
C. Qian, E. C. Acikgoz, Q. He, H. Wang, X. Chen, D. Hakkani-Tür, G. Tur, and H. Ji · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
H. Shen, P. Liu, J. Li, C. Fang, Y. Ma, J. Liao, Q. Shen, Z. Zhang, K. Zhao, Q. Zhang, et al · 2025
Closest in time.
Gui-r1: A generalist r1-style vision-language action model for gui agents
X. Xia and R. Luo · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Y. Yang, X. He, H. Pan, X. Jiang, Y. Deng, X. Yang, H. Lu, D. Yin, F. Rao, M. Zhu, et al · 2025
Closest in time.
Y. Zhao, F. Xue, S. Reed, L. Fan, Y. Zhu, J. Kautz, Z. Yu, P. Krähenbühl, and D.-A. Huang · 2025
Closest in time.
X. Zhuang, Y. Xie, Y. Deng, D. Yang, L. Liang, J. Ru, Y. Yin, and Y. Zou · 2025
Closest in time.