Fetching the paper…

Image2Struct: Benchmarking Structure Extraction for Vision-Language Models · Around