Fetching the paper…
Reading the bibliography…
Computer vision often treats human perception as homogeneous: an implicit assumption that visual stimuli are perceived similarly by everyone.
Experimentelle studien uber das sehen von bewegung
Max Wertheimer · 1912
Earlier work this paper cites.
Philosophical Investigations
Ludwig Wittgenstein · 1953
Earlier work this paper cites.
Gestalt psychology
Wolfgang Köhler · 1967
Earlier work this paper cites.
The influence of culture on visual perception
Marshall H. Segall, Donald T. Campbell, and Melville J. Herskovits · 1967
Earlier work this paper cites.
Developmental influences on geometric illusion susceptibility among hong kong chinese children
John L. M. Dawson, Brian M. Young, and Peter P. C. Choi · 1973
Earlier work this paper cites.
Pictorial depth perception: a developmental study
Gustav Jahoda and Harry McGurk · 1974
Earlier work this paper cites.
The müller-lyer illusion among navajos
Darhl M. Pedersen and Jay Wheeler · 1983
Earlier work this paper cites.
Fire, Women, and Dangerous Things
George Lakoff · 1987
Earlier work this paper cites.
Force dynamics in language and cognition
Leonard Talmy · 1988
Earlier work this paper cites.
Universals of construal
Ronald W Langacker · 1993
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller · 1995
Earlier work this paper cites.
Construal operations in linguistics and artificial intelligence
William Croft and Esther J Wood · 2000
Earlier work this paper cites.
Attending holistically versus analytically: comparing the context sensitivity of japanese and americans
Takahiko Masuda and Richard E. Nisbett · 2001
Earlier work this paper cites.
Sex, syntax, and semantics
Lera Boroditsky, Lauren A Schmidt, and Webb Phillips · 2003
Earlier work this paper cites.
Culture and point of view
Richard E. Nisbett and Takahiko Masuda · 2003
Earlier work this paper cites.
Linguistic relativity
Lera Boroditsky · 2006
Earlier work this paper cites.
Usage-based linguistics
Michael Tomasello · 2006
Earlier work this paper cites.
Measuring emotional expression with the linguistic inquiry and word count
Jeffrey H. Kahn, Renée M. Tobin, Audra E Massey, and Jemima Asabea Anderson · 2007
Earlier work this paper cites.
Interesting objects are visually salient
Lior Elazary and Laurent Itti · 2008
Earlier work this paper cites.
Linguistic dimensions of psychopathology: A quantitative analysis
Doerte U Junghaenel, Joshua M Smyth, and Laura Santner · 2008
Earlier work this paper cites.
Metaphors we live by
George Lakoff and Mark Johnson · 2008
Earlier work this paper cites.
Parsing three German treebanks: Lexicalized and unlexicalized baselines
Anna Rafferty and Christopher D. Manning · 2008
Earlier work this paper cites.
Some objects are more equal than others: Measuring and predicting importance
Merrielle Spain and Pietro Perona · 2008
Earlier work this paper cites.
Discriminative reordering with Chinese grammatical relations features
Pi-Chuan Chang, Huihsin Tseng, Dan Jurafsky, and Christopher D. Manning · 2009
Earlier work this paper cites.
The psychological meaning of words: Liwc and computerized text analysis methods
Yla R. Tausczik and James W. Pennebaker · 2010
Earlier work this paper cites.
Understanding and predicting importance in images
Alexander C. Berg, Tamara L. Berg, Hal Daumé, Jesse Dodge, Amit Goyal, Xufeng Han, Alyssa C. Mensch, Margaret Mitchell, Aneesh Sood, Karl Stratos, and Kota Yamaguchi · 2012
Earlier work this paper cites.
Cartesian Meditations: An Introduction to Phenomenology
Edmund Husserl · 2012
Earlier work this paper cites.
The expressive power of word embeddings, 2013
Yanqing Chen, Bryan Perozzi, Rami Al-Rfou, and Steven Skiena · 2013
Earlier work this paper cites.
Studying relationships between human gaze, description, and computer vision
Kiwon Yun, Yifan Peng, Dimitris Samaras, Gregory J. Zelinsky, and Tamara L. Berg · 2013
Earlier work this paper cites.
Concreteness ratings for 40 thousand generally known english word lemmas
Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Motion encoding in russian and english: moving beyond talmy’s typology
A. Pavlenko and M. Volynsky · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Multimodal pivots for image caption translation
Julian Hitschler, Shigehiko Schamoni, and Stefan Riezler · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Image captioning in the wild: How people caption images on flickr
Philipp Blandfort, Tushar Karayil, Damian Borth, and Andreas R. Dengel · 2017
Earlier work this paper cites.
STAIR captions: Constructing a large-scale Japanese image caption dataset
Yuya Yoshikawa et al · 2017
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Alexander J. Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason Alan Fries, Sen Wu, and Christopher Ré · 2017
Earlier work this paper cites.
Paying attention to descriptions generated by image captioning models
Hamed Rezazadegan Tavakoli, Rakshith Shetty, Ali Borji, and Jorma T. Laaksonen · 2017
Earlier work this paper cites.
Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation
Mikel Artetxe, Gorka Labaka, Iñigo Lopez-Gazpio, and Eneko Agirre · 2018
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties, 2018
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Cited alongside, same era.
Analytic versus holistic cognition: Constructs and measurement
Minkyung Koo, Jong An Choi, and Incheol Choi · 2018
Cited alongside, same era.
Fictive motion in language and ‘ception’
Leonard Talmy · 2018
Cited alongside, same era.
Improving adversarial robustness of ensembles with diversity training, 2019
Sanjay Kariyappa and Moinuddin K. Qureshi · 2019
Cited alongside, same era.
Construal
Ronald W Langacker · 2019
Towards multi-lingual visual question answering
Soravit Changpinyo et al · 2022
Later among the works it cites.
There’s a time and place for reasoning beyond the image
Xingyu Fu, Ben Zhou, Ishaan Preetam Chandratreya, Carl Vondrick, and Dan Roth · 2022
Later among the works it cites.
Unison: Unpaired cross-lingual image captioning, 2022
Jiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq Joty, and Jiuxiang Gu · 2022
Later among the works it cites.
Perceptual experience
Christopher S Hill · 2022
Later among the works it cites.
Underspecification in scene description-to-depiction tasks
Ben Hutchinson, Jason Baldridge, and Vinodkumar Prabhakaran · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Coco-cn for cross-lingual image tagging, captioning and retrieval, 2019
Xirong Li, Chaoxi Xu, Xiaoxu Wang, Weiyu Lan, Zhengxiong Jia, Gang Yang, and Jieping Xu · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Cited alongside, same era.
Large-scale datasets for going deeper in image understanding
Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, Yizhou Wang, and Yonggang Wang · 2019
Cited alongside, same era.
The inclusive images competition
James Atwood, Yoni Halpern, Pallavi Baljekar, Eric Breck, D Sculley, Pavel Ostyakov, Sergey I Nikolenko, Igor Ivanov, Roman Solovyev, Weimin Wang, et al · 2020
Cited alongside, same era.
Cross-cultural differences in visual perception
Jiři Ceněk and Ceněk Sasinka · 2020
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Ward Church, and Mohamed Elhoseiny · 2022
Later among the works it cites.
Cross-linguistic comparison of linguistic feature encoding in BERT models for typologically different languages
Yulia Otmakhova, Karin Verspoor, and Jey Han Lau · 2022
Later among the works it cites.
Multilingual multimodal learning with machine translated text
Chen Qiu, Dan Oneaţă, Emanuele Bugliarello, Stella Frank, and Desmond Elliott · 2022
Later among the works it cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Later among the works it cites.
Crossmodal-3600: A massively multilingual multimodal evaluation dataset
Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, and Radu Soricut · 2022
Later among the works it cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Later among the works it cites.
Babel-imagenet: Massively multilingual evaluation of vision-and-language representations
Gregor Geigle, Radu Timofte, and Goran Glavas · 2023
Closest in time.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A. Smith · 2023
Closest in time.
Is chatgpt a good translator? yes with gpt-4 as the engine, 2023
Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu · 2023
Closest in time.
From scarcity to efficiency: Improving clip training via visual-enriched captions
Zhengfeng Lai, Haotian Zhang, Wentao Wu, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang, et al · 2023
Closest in time.
That was the last straw, we need more: Are translation systems sensitive to disambiguating context?, 2023
Jaechan Lee, Alisa Liu, Orevaoghene Ahia, Hila Gonen, and Noah A. Smith · 2023
Closest in time.
Evjvqa challenge: Multilingual visual question answering
Ngan Luu-Thuy Nguyen, Nghia Hieu Nguyen, Duong T.D. Vo, Khanh Quoc Tran, and Kiet Van Nguyen · 2023
Closest in time.
Linguistically motivated evaluation of the 2023 state-of-the-art machine translation: Can ChatGPT outperform NMT?
Shushen Manakhimova, Eleftherios Avramidis, Vivien Macketanz, Ekaterina Lapshinova-Koltunski, Sergei Bagdasarov, and Sebastian Möller · 2023
Closest in time.
Improving multimodal datasets with image captioning
Thao Nguyen, Samir Yitzhak Gadre, Gabriel Ilharco, Sewoong Oh, and Ludwig Schmidt · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Identifying the correlation between language distance and cross-lingual transfer in a multilingual representation space, 2023
Fred Philippy, Siwen Guo, and Shohreh Haddadan · 2023
Closest in time.
On the connection between pre-training data diversity and fine-tuning robustness, 2023
Vivek Ramanujan, Thao Nguyen, Sewoong Oh, Ludwig Schmidt, and Ali Farhadi · 2023
Closest in time.
Leveraging gpt-4 for automatic translation post-editing, 2023
Vikas Raunak, Amr Sharaf, Yiren Wang, Hany Hassan Awadallah, and Arul Menezes · 2023
Closest in time.
Nlpositionality: Characterizing design biases of datasets and models
Sebastin Santy, Jenny T Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap · 2023
Closest in time.
Literature adventures with linguistic inquiry and word count
Kristin L Schaefer, Jorge Rosales, and Jerrod A Henderson · 2023
Closest in time.
Agile modeling: From concept to classifier in minutes
Otilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan, Vittorio Ferrari, Sasan Tavakkol, Wenlei Zhou, Aditya Avinash, Enming Luo, Neil Gordon Alldrin, MohammadHossein Bateni, Gabriel Berger, Andrew Bunner, Chun-Ta Lu, Javier A Rey, Giulia DeSalvo, Ranjay Krishna, and Ariel Fuxman · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Deep image captioning: A review of methods, trends and future challenges
Liming Xu, Quan Tang, Jiancheng Lv, Bochuan Zheng, Xianhua Zeng, and Weisheng Li · 2023
Closest in time.
Givl: Improving geographical inclusivity of vision-language models with pre-training methods
Da Yin, Feng Gao, Govind Thattai, Michael F. Johnston, and Kai-Wei Chang · 2023
Closest in time.
Investigating cultural alignment of large language models
Badr AlKhamissi, Muhammad N. ElNokrashy, Mai AlKhamissi, and Mona Diab · 2024
Closest in time.
See it from my perspective: Diagnosing the western cultural bias of large vision-language models in image understanding, 2024
Amith Ananthram, Elias Stengel-Eskin, Carl Vondrick, Mohit Bansal, and Kathleen McKeown · 2024
Closest in time.
Constructing multilingual visual-text datasets revealing visual multilingual ability of vision language models, 2024
Jesse Atuhurra, Iqra Ali, Tatsuya Hiraoka, Hidetaka Kamigaito, Tomoya Iwakura, and Taro Watanabe · 2024
Closest in time.
Multilingual diversity improves vision-language representations, 2024
Thao Nguyen et al · 2024
Closest in time.
mblip: Efficient bootstrapping of multilingual vision-llms, 2024
Gregor Geigle, Abhay Jain, Radu Timofte, and Goran Glavaš · 2024
Closest in time.
Why do llava vision-language models reply to images in english?
Musashi Hinck, Carolin Holtermann, Matthew Lyle Olson, Florian Schneider, Sungduk Yu, Anahita Bhiwandiwalla, Anne Lauscher, Shao-Yen Tseng, and Vasudev Lal · 2024
Closest in time.
Who’s in and who’s out? a case study of multimodal clip-filtering in datacomp, 2024
Rachel Hong, William Agnew, Tadayoshi Kohno, and Jamie Morgenstern · 2024
Closest in time.
Scoft: Self-contrastive fine-tuning for equitable image generation
Zhixuan Liu, Peter Schaldenbrand, Beverley-Claire Okogwu, Wenxuan Peng, Youngsik Yun, Andrew Hundt, Jihie Kim, and Jean Oh · 2024
Closest in time.
Modeling collaborator: Enabling subjective vision classification with minimal human effort via llm tool-use
Imad Eddine Toubal, Aditya Avinash, Neil Gordon Alldrin, Jan Dlabal, Wenlei Zhou, Enming Luo, Otilia Stretcu, Hao Xiong, Chun-Ta Lu, Howard Zhou, Ranjay Krishna, Ariel Fuxman, and Tom Duerig · 2024
Closest in time.
Do llamas work in english? on the latent language of multilingual transformers
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West · 2024
Closest in time.
Multilingual machine translation with large language models: Empirical results and analysis, 2024
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li · 2024
Closest in time.