Fetching the paper…
Reading the bibliography…
Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents.
Causal diagrams for empirical research
Judea Pearl · 1995
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Keisuke Hirano, Guido W Imbens, and Geert Ridder · 2003
Earlier work this paper cites.
Causality and causation in tort law
Robert Young, Michael Faure, and Paul Fenn · 2004
Earlier work this paper cites.
Causation and causal inference in epidemiology
Kenneth J Rothman and Sander Greenland · 2005
Earlier work this paper cites.
Estimating high-dimensional directed acyclic graphs with the pc-algorithm
Markus Kalisch and Peter Bühlman · 2007
Earlier work this paper cites.
Minimally supervised event causality identification
Quang Do, Yee Seng Chan, and Dan Roth · 2011
Earlier work this paper cites.
Estimating the causal effect of gun prevalence on homicide rates: A local average treatment effect approach
Tomislav Kovandzic, Mark E Schaffer, and Gary Kleck · 2013
Earlier work this paper cites.
On the joint use of propensity and prognostic scores in estimation of the average treatment effect on the treated: a simulation study
Finbarr P Leacy and Elizabeth A Stuart · 2014
Earlier work this paper cites.
Causal Inference in Statistics: A Primer
Judea Pearl, Madelyn Glymour, and Nicholas P. Jewell · 2016
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
The book of why: the new science of cause and effect
Judea Pearl and Dana Mackenzie · 2018
Earlier work this paper cites.
The global landscape of ai ethics guidelines
Anna Jobin, Marcello Ienca, and Effy Vayena · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum · 2019
Earlier work this paper cites.
Uncertainty-based traffic accident anticipation with spatio-temporal relational learning
Wentao Bao, Qi Yu, and Yu Kong · 2020
Earlier work this paper cites.
Poverty, depression, and anxiety: Causal evidence and mechanisms
Matthew Ridley, Gautam Rao, Frank Schilbach, and Vikram Patel · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al · 2021
Earlier work this paper cites.
e-care: a new dataset for exploring explainable causal reasoning
Li Du, Xiao Ding, Kai Xiong, Ting Liu, and Bing Qin · 2022
Earlier work this paper cites.
Crass: A novel data set and benchmark to test counterfactual reasoning of large language models
Jörg Frohberg and Frank Binder · 2022
Earlier work this paper cites.
Trustworthy ai: A computational perspective
Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil Jain, and Jiliang Tang · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, , and Jason Wei · 2022
Cited alongside, same era.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Cited alongside, same era.
Maven-ere: A unified large-scale dataset for event coreference, temporal, causal, and subevent relation extraction
Xiaozhi Wang, Yulin Chen, Ning Ding, Hao Peng, Zimu Wang, Yankai Lin, Xu Han, Lei Hou, Juanzi Li, Zhiyuan Liu, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Wordcraft: story writing with large language models
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito · 2022
Videochat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao · 2023
Later among the works it cites.
Mvbench: A comprehensive multi-modal video understanding benchmark, 2023
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2023
Later among the works it cites.
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2023
Later among the works it cites.
Mixtral of experts
Mistral AI Team · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2023
Cited alongside, same era.
Holistic analysis of hallucination in gpt-4v(ision): Bias and interference challenges, 2023
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao · 2023
Cited alongside, same era.
How it’s made: Gemini multimodal prompting, 2023
Google Developers · 2023
Cited alongside, same era.
Rh20t: A robotic dataset for learning diverse skills in one-shot
Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Junbo Wang, Haoyi Zhu, and Cewu Lu · 2023
Cited alongside, same era.
Is chatgpt a good causal reasoner? a comprehensive evaluation
Jinglong Gao, Xiao Ding, Bing Qin, and Ting Liu · 2023
Cited alongside, same era.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Oscar Sainz, Jon Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre · 2023
Later among the works it cites.
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre · 2023
Later among the works it cites.
Mapping and pre-and post-failure analyses of the april 2019 kantutani landslide in la paz, bolivia, using synthetic aperture radar data
Monan Shan, Federico Raspini, Matteo Del Soldato, Abel Cruz, and Nicola Casagli · 2023
Later among the works it cites.
Zhelun Shi, Zhipin Wang, Hongxing Fan, Zhenfei Yin, Lu Sheng, Yu Qiao, and Jing Shao · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities
InternLM Team · 2023
Later among the works it cites.
Privacy-aware document visual question answering
Rubèn Tito, Khanh Nguyen, Marlon Tobaben, Raouf Kerkouche, Mohamed Ali Souibgui, Kangsoo Jung, Lei Kang, Ernest Valveny, Antti Honkela, Mario Fritz, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Fake alignment: Are llms really aligned well?
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, and Yingchun Wang · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Skywork: A more open bilingual foundation model
Tianwen Wei, Liang Zhao, Lichang Zhang, Bo Zhu, Lijie Wang, Haihua Yang, Biye Li, Cheng Cheng, Weiwei Lü, Rui Hu, et al · 2023
Later among the works it cites.
Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao, Weixiong Lin, Xiaoman Zhang, Xiao Zhou, Ziheng Zhao, Ya Zhang, Yanfeng Wang, et al · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models, 2023
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen · 2023
Later among the works it cites.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Zhenfei Yin, Jiong Wang, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Lu Sheng, Lei Bai, Xiaoshui Huang, Zhiyong Wang, et al · 2023
Later among the works it cites.
Benchmarking large language models for news summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Sentiment analysis in the era of large language models: A reality check
Wenxuan Zhang, Yue Deng, Bing Liu, Sinno Jialin Pan, and Lidong Bing · 2023
Later among the works it cites.
Llmeval: A preliminary study on how to evaluate large language models
Yue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu, Yongyao Shi, Tao Gui, Qi Zhang, and Xuanjing Huang · 2023
Later among the works it cites.
Don’t make your llm an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han · 2023
Later among the works it cites.