Fetching the paper…
Reading the bibliography…
This research introduces DesignQA, a novel benchmark aimed at evaluating the proficiency of multimodal large language models (MLLMs) in comprehending and applying engineering requirements in technical documentation.
“Large language models in medicine.”
Thirunavukarasu, Arun James, Ting, Darren Shu Jeng, Elangovan, Kabilan, Gutierrez, Laura, Tan, Ting Fang and Ting, Daniel Shu Wei · 1940
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation.”
Papineni, Kishore, Roukos, Salim, Ward, Todd and Zhu, Wei-Jing · 2002
Earlier work this paper cites.
“Rouge: A package for automatic evaluation of summaries.”
Lin, Chin-Yew · 2004
Earlier work this paper cites.
“Mctest: A challenge dataset for the open-domain machine comprehension of text.”
Richardson, Matthew, Burges, Christopher JC and Renshaw, Erin · 2013
Earlier work this paper cites.
“Wikiqa: A challenge dataset for open-domain question answering.”
Yang, Yi, Yih, Wen-tau and Meek, Christopher · 2015
Earlier work this paper cites.
Product design and development
Ulrich, Karl T and Eppinger, Steven D · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text.”
Rajpurkar, Pranav, Zhang, Jian, Lopyrev, Konstantin and Liang, Percy · 2016
Earlier work this paper cites.
“Attention is all you need.”
Vaswani, Ashish, Shazeer, Noam, Parmar, Niki, Uszkoreit, Jakob, Jones, Llion, Gomez, Aidan N, Kaiser, Łukasz and Polosukhin, Illia · 2017
Earlier work this paper cites.
“Pure: A dataset of public requirements documents.”
Ferrari, Alessio, Spagnolo, Giorgio Oronzo and Gnesi, Stefania · 2017
Earlier work this paper cites.
“Sentence-bert: Sentence embeddings using siamese bert-networks.”
Reimers, Nils and Gurevych, Iryna · 2019
Earlier work this paper cites.
“Bertscore: Evaluating text generation with bert.”
Zhang, Tianyi, Kishore, Varsha, Wu, Felix, Weinberger, Kilian Q and Artzi, Yoav · 2019
Earlier work this paper cites.
“Retrieval-augmented generation for knowledge-intensive nlp tasks.”
Lewis, Patrick, Perez, Ethan, Piktus, Aleksandra, Petroni, Fabio, Karpukhin, Vladimir, Goyal, Naman, Küttler, Heinrich, Lewis, Mike, Yih, Wen-tau, Rocktäschel, Tim et al · 2020
Earlier work this paper cites.
“A dataset of information-seeking questions and answers anchored in research papers.”
Dasigi, Pradeep, Lo, Kyle, Beltagy, Iz, Cohan, Arman, Smith, Noah A and Gardner, Matt · 2021
Earlier work this paper cites.
“Adapting natural language processing for technical text.”
Dima, Alden, Lukens, Sarah, Hodkiewicz, Melinda, Sexton, Thurston and Brundage, Michael P · 2021
Earlier work this paper cites.
“Technical language processing: Unlocking maintenance knowledge.”
Brundage, Michael P, Sexton, Thurston, Hodkiewicz, Melinda, Dima, Alden and Lukens, Sarah · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models.”
Hu, Edward J, Shen, Yelong, Wallis, Phillip, Allen-Zhu, Zeyuan, Li, Yuanzhi, Wang, Shean, Wang, Lu and Chen, Weizhu · 2021
Earlier work this paper cites.
“Learn to explain: Multimodal reasoning via thought chains for science question answering.”
Lu, Pan, Mishra, Swaroop, Xia, Tanglin, Qiu, Liang, Chang, Kai-Wei, Zhu, Song-Chun, Tafjord, Oyvind, Clark, Peter and Kalyan, Ashwin · 2022
Cited alongside, same era.
“Deep learning for technical document classification.”
Jiang, Shuo, Hu, Jie, Magee, Christopher L and Luo, Jianxi · 2022
Cited alongside, same era.
“Self-instruct: Aligning language models with self-generated instructions.”
Wang, Yizhong, Kordi, Yeganeh, Mishra, Swaroop, Liu, Alisa, Smith, Noah A, Khashabi, Daniel and Hajishirzi, Hannaneh · 2022
Cited alongside, same era.
“LlamaIndex.” (2022)
Liu, Jerry · 2022
Cited alongside, same era.
“A survey on retrieval-augmented text generation.”
Li, Huayang, Su, Yixuan, Cai, Deng, Wang, Yan and Liu, Lemao · 2022
Cited alongside, same era.
“Visual Instruction Tuning.” (2023)
Liu, Haotian, Li, Chunyuan, Wu, Qingyang and Lee, Yong Jae · 2023
Later among the works it cites.
“MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.”
Fu, Chaoyou, Chen, Peixian, Shen, Yunhang, Qin, Yulei, Zhang, Mengdan, Lin, Xu, Qiu, Zhenyu, Lin, Wei, Yang, Jinrui, Zheng, Xiawu et al · 2023
Later among the works it cites.
“Mmbench: Is your multi-modal model an all-around player?”
Liu, Yuan, Duan, Haodong, Zhang, Yuanhan, Li, Bo, Zhang, Songyang, Zhao, Wangbo, Yuan, Yike, Wang, Jiaqi, He, Conghui, Liu, Ziwei et al · 2023
Later among the works it cites.
“Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.”
Yue, Xiang, Ni, Yuansheng, Zhang, Kai, Zheng, Tianyu, Liu, Ruoqi, Zhang, Ge, Stevens, Samuel, Jiang, Dongfu, Ren, Weiming, Sun, Yuxuan et al · 2023
Later among the works it cites.
“Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“GPT-4V(ision) system card.” (2023)
OpenAI · 2023
Cited alongside, same era.
“Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv.”
Bubeck, Sébastien, Chandrasekaran, Varun, Eldan, Ronen, Gehrke, Johannes, Horvitz, Eric, Kamar, Ece, Lee, Peter, Lee, Yin Tat, Li, Yuanzhi, Lundberg, Scott et al · 2023
Cited alongside, same era.
“Introduction to Artificial Intelligence: Current Developments, Concerns and Possibilities for Education.”
Barbhuiya, Rejaul Karim · 2023
Cited alongside, same era.
“The future landscape of large language models in medicine.”
Clusmann, Jan, Kolbinger, Fiona R, Muti, Hannah Sophie, Carrero, Zunamys I, Eckardt, Jan-Niklas, Laleh, Narmin Ghaffari, Löffler, Chiara Maria Lavinia, Schwarzkopf, Sophie-Caroline, Unger, Michaela, Veldhuizen, Gregory P et al · 2023
Cited alongside, same era.
“ChatGPT for good? On opportunities and challenges of large language models for education.”
Kasneci, Enkelejda, Seßler, Kathrin, Küchemann, Stefan, Bannert, Maria, Dementieva, Daryna, Fischer, Frank, Gasser, Urs, Groh, Georg, Günnemann, Stephan, Hüllermeier, Eyke et al · 2023
Cited alongside, same era.
“How Can Large Language Models Help Humans in Design and Manufacturing?”
Makatura, Liane, Foshey, Michael, Wang, Bohan, HähnLein, Felix, Ma, Pingchuan, Deng, Bolei, Tjandrasuwita, Megan, Spielberg, Andrew, Owens, Crystal Elaine, Chen, Peter Yichen et al · 2023
Cited alongside, same era.
“From Concept to Manufacturing: Evaluating Vision-Language Models for Engineering Design.”
Picard, Cyril, Edwards, Kristen M, Doris, Anna C, Man, Brandon, Giannone, Giorgio, Alam, Md Ferdous and Ahmed, Faez · 2023
Cited alongside, same era.
Chen, Yang, Hu, Hexiang, Luan, Yi, Sun, Haitian, Changpinyo, Soravit, Ritter, Alan and Chang, Ming-Wei · 2023
Later among the works it cites.
“Generative transformers for design concept generation.”
Zhu, Qihao and Luo, Jianxi · 2023
Later among the works it cites.
“A survey on evaluation of large language models.”
Chang, Yupeng, Wang, Xu, Wang, Jindong, Wu, Yuan, Zhu, Kaijie, Chen, Hao, Yang, Linyi, Yi, Xiaoyuan, Wang, Cunxiang, Wang, Yidong et al · 2023
Later among the works it cites.
“Human Landing System (HLS) Program Extravehicular Activity (EVA) Compatibility Interface Requirements Document (IRD).”
Kovich, Christine N · 2023
Later among the works it cites.
“Improved baselines with visual instruction tuning.”
Liu, Haotian, Li, Chunyuan, Li, Yuheng and Lee, Yong Jae · 2023
Later among the works it cites.
“Visual instruction tuning.”
Liu, Haotian, Li, Chunyuan, Wu, Qingyang and Lee, Yong Jae · 2024
Closest in time.
“GPT-4o System Card.” (2024)
OpenAI · 2024
Closest in time.
“Multi-modal machine learning in engineering design: A review and future directions.”
Song, Binyang, Zhou, Rui and Ahmed, Faez · 2024
Closest in time.
“2024 Formula 1 Technical Regulations.”
de l’Automobile, 2023 Federation Internationale · 2024
Closest in time.
“Qlora: Efficient finetuning of quantized llms.”
Dettmers, Tim, Pagnoni, Artidoro, Holtzman, Ari and Zettlemoyer, Luke · 2024
Closest in time.
“Lab: Large-scale alignment for chatbots.”
Sudalairaj, Shivchander, Bhandwaldar, Abhishek, Pareja, Aldo, Xu, Kai, Cox, David D and Srivastava, Akash · 2024
Closest in time.
“Inadequacies of large language model benchmarks in the era of generative artificial intelligence.”
McIntosh, Timothy R, Susnjak, Teo, Liu, Tong, Watters, Paul and Halgamuge, Malka N · 2024
Closest in time.