Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are advancing at an amazing speed and have become indispensable across academia, industry, and daily applications.
1912
Earlier work this paper cites.
C. Elkan and R. Greiner, “D. b. lenat and r. v. guha, building large knowledge-based systems: Representation and inference in the cyc project,” Artif. Intell. , vol. 61, no. 1, pp. 41–52, 1993. [Online]. Available: https://doi.org/10.1016/0004-3702(93)90092-P
1993
Earlier work this paper cites.
P. Singh, T. Lin, E. T. Mueller, G. Lim, T. Perkins, and W. L. Zhu, “Open mind common sense: Knowledge acquisition from the general public,” in On the Move to Meaningful Internet Systems, 2002 - DOA/CoopIS/ODBASE 2002 Confederated International Conferences DOA, CoopIS and ODBASE 2002 Irvine, California, USA, October 30 - November 1, 2002, Proceedings , ser. Lecture Notes in Computer Science, R. Meersman and Z. Tari, Eds., vol. 2519. Springer, 2002, pp. 1223–1237. [Online]. Available: https://doi.org/10.1007/3-540-36124-3\_77
2002
Earlier work this paper cites.
2002
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
M. Sicilia, E. García-Barriocanal, S. Sánchez-Alonso, and M. E. Rodríguez, “On integrating learning object metadata inside the opencyc knowledge base,” in Proceedings of the IEEE International Conference on Advanced Learning Technologies, ICALT 2004, Joensuu, Finland, August 30 - September 1, 2004 , Kinshuk, C. Looi, E. Sutinen, D. G. Sampson, I. Aedo, L. Uden, and E. Kähkönen, Eds. IEEE Computer Society, 2004. [Online]. Available: https://doi.org/10.1109/ICALT.2004.1357711
2004
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Information processing & management , vol. 45, no. 4, pp. 427–437, 2009
2009
Earlier work this paper cites.
E. Cambria, R. Speer, C. Havasi, and A. Hussain, “Senticnet: A publicly available semantic resource for opinion mining,” in Commonsense Knowledge, Papers from the 2010 AAAI Fall Symposium, Arlington, Virginia, USA, November 11-13, 2010 , ser. AAAI Technical Report, vol. FS-10-02. AAAI, 2010. [Online]. Available: http://www.aaai.org/ocs/index.php/FSS/FSS10/paper/view/2216
2010
Earlier work this paper cites.
A. S. Gordon, Z. Kozareva, and M. Roemmele, “Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning,” in Proceedings of the 6th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2012, Montréal, Canada, June 7-8, 2012 , E. Agirre, J. Bos, and M. T. Diab, Eds. The Association for Computer Linguistics, 2012, pp. 394–398. [Online]. Available: https://aclanthology.org/S12-1052/
2012
Earlier work this paper cites.
J. Berant, A. Chou, R. Frostig, and P. Liang, “Semantic parsing on freebase from question-answer pairs,” in Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , 2013, pp. 1533–1544
2013
Earlier work this paper cites.
R. Alur, R. Bodík, G. Juniwal, M. M. K. Martin, M. Raghothaman, S. A. Seshia, R. Singh, A. Solar-Lezama, E. Torlak, and A. Udupa, “Syntax-guided synthesis,” in Formal Methods in Computer-Aided Design, FMCAD 2013, Portland, OR, USA, October 20-23, 2013 . IEEE, 2013, pp. 1–8. [Online]. Available: https://ieeexplore.ieee.org/document/6679385/
2013
Earlier work this paper cites.
Y. Wang, L. Wang, Y. Li, D. He, and T.-Y. Liu, “A theoretical analysis of ndcg type ranking measures,” in Conference on learning theory . PMLR, 2013, pp. 25–54
2013
Earlier work this paper cites.
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 961–970, 2015. [Online]. Available: https://api.semanticscholar.org/CorpusID:1710722
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in ECCV , 2016
2016
Earlier work this paper cites.
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer, “TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Vancouver, Canada: Association for Computational Linguistics, Jul. 2017, pp. 1601–1611
2017
Earlier work this paper cites.
Y. Wang, X. Liu, and S. Shi, “Deep neural solver for math word problems,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , M. Palmer, R. Hwa, and S. Riedel, Eds. Copenhagen, Denmark: Association for Computational Linguistics, Sep. 2017, pp. 845–854. [Online]. Available: https://aclanthology.org/D17-1088
2017
Earlier work this paper cites.
S. Zhang, R. Rudinger, K. Duh, and B. V. Durme, “Ordinal common-sense inference,” Trans. Assoc. Comput. Linguistics , vol. 5, pp. 379–395, 2017. [Online]. Available: https://doi.org/10.1162/tacl\_a\_00068
2017
Earlier work this paper cites.
R. Speer, J. Chin, and C. Havasi, “Conceptnet 5.5: An open multilingual graph of general knowledge,” in Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA , S. Singh and S. Markovitch, Eds. AAAI Press, 2017, pp. 4444–4451. [Online]. Available: https://doi.org/10.1609/aaai.v31i1.11164
2017
Earlier work this paper cites.
H. Rashkin, A. Bosselut, M. Sap, K. Knight, and Y. Choi, “Modeling naive psychology of characters in simple commonsense stories,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers , I. Gurevych and Y. Miyao, Eds. Association for Computational Linguistics, 2018, pp. 2289–2299. [Online]. Available: https://aclanthology.org/P18-1213/
2018
Earlier work this paper cites.
R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi, “SWAG: A large-scale adversarial dataset for grounded commonsense inference,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds. Association for Computational Linguistics, 2018, pp. 93–104. [Online]. Available: https://doi.org/10.18653/v1/d18-1009
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Côté, Á. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, M. J. Hausknecht, L. E. Asri, M. Adada, W. Tay, and A. Trischler, “Textworld: A learning environment for text-based games,” in Computer Games - 7th Workshop, CGW 2018, Held in Conjunction with the 27th International Conference on Artificial Intelligence, IJCAI 2018, Stockholm, Sweden, July 13, 2018, Revised Selected Papers , ser. Communications in Computer and Information Science, T. Cazenave, A. Saffidine, and N. R. Sturtevant, Eds., vol. 1017. Springer, 2018, pp. 41–75. [Online]. Available: https://doi.org/10.1007/978-3-030-24337-1\_3
2018
Earlier work this paper cites.
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning, “Hotpotqa: A dataset for diverse, explainable multi-hop question answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , 2018, pp. 2369–2380
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP , T. Linzen, G. Chrupała, and A. Alishahi, Eds. Brussels, Belgium: Association for Computational Linguistics, Nov. 2018, pp. 353–355. [Online]. Available: https://aclanthology.org/W18-5446/
2018
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in CVPR , 2018
2018
Earlier work this paper cites.
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller, “Language models as knowledge bases?” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 2463–2473
2019
Earlier work this paper cites.
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov, “Natural questions: A benchmark for question answering research,” Transactions of the Association for Computational Linguistics , vol. 7, pp. 452–466, 2019
2019
Earlier work this paper cites.
A. Talmor, J. Herzig, N. Lourie, and J. Berant, “CommonsenseQA: A question answering challenge targeting commonsense knowledge,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4149–4158. [Online]. Available: https://aclanthology.org/N19-1421/
2019
Earlier work this paper cites.
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi, “MathQA: Towards interpretable math word problem solving with operation-based formalisms,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 2357–2367. [Online]. Available: https://aclanthology.org/N19-1245
2019
Earlier work this paper cites.
F. Chollet, “On the measure of intelligence,” arXiv preprint arXiv:1911.01547 , 2019
2019
Earlier work this paper cites.
K. Sinha, S. Sodhani, J. Dong, J. Pineau, and W. L. Hamilton, “CLUTRR: A diagnostic benchmark for inductive reasoning from text,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 , K. Inui, J. Jiang, V. Ng, and X. Wan, Eds. Association for Computational Linguistics, 2019, pp. 4505–4514. [Online]. Available: https://doi.org/10.18653/v1/D19-1458
2019
Earlier work this paper cites.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 . Computer Vision Foundation / IEEE, 2019, pp. 6720–6731. [Online]. Available: http://openaccess.thecvf.com/content\_CVPR\_2019/html/Zellers\_From\_Recognition\_to\_Cognition\_Visual\_Commonsense\_Reasoning\_CVPR\_2019\_paper.html
2019
Earlier work this paper cites.
Q. N. Ben Zhou, Daniel Khashabi and D. Roth, ““going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding,” in EMNLP , 2019
2019
Earlier work this paper cites.
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “Hellaswag: Can a machine really finish your sentence?” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019
2019
Earlier work this paper cites.
M. Sap, R. L. Bras, E. Allaway, C. Bhagavatula, N. Lourie, H. Rashkin, B. Roof, N. A. Smith, and Y. Choi, “ATOMIC: an atlas of machine commonsense for if-then reasoning,” in The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019 . AAAI Press, 2019, pp. 3027–3035. [Online]. Available: https://doi.org/10.1609/aaai.v33i01.33013027
2019
Earlier work this paper cites.
M. Savva, J. Malik, D. Parikh, D. Batra, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, and V. Koltun, “Habitat: A platform for embodied AI research,” in 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 . IEEE, 2019, pp. 9338–9346. [Online]. Available: https://doi.org/10.1109/ICCV.2019.00943
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in CVPR , 2019
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , 2019, pp. 3195–3204
2019
Earlier work this paper cites.
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Superglue: A stickier benchmark for general-purpose language understanding systems,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli, “ELI5: Long form question answering,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , A. Korhonen, D. Traum, and L. Màrquez, Eds. Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 3558–3567. [Online]. Available: https://aclanthology.org/P19-1346/
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Aslanyan and U. Porwal, “Position bias estimation for unbiased learning-to-rank in ecommerce search,” in String Processing and Information Retrieval: 26th International Symposium, SPIRE 2019, Segovia, Spain, October 7–9, 2019, Proceedings 26 . Springer, 2019, pp. 47–64
2019
Earlier work this paper cites.
S.-y. Miao, C.-C. Liang, and K.-Y. Su, “A diverse corpus for evaluating and developing English math word problem solvers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computational Linguistics, Jul. 2020, pp. 975–984. [Online]. Available: https://aclanthology.org/2020.acl-main.92
2020
Earlier work this paper cites.
J. Liu, L. Cui, H. Liu, D. Huang, Y. Wang, and Y. Zhang, “Logiqa: A challenge dataset for machine reading comprehension with logical reasoning,” in Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020 , C. Bessiere, Ed. ijcai.org, 2020, pp. 3622–3628. [Online]. Available: https://doi.org/10.24963/ijcai.2020/501
2020
Earlier work this paper cites.
W. Yu, Z. Jiang, Y. Dong, and J. Feng, “Reclor: A reading comprehension dataset requiring logical reasoning,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020. [Online]. Available: https://openreview.net/forum?id=HJgJtT4tvB
2020
Earlier work this paper cites.
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi, “PIQA: reasoning about physical commonsense in natural language,” in The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020 . AAAI Press, 2020, pp. 7432–7439. [Online]. Available: https://doi.org/10.1609/aaai.v34i05.6239
2020
Earlier work this paper cites.
H. Zhang, X. Liu, H. Pan, Y. Song, and C. W. Leung, “ASER: A large-scale eventuality knowledge graph,” in WWW ’20: The Web Conference 2020, Taipei, Taiwan, April 20-24, 2020 , Y. Huang, I. King, T. Liu, and M. van Steen, Eds. ACM / IW3C2, 2020, pp. 201–211. [Online]. Available: https://doi.org/10.1145/3366423.3380107
2020
Earlier work this paper cites.
J. S. Park, C. Bhagavatula, R. Mottaghi, A. Farhadi, and Y. Choi, “Visualcomet: Reasoning about the dynamic context of a still image,” in Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part V , ser. Lecture Notes in Computer Science, A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, Eds., vol. 12350. Springer, 2020, pp. 508–524. [Online]. Available: https://doi.org/10.1007/978-3-030-58558-7\_30
2020
Earlier work this paper cites.
S. Heindorf, Y. Scholten, H. Wachsmuth, A. N. Ngomo, and M. Potthast, “Causenet: Towards a causality graph extracted from the web,” in CIKM ’20: The 29th ACM International Conference on Information and Knowledge Management, Virtual Event, Ireland, October 19-23, 2020 , M. d’Aquin, S. Dietze, C. Hauff, E. Curry, and P. Cudré-Mauroux, Eds. ACM, 2020, pp. 3023–3030. [Online]. Available: https://doi.org/10.1145/3340531.3412763
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020. [Online]. Available: https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
2020
Earlier work this paper cites.
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “RealToxicityPrompts: Evaluating neural toxic degeneration in language models,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 3356–3369. [Online]. Available: https://aclanthology.org/2020.findings-emnlp.301/
2020
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, “Aligning ai with shared human values,” CoRR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “Alfred: A benchmark for interpreting grounded instructions for everyday tasks,” in CVPR , 2020
2020
Earlier work this paper cites.
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel, “Reverie: Remote embodied visual referring expression in real indoor environments,” in CVPR , 2020
2020
Earlier work this paper cites.
W. Chen, H. Zha, Z. Chen, W. Xiong, H. Wang, and W. Y. Wang, “HybridQA: A dataset of multi-hop question answering over tabular and textual data,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 1026–1036. [Online]. Available: https://aclanthology.org/2020.findings-emnlp.91/
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Zhang and E. Choi, “SituatedQA: Incorporating extra-linguistic contexts into QA,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 7371–7387. [Online]. Available: https://aclanthology.org/2021.emnlp-main.586/
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the math dataset,” NeurIPS , 2021
2021
Earlier work this paper cites.
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, “Measuring coding challenge competence with apps,” NeurIPS , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating large language models trained on code,” CoRR , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
O. Tafjord, B. Dalvi, and P. Clark, “Proofwriter: Generating implications, proofs, and abductive statements over natural language,” in Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 , ser. Findings of ACL, C. Zong, F. Xia, W. Li, and R. Navigli, Eds., vol. ACL/IJCNLP 2021. Association for Computational Linguistics, 2021, pp. 3621–3634. [Online]. Available: https://doi.org/10.18653/v1/2021.findings-acl.317
2021
Earlier work this paper cites.
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM , vol. 64, no. 9, pp. 99–106, 2021
2021
Earlier work this paper cites.
M. T. Phu, M. V. Nguyen, and T. H. Nguyen, “Fine-grained temporal relation extraction with ordered-neuron LSTM and graph convolutional networks,” in Proceedings of the Seventh Workshop on Noisy User-generated Text, W-NUT 2021, Online, November 11, 2021 , W. Xu, A. Ritter, T. Baldwin, and A. Rahimi, Eds. Association for Computational Linguistics, 2021, pp. 35–45. [Online]. Available: https://doi.org/10.18653/v1/2021.wnut-1.5
2021
Earlier work this paper cites.
Q. Lyu, L. Zhang, and C. Callison-Burch, “Goal-oriented script construction,” in Proceedings of the 14th International Conference on Natural Language Generation, INLG 2021, Aberdeen, Scotland, UK, 20-24 September, 2021 , A. Belz, A. Fan, E. Reiter, and Y. Sripada, Eds. Association for Computational Linguistics, 2021, pp. 184–200. [Online]. Available: https://doi.org/10.18653/v1/2021.inlg-1.19
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Shridhar, X. Yuan, M. Côté, Y. Bisk, A. Trischler, and M. J. Hausknecht, “Alfworld: Aligning text and embodied environments for interactive learning,” in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net, 2021. [Online]. Available: https://openreview.net/forum?id=0IOX0YcCdTn
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
M. Mathew, D. Karatzas, and C. Jawahar, “Docvqa: A dataset for vqa on document images,” in WACV , 2021
2021
Earlier work this paper cites.
J. Xu, D. Ju, M. Li, Y.-L. Boureau, J. Weston, and E. Dinan, “Bot-adversarial dialogue for safe conversational agents,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, Eds. Online: Association for Computational Linguistics, Jun. 2021, pp. 2950–2968. [Online]. Available: https://aclanthology.org/2021.naacl-main.235/
2021
Earlier work this paper cites.
M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant, “Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 346–361, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych, “BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , 2021. [Online]. Available: https://openreview.net/forum?id=wCu6T5xFjeJ
2021
Earlier work this paper cites.
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang, “FinQA: A dataset of numerical reasoning over financial data,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, Eds. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 3697–3711. [Online]. Available: https://aclanthology.org/2021.emnlp-main.300/
2021
Earlier work this paper cites.
W. Chen, M. wei Chang, E. Schlinger, W. Wang, and W. Cohen, “Open question answering over tables and text,” Proceedings of ICLR 2021 , 2021
2021
Earlier work this paper cites.
A. Talmor, O. Yoran, A. Catav, D. Lahav, Y. Wang, A. Asai, G. Ilharco, H. Hajishirzi, and J. Berant, “Multimodal{qa}: complex question answering over text, tables and images,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=ee6W5UgQLa
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learning . Pmlr, 2021, pp. 8821–8831
2021
Earlier work this paper cites.
Y. Zhang, P. Ren, and M. de Rijke, “A human-machine collaborative framework for evaluating malevolence in dialogues,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , C. Zong, F. Xia, W. Li, and R. Navigli, Eds. Online: Association for Computational Linguistics, Aug. 2021, pp. 5612–5623. [Online]. Available: https://aclanthology.org/2021.acl-long.436/
2021
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 3214–3252. [Online]. Available: https://aclanthology.org/2022.acl-long.229/
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
N. Young, Q. Bao, J. Bensemann, and M. Witbrock, “Abductionrules: Training transformers to explain unexpected inputs,” in Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022 , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Association for Computational Linguistics, 2022, pp. 218–227. [Online]. Available: https://doi.org/10.18653/v1/2022.findings-acl.19
2022
Earlier work this paper cites.
L. Du, X. Ding, K. Xiong, T. Liu, and B. Qin, “e-care: a new dataset for exploring explainable causal reasoning,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Association for Computational Linguistics, 2022, pp. 432–446. [Online]. Available: https://doi.org/10.18653/v1/2022.acl-long.33
2022
Earlier work this paper cites.
X. Wang, Y. Chen, N. Ding, H. Peng, Z. Wang, Y. Lin, X. Han, L. Hou, J. Li, Z. Liu, P. Li, and J. Zhou, “MAVEN-ERE: A unified large-scale dataset for event coreference, temporal, causal, and subevent relation extraction,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 926–941. [Online]. Available: https://aclanthology.org/2022.emnlp-main.60/
2022
Earlier work this paper cites.
Y. Ma, Z. Wang, M. Li, Y. Cao, M. Chen, X. Li, W. Sun, K. Deng, K. Wang, A. Sun, and J. Shao, “MMEKG: multi-modal event knowledge graph towards universal representation across modalities,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, ACL 2022 - System Demonstrations, Dublin, Ireland, May 22-27, 2022 , V. Basile, Z. Kozareva, and S. Stajner, Eds. Association for Computational Linguistics, 2022, pp. 231–239. [Online]. Available: https://doi.org/10.18653/v1/2022.acl-demo.23
2022
Earlier work this paper cites.
S. Yao, H. Chen, J. Yang, and K. Narasimhan, “Webshop: Towards scalable real-world web interaction with grounded language agents,” in Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., 2022. [Online]. Available: http://papers.nips.cc/paper\_files/paper/2022/hash/82ad13ec01f9fe44c01cb91814fd7b8c-Abstract-Conference.html
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
S. Mishra, D. Khashabi, C. Baral, and H. Hajishirzi, “Cross-task generalization via natural language crowdsourcing instructions,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 3470–3487. [Online]. Available: https://aclanthology.org/2022.acl-long.244
2022
Earlier work this paper cites.
V. Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, and A. M. Rush, “Multitask prompted training enables zero-shot task generalization,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=9Vrb9D0WI4
2022
Earlier work this paper cites.
Y. Wang, S. Mishra, P. Alipoormolabashi, Y. Kordi, A. Mirzaei, A. Naik, A. Ashok, A. S. Dhanasekaran, A. Arunkumar, D. Stap, E. Pathak, G. Karamanolakis, H. Lai, I. Purohit, I. Mondal, J. Anderson, K. Kuznia, K. Doshi, K. K. Pal, M. Patel, M. Moradshahi, M. Parmar, M. Purohit, N. Varshney, P. R. Kaza, P. Verma, R. S. Puri, R. Karia, S. Doshi, S. K. Sampat, S. Mishra, S. Reddy A, S. Patro, T. Dixit, and X. Shen, “Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 5085–5109. [Online]. Available: https://aclanthology.org/2022.emnlp-main.340
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Sun, G. Xu, J. Deng, J. Cheng, C. Zheng, H. Zhou, N. Peng, X. Zhu, and M. Huang, “On the safety of conversational models: Taxonomy, dataset, and benchmark,” in Findings of the Association for Computational Linguistics: ACL 2022 , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 3906–3923. [Online]. Available: https://aclanthology.org/2022.findings-acl.308/
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar et al. , “Holistic evaluation of language models,” Transactions on Machine Learning Research , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. T. Ribeiro and S. Lundberg, “Adaptive testing and debugging of NLP models,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 3253–3267. [Online]. Available: https://aclanthology.org/2022.acl-long.230/
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
B. bench authors, “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” Transactions on Machine Learning Research , 2023. [Online]. Available: https://openreview.net/forum?id=uyTL5Bvosj
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen, “HaluEval: A large-scale hallucination evaluation benchmark for large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 6449–6464. [Online]. Available: https://aclanthology.org/2023.emnlp-main.397/
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “Multipl-e: A scalable and polyglot approach to benchmarking neural code generation,” IEEE Trans. Software Eng. , vol. 49, no. 7, pp. 3675–3691, 2023. [Online]. Available: https://doi.org/10.1109/TSE.2023.3267446
2023
Earlier work this paper cites.
A. Saparov and H. He, “Language models are greedy reasoners: A systematic formal analysis of chain-of-thought,” in The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. [Online]. Available: https://openreview.net/forum?id=qFVVBzXxR2V
2023
Earlier work this paper cites.
W. Zhao, J. T. Chiu, C. Cardie, and A. M. Rush, “Abductive commonsense reasoning exploiting mutually exclusive explanations,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , A. Rogers, J. L. Boyd-Graber, and N. Okazaki, Eds. Association for Computational Linguistics, 2023, pp. 14 883–14 896. [Online]. Available: https://doi.org/10.18653/v1/2023.acl-long.831
2023
Earlier work this paper cites.
M. Del and M. Fishel, “True detective: A deep abductive reasoning benchmark undoable for GPT-3 and challenging for GPT-4,” in Proceedings of the The 12th Joint Conference on Lexical and Computational Semantics, *SEM@ACL 2023, Toronto, Canada, July 13-14, 2023 , A. Palmer and J. Camacho-Collados, Eds. Association for Computational Linguistics, 2023, pp. 314–322. [Online]. Available: https://doi.org/10.18653/v1/2023.starsem-1.28
2023
Earlier work this paper cites.
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, A. Kluska, A. Lewkowycz, A. Agarwal, A. Power, A. Ray, A. Warstadt, A. W. Kocurek, A. Safaya, A. Tazarv, A. Xiang, and et al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” Trans. Mach. Learn. Res. , vol. 2023, 2023. [Online]. Available: https://openreview.net/forum?id=uyTL5Bvosj
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
C. Yu, W. Wang, X. Liu, J. Bai, Y. Song, Z. Li, Y. Gao, T. Cao, and B. Yin, “Folkscope: Intention knowledge graph construction for e-commerce commonsense discovery,” in Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , A. Rogers, J. L. Boyd-Graber, and N. Okazaki, Eds. Association for Computational Linguistics, 2023, pp. 1173–1191. [Online]. Available: https://doi.org/10.18653/v1/2023.findings-acl.76
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
C. An, S. Gong, M. Zhong, M. Li, J. Zhang, L. Kong, and X. Qiu, “L-eval: Instituting standardized evaluation for long context language models,” 2023
2023
Earlier work this paper cites.
X. Li, Y. Cao, M. Chen, and A. Sun, “Take a break in the middle: Investigating subgoals towards hierarchical script generation,” in Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , A. Rogers, J. L. Boyd-Graber, and N. Okazaki, Eds. Association for Computational Linguistics, 2023, pp. 10 129–10 147. [Online]. Available: https://doi.org/10.18653/v1/2023.findings-acl.644
2023
Earlier work this paper cites.
K. Valmeekam, M. Marquez, A. O. Hernandez, S. Sreedharan, and S. Kambhampati, “Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [Online]. Available: http://papers.nips.cc/paper\_files/paper/2023/hash/7a92bcdede88c7afd108072faf5485c8-Abstract-Datasets\_and\_Benchmarks.html
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,” Advances in Neural Information Processing Systems , vol. 36, pp. 24 678–24 704, 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin, “Free dolly: Introducing the world’s first truly open instruction-tuned llm,” Company Blog of Databricks , 2023
2023
Cited alongside, same era.
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, D. M. Nguyen, O. Stanley, R. Nagyfi, S. ES, S. Suri, D. A. Glushkov, A. V. Dantuluri, A. Maguire, C. Schuhmann, H. Nguyen, and A. J. Mattick, “Openassistant conversations - democratizing large language model alignment,” in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2023. [Online]. Available: https://openreview.net/forum?id=VSJotgbPHF
2023
Cited alongside, same era.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” Advances in Neural Information Processing Systems , vol. 36, pp. 46 595–46 623, 2023
2023
Cited alongside, same era.
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto, “Alpacaeval: An automatic evaluator of instruction-following models,” https://github.com/tatsu-lab/alpaca_eval
2023
Cited alongside, same era.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhang, L. Lei, L. Wu, R. Sun, Y. Huang, C. Long, X. Liu, X. Lei, J. Tang, and M. Huang, “Safetybench: Evaluating the safety of large language models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024 , L. Ku, A. Martins, and V. Srikumar, Eds. Association for Computational Linguistics, 2024, pp. 15 537–15 553. [Online]. Available: https://doi.org/10.18653/v1/2024.acl-long.830
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao, “Salad-bench: A hierarchical and comprehensive safety benchmark for large language models,” in Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 , L. Ku, A. Martins, and V. Srikumar, Eds. Association for Computational Linguistics, 2024, pp. 3923–3954. [Online]. Available: https://doi.org/10.18653/v1/2024.findings-acl.235
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Wang, H. Li, X. Han, P. Nakov, and T. Baldwin, “Do-not-answer: Evaluating safeguards in llms,” in EACL (Findings) , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Pang, S. Tang, R. Ye, Y. Xiong, B. Zhang, Y. Wang, and S. Chen, “Self-alignment of large language models via monopolylogue-based social scene simulation,” in Proceedings of the 41st International Conference on Machine Learning , 2024, pp. 39 416–39 447
2024
Later among the works it cites.
2024
Later among the works it cites.
N. Wichers, C. Denison, and A. Beirami, “Gradient-based language model red teaming,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 2862–2881
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security , 2024, pp. 1671–1685
2024
Later among the works it cites.
S. Tedeschi, F. Friedrich, P. Schramowski, K. Kersting, R. Navigli, H. Nguyen, and B. Li, “Alert: A comprehensive benchmark for assessing large language models’ safety through red teaming,” 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2404.08676
2024
Later among the works it cites.
X. Liu, N. Xu, M. Chen, and C. Xiao, “Autodan: Generating stealthy jailbreak prompts on aligned large language models,” in ICLR , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Shi, R. Xu, Y. Zhuang, Y. Yu, J. Zhang, H. Wu, Y. Zhu, J. Ho, C. Yang, and M. D. Wang, “Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 22 315–22 339
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
T. Yuan, Z. He, L. Dong, Y. Wang, R. Zhao, T. Xia, L. Xu, B. Zhou, F. Li, Z. Zhang et al. , “R-judge: Benchmarking safety risk awareness for llm agents,” in Findings of the Association for Computational Linguistics: EMNLP 2024 , 2024, pp. 1467–1490
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Xu, C. Feng, R. Shao, T. Ashby, Y. Shen, D. Jin, Y. Cheng, Q. Wang, and L. Huang, “Vision-flan: Scaling human-labeled tasks in visual instruction tuning,” in Findings of the Association for Computational Linguistics ACL 2024 , 2024, pp. 15 271–15 342
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
T. Lee, H. Tu, C. H. Wong, W. Zheng, Y. Zhou, Y. Mai, J. Roberts, M. Yasunaga, H. Yao, C. Xie et al. , “Vhelm: A holistic evaluation of vision language models,” Advances in Neural Information Processing Systems , vol. 37, pp. 140 632–140 666, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Cheng, Y. Lu, X. Gu, P. Ke, X. Liu, Y. Dong, H. Wang, J. Tang, and M. Huang, “Autodetect: Towards a unified framework for automated weakness detection in large language models,” in Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024 , 2024. [Online]. Available: https://aclanthology.org/2024.findings-emnlp.397
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu, “Chateval: Towards better llm-based evaluators through multi-agent debate,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. [Online]. Available: https://openreview.net/forum?id=FQepisCUWu
2024
Later among the works it cites.
Y. Li, S. Zhang, R. Wu, X. Huang, Y. Chen, W. Xu, G. Qi, and D. Min, “Mateval: A multi-agent discussion framework for advancing open-ended text evaluation,” in International Conference on Database Systems for Advanced Applications , ser. Lecture Notes in Computer Science, M. Onizuka, J. Lee, Y. Tong, C. Xiao, Y. Ishikawa, S. Amer-Yahia, H. V. Jagadish, and K. Lu, Eds., vol. 14856, Springer. Springer, 2024, pp. 415–426. [Online]. Available: https://doi.org/10.1007/978-981-97-5575-2\_31
2024
Later among the works it cites.
W. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. I. Jordan, J. E. Gonzalez, and I. Stoica, “Chatbot arena: An open platform for evaluating llms by human preference,” in Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. [Online]. Available: https://openreview.net/forum?id=3MW8GKNyzI
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Wang, Z. Hu, P. Lu, Y. Zhu, J. Zhang, S. Subramaniam, A. R. Loomba, S. Zhang, Y. Sun, and W. Wang, “SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models,” in Proceedings of the Forty-First International Conference on Machine Learning , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Ying, Y. Cao, K. Xiong, L. Cui, Y. He, and Y. Liu, “Intuitive or dependent? investigating llms’ behavior style to conflicting prompts,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 4221–4246
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Ma, Z. Gou, J. Hao, R. Xu, S. Wang, L. Pan, Y. Yang, Y. Cao, and A. Sun, “SciAgent: Tool-augmented language models for scientific reasoning,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 15 701–15 736. [Online]. Available: https://aclanthology.org/2024.emnlp-main.880
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su, “Gpt-4v(ision) is a generalist web agent, if grounded,” in Forty-first International Conference on Machine Learning , 2024. [Online]. Available: https://openreview.net/forum?id=piecKJ2DlB
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Li, H. Peng, X. Wang, Y. Qi, L. Hou, B. Xu, and J. Li, “MAVEN-FACT: A large-scale event factuality detection dataset,” in Findings of the Association for Computational Linguistics: EMNLP 2024 , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 11 140–11 158. [Online]. Available: https://aclanthology.org/2024.findings-emnlp.651/
2024
Later among the works it cites.
2024
Later among the works it cites.
N. Wang, Z. Peng, H. Que, J. Liu, W. Zhou, Y. Wu, H. Guo, R. Gan, Z. Ni, J. Yang, M. Zhang, Z. Zhang, W. Ouyang, K. Xu, W. Huang, J. Fu, and J. Peng, “RoleLLM: Benchmarking, eliciting, and enhancing role-playing abilities of large language models,” in Findings of the Association for Computational Linguistics: ACL 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 14 743–14 777. [Online]. Available: https://aclanthology.org/2024.findings-acl.878/
2024
Later among the works it cites.
G. Li, H. A. Al Kader Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: communicative agents for ”mind” exploration of large language model society,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2024
2024
Later among the works it cites.
G. Bai, J. Liu, X. Bu, Y. He, J. Liu, Z. Zhou, Z. Lin, W. Su, T. Ge, B. Zheng, and W. Ouyang, “MT-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 7421–7454. [Online]. Available: https://aclanthology.org/2024.acl-long.401/
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Mishra, A. Asai, V. Balachandran, Y. Wang, G. Neubig, Y. Tsvetkov, and H. Hajishirzi, “Fine-grained hallucination detection and editing for language models,” in First Conference on Language Modeling , 2024. [Online]. Available: https://openreview.net/forum?id=dJMTn3QOWO
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, Q. Lin, and D. Jiang, “Wizardlm: Empowering large pre-trained language models to follow complex instructions,” in The Twelfth International Conference on Learning Representations , 2024
2024
Later among the works it cites.
J. Ying, Y. Cao, Y. Bai, Q. Sun, B. Wang, W. Tang, Z. Ding, Y. Yang, X. Huang, and S. Yan, “Automating dataset updates towards reliable and timely evaluation of large language models,” in Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 2024, pp. 17 106–17 132. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2024/file/1e89c12621c0315373f20f0aeabe5dbe-Paper-Datasets_and_Benchmarks_Track.pdf
2024
Later among the works it cites.
Y. Bai, J. Ying, Y. Cao, X. Lv, Y. He, X. Wang, J. Yu, K. Zeng, Y. Xiao, H. Lyu et al. , “Benchmarking foundation models with language-model-as-an-examiner,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang, “Agentbench: Evaluating llms as agents,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. [Online]. Available: https://openreview.net/forum?id=zAdUB0aCTQ
2024
Later among the works it cites.
P. Laban, A. Fabbri, C. Xiong, and C.-S. Wu, “Summary of a haystack: A challenge to long-context LLMs and RAG systems,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 9885–9903. [Online]. Available: https://aclanthology.org/2024.emnlp-main.552/
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction-finetuned language models,” Journal of Machine Learning Research , vol. 25, no. 70, pp. 1–53, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
T. Y. Zhuo, “Ice-score: Instructing large language models to evaluate code,” in Findings of the Association for Computational Linguistics: EACL 2024 , 2024, pp. 2232–2242
2024
Later among the works it cites.
S. Saha, O. Levy, A. Celikyilmaz, M. Bansal, J. Weston, and X. Li, “Branch-solve-merge improves large language model evaluation and generation,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 8352–8370. [Online]. Available: https://aclanthology.org/2024.naacl-long.462
2024
Later among the works it cites.
H. He, H. Zhang, and D. Roth, “SocREval: Large language models with the socratic method for reference-free reasoning evaluation,” in Findings of the Association for Computational Linguistics: NAACL 2024 , K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 2736–2764. [Online]. Available: https://aclanthology.org/2024.findings-naacl.175/
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Song, H. Su, I. Shalyminov, J. Cai, and S. Mansour, “Finesure: Fine-grained summarization evaluation using llms,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 906–922
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Liu, T. Yang, S. Huang, Z. Zhang, H. Huang, F. Wei, W. Deng, F. Sun, and Q. Zhang, “Calibrating llm-based evaluator,” in LREC/COLING , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Gao, G. Xu, D. Z. Wang, and A. Cohan, “Bayesian calibration of win rate estimation with llm evaluators,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 4757–4769
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Liang, B. Zhang, J. Zhao, and K. Liu, “Abseval: An agent-based framework for script evaluation,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 12 418–12 434
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Huang, Y. Qu, J. Liu, M. Yang, and T. Zhao, “An empirical study of llm-as-a-judge for llm evaluation: Fine-tuned judge models are task-specific classifiers,” arXiv e-prints , pp. arXiv–2403, 2024
2024
Later among the works it cites.
Q. Pan, Z. Ashktorab, M. Desmond, M. S. Cooper, J. Johnson, R. Nair, E. Daly, and W. Geyer, “Human-centered design recommendations for llm-as-a-judge,” in Proceedings of the 1st Human-Centered Large Language Modeling Workshop , 2024, pp. 16–29
2024
Later among the works it cites.
S. Shankar, J. Zamfirescu-Pereira, B. Hartmann, A. Parameswaran, and I. Arawjo, “Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology , 2024, pp. 1–14
2024
Later among the works it cites.
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, L. Kong, Q. Liu, T. Liu et al. , “Large language models are not fair evaluators,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 9440–9450
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui, “Math-shepherd: Verify and reinforce llms step-by-step without human annotations,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 9426–9439
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
D. Eccleston, “Sharegpt,” 2023, accessed: 2025-01-22. [Online]. Available: https://sharegpt.com/
2025
Closest in time.
F. Wang, X. Fu, J. Y. Huang, Z. Li, Q. Liu, X. Liu, M. D. Ma, N. Xu, W. Zhou, K. Zhang et al. , “Muirbench: A comprehensive benchmark for robust multi-image understanding,” in ICLR , 2025
2025
Closest in time.
2025
Closest in time.
X. Liu, P. Li, E. Suh, Y. Vorobeychik, Z. Mao, S. Jha, P. McDaniel, H. Sun, B. Li, and C. Xiao, “Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms,” in ICLR , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Chen, J. Benton, A. Radhakrishnan, J. U. C. Denison, J. Schulman, A. Somani, P. Hase, M. W. F. R. V. Mikulik, S. Bowman, J. L. J. Kaplan et al. , “Reasoning models don’t always say what they think,” Anthropic Research , 2025
2025
Closest in time.
J. Kasai, K. Choi et al. , “REALTIMEQA: A dynamic benchmark for real-time question answering,” in Proceedings of the 2022 Annual Meeting of the Association for Computational Linguistics (ACL) , 2022, pp. 2034–2045
2045
Closest in time.