Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are conversational interfaces.
Human behavior and the principle of least effort: An introduction to human eoclogy
G. K. Zipf · 1949
Earlier work this paper cites.
An intensional parametric semantics for vague quantifiers
S. Lappin · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Ambiguity, accessibility, and a division of labor for communicative success
V. S. Ferreira · 2008
Earlier work this paper cites.
Semantic underspecification in language processing
S. Frisson · 2009
Earlier work this paper cites.
Quac: Question answering in context
E. Choi, H. He, M. Iyyer, M. Yatskar, W.-t. Yih, Y. Choi, P. Liang, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Conversational ai: The science behind the alexa prize
A. Ram, R. Prasad, C. Khatri, A. Venkatesh, R. Gabriel, Q. Liu, J. Nunn, B. Hedayatnia, M. Cheng, A. Nagar, et al · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
T. Yu, R. Zhang, K. Yang, M. Yasunaga, D. Wang, Z. Li, J. Ma, I. Li, Q. Yao, S. Roman, et al · 2018
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
S. Reddy, D. Chen, and C. D. Manning · 2019
Earlier work this paper cites.
Analysing concatenation approaches to document-level nmt in two different domains
Y. Scherrer, J. Tiedemann, and S. Loáiciga · 2019
Earlier work this paper cites.
Totto: A controlled table-to-text generation dataset
A. P. Parikh, X. Wang, S. Gehrmann, M. Faruqui, B. Dhingra, D. Yang, and D. Das · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Semantic evaluation for text-to-sql with distilled test suites
R. Zhong, T. Yu, and D. Klein · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Decontextualization: Making sentences stand-alone
E. Choi, J. Palomaki, M. Lamm, T. Kwiatkowski, D. Das, and M. Collins · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Survey on evaluation methods for dialogue systems
J. Deriu, A. Rodrigo, A. Otegi, G. Echegoyen, S. Rosset, E. Agirre, and M. Cieliebak · 2021
Earlier work this paper cites.
Alquist 4.0: Towards social intelligence using generative models and dialogue personalization
J. Konrád, J. Pichl, P. Marek, P. Lorenc, V. D. Ta, O. Kobza, L. Hỳlová, and J. Šedivỳ · 2021
Earlier work this paper cites.
What’s the latest? a question-driven news chatbot
P. Laban, J. Canny, and M. A. Hearst · 2021
Earlier work this paper cites.
A comparison of approaches to document-level machine translation
Z. Ma, S. Edunov, and M. Auli · 2021
Earlier work this paper cites.
Langchain, October 2022
H. Chase · 2022
Earlier work this paper cites.
Evaluating human-language model interaction
M. Lee, M. Srivastava, A. Hardy, J. Thickstun, E. Durmus, A. Paranjape, I. Gerard-Ursin, X. L. Li, F. Ladhak, F. Rong, et al · 2022
Earlier work this paper cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Y. Liu, A. R. Fabbri, P. Liu, Y. Zhao, L. Nan, R. Han, S. Han, S. Joty, C.-S. Wu, C. Xiong, et al · 2022
Earlier work this paper cites.
What does the public think about artificial intelligence?—a criticality map to understand bias in the public perception of ai
P. Brauner, A. Hick, R. Philipsen, and M. Ziefle · 2023
Earlier work this paper cites.
Don’t forget your abc’s: Evaluating the state-of-the-art in chat-oriented dialogue systems
S. E. Finch, J. D. Finch, and J. D. Choi · 2023
Earlier work this paper cites.
K.-H. Huang, P. Laban, A. R. Fabbri, P. K. Choubey, S. Joty, C. Xiong, and C.-S. Wu · 2023
Earlier work this paper cites.
Are you sure? challenging llms leads to performance drops in the flipflop experiment
P. Laban, L. Murakhovs’ ka, C. Xiong, and C.-S. Wu · 2023
Cited alongside, same era.
Instruction-following evaluation through verbalizer manipulation
S. Li, J. Yan, H. Wang, Z. Tang, X. Ren, V. Srinivasan, and H. Jin · 2023
Cited alongside, same era.
We’re afraid language models aren’t modeling ambiguity
A. Liu, Z. Wu, J. Michael, A. Suhr, P. West, A. Koller, S. Swayamdipta, N. A. Smith, and Y. Choi · 2023
Cited alongside, same era.
L. Murakhovs’ ka, P. Laban, T. Xie, C. Xiong, and C.-S. Wu · 2023
Cited alongside, same era.
Artificial intelligence (ai) trust framework and maturity model: Applying an entropy lens to improve security, privacy, and ethical ai
Mt-eval: A multi-turn capabilities evaluation benchmark for large language models
W.-C. Kwan, X. Zeng, Y. Jiang, Y. Wang, L. Li, L. Shang, X. Jiang, Q. Liu, and K.-F. Wong · 2024
Later among the works it cites.
Summary of a haystack: A challenge to long-context llms and rag systems
P. Laban, A. R. Fabbri, C. Xiong, and C.-S. Wu · 2024
Later among the works it cites.
One vs. many: Comprehending accurate information from multiple erroneous and inconsistent ai generations
Y. Lee, K. Son, T. S. Kim, J. Kim, J. J. Y. Chung, E. Adar, and J. Kim · 2024
Later among the works it cites.
Spider 2.0: Evaluating language models on real-world enterprise text-to-sql workflows
F. Lei, J. Chen, Y. Ye, R. Cao, D. Shin, H. Su, Z. Suo, H. Gao, W. Hu, P. Yin, et al · 2024
Later among the works it cites.
Iqa-eval: Automatic evaluation of human-model interactive question answering
R. Li, R. Li, B. Wang, and X. Du · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Mylrea and N. Robinson · 2023
Cited alongside, same era.
When does in-context learning fall short and why? a study on specification-heavy tasks
H. Peng, X. Wang, J. Chen, W. Li, Y. P. Qi, Z. Wang, Z. Wu, K. Zeng, B. Xu, L. Hou, and J. Li · 2023
Cited alongside, same era.
Dealing with semantic underspecification in multimodal nlp
S. Pezzelle · 2023
Cited alongside, same era.
Escaping the sentence-level paradigm in machine translation
M. Post and M. Junczys-Dowmunt · 2023
Cited alongside, same era.
Developing a model for ai across the curriculum: Transforming the higher education landscape via innovation in ai literacy
J. Southworth, K. Migliaccio, J. Glover, J. Glover, D. Reed, C. McCarty, J. Brendemuhl, and A. Thomas · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Cited alongside, same era.
Interactive ai alignment: specification, process, and evaluation alignment
M. Terry, C. Kulkarni, M. Wattenberg, L. Dixon, and M. R. Morris · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Mathchat: Benchmarking mathematical reasoning and instruction following in multi-turn interactions
Z. Liang, D. Yu, W. Yu, W. Yao, Z. Zhang, X. Zhang, and D. Yu · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2024
Later among the works it cites.
Contextualized evaluations: Taking the guesswork out of language model evaluations
C. Malaviya, J. C. Chang, D. Roth, M. Iyyer, M. Yatskar, and K. Lo · 2024
Later among the works it cites.
T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan, et al · 2024
Later among the works it cites.
Parrot: Enhancing multi-turn instruction following for large language models
Y. Sun, C. Liu, K. Zhou, J. Huang, R. Song, W. X. Zhao, F. Zhang, D. Zhang, and K. Gai · 2024
Later among the works it cites.
Search engines in an ai era: The false promise of factual and verifiable source-cited responses
P. N. Venkit, P. Laban, Y. Zhou, Y. Mao, and C.-S. Wu · 2024
Later among the works it cites.
Mint: Evaluating llms in multi-turn interaction with tools and language feedback
X. Wang, Z. Wang, J. Liu, Y. Chen, L. Yuan, H. Peng, and H. Ji · 2024
Later among the works it cites.
Design principles for generative ai applications
J. D. Weisz, J. He, M. Muller, G. Hoefer, R. Miles, and W. Geyer · 2024
Later among the works it cites.
“as an ai language model, i cannot”: Investigating llm denials of user requests
J. Wester, T. Schrills, H. Pohl, and N. van Berkel · 2024
Later among the works it cites.
Do pre-trained language models detect and understand semantic underspecification? ask the dust!
F. Wildenburg, M. Hanna, and S. Pezzelle · 2024
Later among the works it cites.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Later among the works it cites.
T. Chakrabarty, P. Laban, and C.-S. Wu · 2025
Closest in time.
Chatbench: From static benchmarks to human-ai evaluation
S. Chang, A. Anderson, and J. M. Hofman · 2025
Closest in time.
Command a: An enterprise-ready large language model
T. Cohere, A. Ahmadian, M. Ahmed, J. Alammar, Y. Alnumay, S. Althammer, A. Arkhangorodsky, V. Aryabumi, D. Aumiller, R. Avalos, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Can language models follow multiple turns of entangled instructions?
C. Han · 2025
Closest in time.
Which economic tasks are performed with ai? evidence from millions of claude conversations
K. Handa, A. Tamkin, M. McCain, S. Huang, E. Durmus, S. Heck, J. Mueller, J. Hong, S. Ritchie, T. Belonax, et al · 2025
Closest in time.
OpenAI o3 and o4-mini System Card — openai.com
OpenAI · 2025
Closest in time.
L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al · 2025
Closest in time.
Synthetic clarification and correction dialogues about data-centric tasks–a teacher-student approach
C. Poelitz and N. McKenna · 2025
Closest in time.
R. Sarkar, B. Sarrafzadeh, N. Chandrasekaran, N. Rangan, P. Resnik, L. Yang, and S. K. Jauhar · 2025
Closest in time.
Navigating rifts in human-llm grounding: Study and benchmark
O. Shaikh, H. Mozannar, G. Bansal, A. Fourney, and E. Horvitz · 2025
Closest in time.
V. Sirdeshmukh, K. Deshpande, J. Mols, L. Jin, E.-Y. Cardona, D. Lee, J. Kritz, W. Primack, S. Yue, and C. Xing · 2025
Closest in time.
Interactive agents to overcome ambiguity in software engineering
S. Vijayvargiya, X. Zhou, A. Yerukola, M. Sap, and G. Neubig · 2025
Closest in time.