Fetching the paper…
Reading the bibliography…
We introduce PaLM 2, a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM.
Nuanced metrics for measuring unintended bias with real data for text classification, 2019
Borkan, D., Dixon, L., Sorensen, J., Thain, N., and Vasserman, L · 1903
Earlier work this paper cites.
SuperGlue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 1905
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail. corr abs/1906.05271 (2019)
Feldman, V · 1906
Earlier work this paper cites.
Prediction and entropy of printed english
Shannon, C. E · 1951
Earlier work this paper cites.
Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics, 1989
Crenshaw, K · 1989
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Kneser, R. and Ney, H · 1995
Earlier work this paper cites.
The world’s writing systems
Daniels, P. T. and Bright, W · 1996
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé, III, H., and Wallach, H · 2005
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks, 2014
Graves, A · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
Mostafazadeh, N., Chambers, N., He, X., Parikh, D., Batra, D., Vanderwende, L., Kohli, P., and Allen, J · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Fairness and machine learning limitations and opportunities
Barocas, S., Hardt, M., and Narayanan, A · 2017
Earlier work this paper cites.
Stereotype threat among girls: Differences by gender identity and math education context, 2017
Casad, B. J., Hale, P., and Wachs, F. L · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E · 2017
Earlier work this paper cites.
Sympy: symbolic computing in python
Meurer, A., Smith, C. P., Paprocki, M., Čertík, O., Kirpichev, S. B., Rocklin, M., Kumar, A., Ivanov, S., Moore, J. K., Singh, S., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Bender, E. M. and Friedman, B · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Think you have solved question answering? Try arc, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Garg, N., Schiebinger, L., Jurafsky, D., and Zou, J · 2018
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models (extended abstract), 2018
Hendricks, L. A., Burns, K., Saenko, K., Darrell, T., and Rohrbach, A · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Earlier work this paper cites.
Toxic comment classification challenge, 2018
Jigsaw · 2018
Earlier work this paper cites.
The misgendering machines: Trans/hci implications of automatic gender recognition
Keyes, O · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Dinan, E., Humeau, S., Chintagunta, B., and Weston, J · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Model cards for model reporting
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T · 2019
Earlier work this paper cites.
Fairness and abstraction in sociotechnical systems
Selbst, A. D., Boyd, D., and Friedler, S. A · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Earlier work this paper cites.
Metrology for ai: From benchmarks to instruments, 2019
Welty, C., Paritosh, P., and Aroyo, L · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Question directed graph attention network for numerical reasoning over text
Chen, K., Xu, W., Cheng, X., Xiaochuan, Z., Zhang, Y., Song, L., Wang, T., Qi, Y., and Chu, W · 2020
Earlier work this paper cites.
TyDiQA: A benchmark for information-seeking question answering in typologically diverse languages
Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J · 2020
Earlier work this paper cites.
Bringing the people back in: Contesting benchmark machine learning datasets, 2020
Denton, E., Hanna, A., Amironesei, R., Smart, A., Nicole, H., and Scheuerman, M. K · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Towards a critical race methodology in algorithmic fairness
Hanna, A., Denton, E., Smart, A., and Smith-Loud, J · 2020
Earlier work this paper cites.
A domain-specific supercomputer for training deep neural networks
Jouppi, N. P., Yoon, D. H., Kurian, G., Li, S., Patil, N., Laudon, J., Young, C., and Patterson, D · 2020
Cited alongside, same era.
WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization
Ladhak, F., Durmus, E., Cardie, C., and McKeown, K · 2020
Cited alongside, same era.
Adversarial NLI: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D · 2020
Cited alongside, same era.
XCOPA: A multilingual dataset for causal commonsense reasoning
Ponti, E. M., Glavaš, G., Majewska, O., Liu, Q., Vulić, I., and Korhonen, A · 2020
Cited alongside, same era.
Large image datasets: A pyrrhic win for computer vision?, 2020
Prabhu, V. U. and Birhane, A · 2020
Cited alongside, same era.
How much knowledge can you pack into the parameters of a language model?
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C · 2022
Later among the works it cites.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Freitag, M., Rei, R., Mathur, N., Lo, C.-k., Stewart, C., Avramidis, E., Kocmi, T., Foster, G., Lavie, A., and Martins, A. F. T · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roberts, A., Raffel, C., and Shazeer, N · 2020
Cited alongside, same era.
Social bias frames: Reasoning about social and power implications of language
Sap, M., Gabriel, S., Qin, L., Jurafsky, D., Smith, N. A., and Choi, Y · 2020
Cited alongside, same era.
Targeting the benchmark: On methodology in current natural language processing research, 2020
Schlangen, D · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A · 2020
Cited alongside, same era.
Persistent anti-muslim bias in large language models
Abid, A., Farooqi, M., and Zou, J · 2021
Cited alongside, same era.
Findings of the 2021 conference on machine translation (WMT21)
Akhbardeh, F., Arkhangorodsky, A., Biesialska, M., Bojar, O., Chatterjee, R., Chaudhary, V., Costa-jussa, M. R., España-Bonet, C., Fan, A., Federmann, C., Freitag, M., Graham, Y., Grundkiewicz, R., Haddow, B., Harter, L., Heafield, K., Homan, C., Huck, M., Amponsah-Kaakyire, K., Kasai, J., Khashabi, D., Knight, K., Kocmi, T., Koehn, P., Lourie, N., Monz, C., Morishita, M., Nagata, M., Nagesh, A., Nakazawa, T., Negri, M., Pal, S., Tapo, A. A., Turchi, M., Vydrin, V., and Zampieri, M · 2021
Cited alongside, same era.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Cited alongside, same era.
Garg, T., Masud, S., Suresh, T., and Chakraborty, T · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements, 2022
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Later among the works it cites.
Is your toxicity my toxicity? Exploring the impact of rater identity on toxicity annotation
Goyal, N., Kivlichan, I., Rosen, R., and Vasserman, L · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., et al · 2022
Later among the works it cites.
Preventing verbatim memorization in language models gives a false sense of privacy
Ippolito, D., Tramèr, F., Nasr, M., Zhang, C., Jagielski, M., Lee, K., Choquette-Choo, C. A., and Carlini, N · 2022
Later among the works it cites.
Measuring forgetting of memorized training examples
Jagielski, M., Thakkar, O., Tramer, F., Ippolito, D., Lee, K., Carlini, N., Wallace, E., Song, S., Thakurta, A., Papernot, N., et al · 2022
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Kreutzer, J., Caswell, I., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A., Subramani, N., Sokolov, A., Sikasote, C., et al · 2022
Later among the works it cites.
Welcome, singular ”they”
Lee, C · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Pax, 2022
Pax · 2022
Later among the works it cites.
Cultural incongruencies in artificial intelligence
Prabhakaran, V., Qadri, R., and Hutchinson, B · 2022
Later among the works it cites.
Data cards: Purposeful and transparent dataset documentation for responsible ai, 2022
Pushkarna, M., Zaldivar, A., and Kjartansson, O · 2022
Later among the works it cites.
How platform-user power relations shape algorithmic accountability: A case study of instant loan platforms and financially stressed users in india
Ramesh, D., Kameswaran, V., Ding, W., and Sambasivan, N · 2022
Later among the works it cites.
Characteristics of harmful text: Towards rigorous benchmarking of language models, 2022
Rauh, M., Mellor, J., Uesato, J., Huang, P.-S., Welbl, J., Weidinger, L., Dathathri, S., Glaese, A., Irving, G., Gabriel, I., Isaac, W., and Hendricks, L. A · 2022
Later among the works it cites.
Square one bias in NLP: Towards a multi-dimensional exploration of the research manifold
Ruder, S., Vulić, I., and Søgaard, A · 2022
Later among the works it cites.
Sax, 2022
Sax · 2022
Later among the works it cites.
“i’m sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset
Smith, E. M., Hall, M., Kambadur, M., Presani, E., and Williams, A · 2022
Later among the works it cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Challenging BIG-Bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Later among the works it cites.
{ML-Enhanced} code completion improves developer productivity
Tabachnyk, M. and Nikolov, S · 2022
Later among the works it cites.
LaMDA: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Writing system and speaker metadata for 2,800+ language varieties
van Esch, D., Lucassen, T., Ruder, S., Caswell, I., and Rivera, C · 2022
Later among the works it cites.
Prompting palm for translation: Assessing strategies and performance
Vilar, D., Freitag, M., Cherry, C., Luo, J., Ratnakar, V., and Foster, G · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Human parity on CommonsenseQA: Augmenting self-attention with external attention
Xu, Y., Zhu, C., Wang, S., Sun, S., Cheng, H., Liu, X., Gao, J., He, P., Zeng, M., and Huang, X · 2022
Later among the works it cites.
Natural language to code generation in interactive data science notebooks
Yin, P., Li, W.-D., Xiao, K., Rao, A., Wen, Y., Shi, K., Howland, J., Bailey, P., Catasta, M., Michalewski, H., Polozov, A., and Sutton, C · 2022
Later among the works it cites.
Guide to fair pay, 2023
Appen · 2023
Closest in time.
Our principles, 2018
Google · 2023
Closest in time.
Generative ai prohibited use policy, 2023a
Google · 2023
Closest in time.
Palm api and makersuite additional terms of service, 2023b
Google · 2023
Closest in time.
Try bard and share your feedback
Hsiao, S. and Collins, E · 2023
Closest in time.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Closest in time.
Pretraining language models with human preferences, 2023
Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C. L., Phang, J., Bowman, S. R., and Perez, E · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., and Roberts, A · 2023
Closest in time.
Coarse race data conceals disparities in clinical risk score performance, 2023
Movva, R., Shanmugam, D., Hou, K., Pathak, P., Guttag, J., Garg, N., and Pierson, E · 2023
Closest in time.
Towards agile text classifiers for everyone, 2023
Mozes, M., Hoffmann, J., Tomanek, K., Kouate, M., Thain, N., Yuan, A., Bolukbasi, T., and Dixon, L · 2023
Closest in time.
Introducing ChatGPT
OpenAI · 2023
Closest in time.
ChatGPT plugins
OpenAI · 2023
Closest in time.
Measuring the impact of programming language distribution
Orlanski, G., Xiao, K., Garcia, X., Hui, J., Howland, J., Malmaud, J., Austin, J., Singh, R., and Catasta, M · 2023
Closest in time.
On the challenges of using black-box apis for toxicity evaluation in research, 2023
Pozzobon, L., Ermis, B., Lewis, P., and Hooker, S · 2023
Closest in time.
Meet replit ghostwriter, your partner in code
Replit · 2023
Closest in time.
Frmt: A benchmark for few-shot region-aware machine translation
Riley, P., Dozat, T., Botha, J. A., Garcia, X., Garrette, D., Riesa, J., Firat, O., and Constant, N · 2023
Closest in time.
Identifying sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction, 2023
Shelby, R., Rismani, S., Henne, K., Moon, A., Rostamzadeh, N., Nicholas, P., Yilla, N., Gallegos, J., Smart, A., Garcia, E., and Virk, G · 2023
Closest in time.
Language Models are Multilingual Chain-of-Thought Reasoners
Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., Das, D., and Wei, J · 2023
Closest in time.
UL2: Unifying language learning paradigms
Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Wei, J., Wang, X., Chung, H. W., Bahri, D., Schuster, T., Zheng, S., Zhou, D., Houlsby, N., and Metzler, D · 2023
Closest in time.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Closest in time.