Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) went from non-existent to ubiquitous in the machine learning discourse within a few years.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam et al. 2020 · 1901
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier and M. Auli. 2019 · 1904
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger and Y. Artzi. 2019 · 1904
Earlier work this paper cites.
Real or Fake? Learning to Discriminate Machine from Human Generated Text
A. Bakhtin, S. Gross, M. Ott, Y. Deng, M. Ranzato and A. Szlam. 2019 · 1906
Earlier work this paper cites.
Neural legal judgment prediction in english
I. Chalkidis, I. Androutsopoulos and N. Aletras. 2019 · 1906
Earlier work this paper cites.
GLTR: Statistical Detection and Visualization of Generated Text
S. Gehrmann, H. Strobelt and A. M. Rush. 2019 · 1906
Earlier work this paper cites.
R. Schwartz, J. Dodge, N. A. Smith and O. Etzioni. 2019 · 1907
Earlier work this paper cites.
Bridging the Gap for Tokenizer-Free Language Models
D. Choe, R. Al-Rfou, M. Guo, H. Lee and N. Constant. 2019 · 1908
Earlier work this paper cites.
Neural text generation with unlikelihood training
S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho and J. Weston. 2019 · 1908
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong and R. Socher. 2019 · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper and B. Catanzaro. 2019 · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano and G. Irving. 2019 · 1909
Earlier work this paper cites.
Sticking to the Facts: Confident Decoding for Faithful Data-to-Text Generation
R. Tian, S. Narayan, T. Sellam and A. P. Parikh. 2020 · 1910
Earlier work this paper cites.
Structured pruning of large language models
Z. Wang, J. Wohlwend and T. Lei. 2019a · 1910
Earlier work this paper cites.
High-dimensional bayesian optimization using low-dimensional feature spaces
R. Moriconi, M. P. Deisenroth and K. Sesh Kumar. 2020 · 1943
Earlier work this paper cites.
“cloze procedure”: A new tool for measuring readability
W. L. Taylor. 1953 · 1953
Earlier work this paper cites.
Behavioral study of obedience
S. Milgram. 1963 · 1963
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal and C. A. Raffel. 2022b · 1965
Earlier work this paper cites.
Creativity: Yesterday, today and tomorrow
J. P. Guilford. 1967 · 1967
Earlier work this paper cites.
The “false consensus effect”: An egocentric bias in social perception and attribution processes
L. Ross, D. Greene and P. House. 1977 · 1977
Earlier work this paper cites.
Continuous automated model evaluation (cameo)—perspectives on the future of fully automated evaluation of structure prediction methods
X. Robin, J. Haas, R. Gumienny, A. Smolinski, G. Tauriello and T. Schwede. 2021 · 1986
Earlier work this paper cites.
What every computer scientist should know about floating-point arithmetic
D. Goldberg. 1991 · 1991
Earlier work this paper cites.
Personality trait structure as a human universal
R. R. McCrae and P. T. Costa Jr. 1997 · 1997
Earlier work this paper cites.
Min-wise independent permutations
A. Z. Broder, M. Charikar, A. M. Frieze and M. Mitzenmacher. 1998 · 1998
Earlier work this paper cites.
Towards a human-like open-domain chatbot
D. Adiwardana, M.-T. Luong, D. R. So, J. Hall, N. Fiedel, R. Thoppilan, Z. Yang, A. Kulshreshtha et al. 2020 · 2001
Earlier work this paper cites.
Thematic roles assigned along the garden path linger
K. Christianson, A. Hollingworth, J. F. Halliwell and F. Ferreira. 2001 · 2001
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford et al. 2020 · 2001
Earlier work this paper cites.
Money, kisses, and electric shocks: On the affective psychology of risk
Y. Rottenstreich and C. K. Hsee. 2001 · 2001
Earlier work this paper cites.
A dataset for statutory reasoning in tax law entailment and question answering
N. Holzenberger, A. Blair-Stanek and B. Van Durme. 2020 · 2005
Earlier work this paper cites.
Why Most Published Research Findings Are False
J. P. A. Ioannidis. 2005 · 2005
Earlier work this paper cites.
Sparse GPU Kernels for Deep Learning
T. Gale, M. Zaharia, C. Young and E. Elsen. 2020 · 2006
Earlier work this paper cites.
A. Elnaggar, M. Heinzinger, C. Dallago, G. Rihawi, Y. Wang, L. Jones, T. Gibbs, T. Feher et al. 2020 · 2007
Earlier work this paper cites.
Linear attention mechanism: An efficient attention for semantic segmentation
R. Li, J. Su, C. Duan and S. Zheng. 2020 · 2007
Earlier work this paper cites.
Practical and sample efficient zero-shot hpo
F. Winkelmolen, N. Ivkin, H. F. Bozkurt and Z. Karnin. 2020 · 2007
Earlier work this paper cites.
Aligning ai with shared human values
D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song and J. Steinhardt. 2020 · 2008
Earlier work this paper cites.
The hexaco–60: A short measure of the major dimensions of personality
M. C. Ashton and K. Lee. 2009 · 2009
Earlier work this paper cites.
Rethinking attention with performers
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis et al. 2020 · 2009
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi and N. A. Smith. 2020 · 2009
Earlier work this paper cites.
Lingering misinterpretations in garden-path sentences: evidence from a paraphrasing task
N. D. Patson, E. S. Darowski, N. Moon and F. Ferreira. 2009 · 2009
Earlier work this paper cites.
Legal-bert: The muppets straight out of law school
I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras and I. Androutsopoulos. 2020 · 2010
Earlier work this paper cites.
Limitations of autoregressive models and their alternatives
C.-C. Lin, A. Jaech, X. Li, M. R. Gormley and J. Eisner. 2020 · 2010
Earlier work this paper cites.
Red teaming with coevolution
P. Hingston and M. Preuss. 2011 · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright and F. Niu. 2011 · 2011
Earlier work this paper cites.
Introduction to Montague semantics , volume 11
D. R. Dowty, R. Wall and S. Peters. 2012 · 2012
Earlier work this paper cites.
Japanese and korean voice search
M. Schuster and K. Nakajima. 2012 · 2012
Earlier work this paper cites.
Social influence and the collective dynamics of opinion formation
M. Moussaïd, J. E. Kämmer, P. P. Analytis and H. Neth. 2013 · 2013
Earlier work this paper cites.
Bayesian optimization in high dimensions via random embeddings
Z. Wang, M. Zoghi, F. Hutter, D. Matheson, N. De Freitas et al. 2013 · 2013
Earlier work this paper cites.
Experimental economics and experimental game theory
D. Houser and K. McCabe. 2014 · 2014
Earlier work this paper cites.
Human values scale (ess)
S. H. Schwartz, B. Breyer and D. Danner. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow and A. Birch. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba and S. Fidler. 2015 · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao et al. 2016 · 2016
Earlier work this paper cites.
Efficient attention using a fixed-size memory representation
D. Britz, M. Y. Guan and M.-T. Luong. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg and D. Amodei. 2017 · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. Patwary, M. Ali et al. 2017 · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
W. Ling, D. Yogatama, C. Dyer and P. Blunsom. 2017 · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
J. Peters, D. Janzing and B. Schölkopf. 2017 · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
A. Radford, R. Jozefowicz and I. Sutskever. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford and O. Klimov. 2017 · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton and J. Dean. 2017 · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying and Q. V. Le. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser and I. Polosukhin. 2017 · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
P. Bajaj, D. Campos, N. Craswell, L. Deng, J. Gao, X. Liu, R. Majumder, A. McNamara et al. 2018 · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre et al. 2018 · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
J. Howard and S. Ruder. 2018 · 2018
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam et al. 2018 · 2018
Earlier work this paper cites.
G. Irving, P. Christiano and D. Amodei. 2018 · 2018
Earlier work this paper cites.
Many labs 2: Investigating variation in replicability across samples and settings
R. A. Klein, M. Vianello, F. Hasselman, B. G. Adams, R. B. Adams Jr, S. Alper, M. Aveyard, J. R. Axt et al. 2018 · 2018
Earlier work this paper cites.
Introduction to reasoning
D. C. Krawczyk. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
T. Kudo. 2018 · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson. 2018 · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
K. Lee, O. Firat, A. Agarwal, C. Fannjiang and D. Sussillo. 2018 · 2018
Earlier work this paper cites.
Object hallucination in image captioning
A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell and K. Saenko. 2018 · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw, J. Uszkoreit and A. Vaswani. 2018 · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
M. Stern, N. Shazeer and J. Uszkoreit. 2018 · 2018
Earlier work this paper cites.
Diverse beam search for improved description of complex scenes
A. Vijayakumar, M. Cogswell, R. Selvaraju, Q. Sun, S. Lee, D. Crandall and D. Batra. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy and S. Bowman. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos and D. Song. 2019 · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le and R. Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee and K. Toutanova. 2019 · 2019
Earlier work this paper cites.
Efficient training of BERT by progressively stacking
L. Gong, D. He, Z. Li, T. Qin, L. Wang and T. Liu. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan and S. Gelly. 2019 · 2019
Earlier work this paper cites.
Music transformer
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman et al. 2019 · 2019
Earlier work this paper cites.
Metalearners for estimating heterogeneous treatment effects using machine learning
S. R. Künzel, J. S. Sekhon, P. J. Bickel and B. Yu. 2019 · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
P. J. Ortiz Su’arez, B. Sagot and L. Romary. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei and I. Sutskever. 2019 · 2019
Earlier work this paper cites.
Do massively pretrained language models make better storytellers?
A. See, A. Pappu, R. Saxena, A. Yerukola and C. D. Manning. 2019 · 2019
Earlier work this paper cites.
Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions
M. Weiss. 2019 · 2019
Earlier work this paper cites.
Towards coherent and engaging spoken dialog response generation using automatic conversation evaluators
S. Yi, R. Goel, C. Khatri, A. Cervone, T. Chung, B. Hedayatnia, A. Venkatesh, R. Gabriel et al. 2019 · 2019
Earlier work this paper cites.
Defending against neural fake news
R. Zellers, A. Holtzman, H. Rashkin, Y. Bisk, A. Farhadi, F. Roesner and Y. Choi. 2019 · 2019
Earlier work this paper cites.
The effects of populism as a social identity frame on persuasion and mobilisation: Evidence from a 15-country experiment
L. Bos, C. Schemer, N. Corbu, M. Hameleers, I. Andreadis, A. Schulz, D. Schmuck, C. Reinemann et al. 2020 · 2020
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown et al. 2020 · 2020
Earlier work this paper cites.
Improving multilingual models with language-clustered vocabularies
H. W. Chung, D. Garrette, K. C. Tan and J. Riesa. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott et al. 2020 · 2020
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
S. Dathathri, A. Madotto, J. Lan, J. Hung, E. Frank, P. Molino, J. Yosinski and R. Liu. 2020 · 2020
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave and A. Joulin. 2020 · 2020
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
V. Feldman. 2020 · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
I. Gabriel. 2020 · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He et al. 2020 · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge and F. A. Wichmann. 2020 · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
K. Guu, K. Lee, Z. Tung, P. Pasupat and M. Chang. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes and Y. Choi. 2020 · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang and Q. Liu. 2020 · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer et al. 2020 · 2020
Earlier work this paper cites.
Legal language modeling with transformers
L. Peric, S. Mijic, D. Stammbach and E. Ash. 2020 · 2020
Earlier work this paper cites.
AdapterHub: A framework for adapting transformers
J. Pfeiffer, A. Rücklé, C. Poth, A. Kamath, I. Vulić, S. Ruder, K. Cho and I. Gurevych. 2020 · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase and Y. He. 2020 · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase and Y. He. 2020 · 2020
Earlier work this paper cites.
Xla : Compiling machine learning for peak performance
A. Sabne. 2020 · 2020
Earlier work this paper cites.
S. Sagawa, P. W. Koh, T. B. Hashimoto and P. Liang. 2020 · 2020
Earlier work this paper cites.
The limitations of stylometry for detecting machine-generated fake news
T. Schuster, R. Schuster, D. J. Shah and R. Barzilay. 2020 · 2020
Earlier work this paper cites.
EDITABLE NEURAL NETWORKS
A. Sinitsin, D. Pyrkin, A. Babenko, V. Plokhotnyuk and S. Popov. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei et al. 2020 · 2020
Earlier work this paper cites.
Authorship Attribution for Neural Text Generation
A. Uchendu, T. Le, K. Shu and D. Lee. 2020 · 2020
Earlier work this paper cites.
Neural machine translation with byte-level subwords
C. Wang, K. Cho and J. Gu. 2020 · 2020
Earlier work this paper cites.
Accelerating training of transformer-based language models with progressive layer dropping
M. Zhang and Y. He. 2020 · 2020
Earlier work this paper cites.
Deep Reinforcement Learning at the Edge of the Statistical Precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville and M. Bellemare. 2021 · 2021
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai et al. 2021 · 2021
Earlier work this paper cites.
Learning in high dimension always amounts to extrapolation
R. Balestriero, J. Pesenti and Y. LeCun. 2021 · 2021
Earlier work this paper cites.
Addressing "documentation debt" in machine learning research: A retrospective datasheet for bookcorpus
J. Bandy and N. Vincent. 2021 · 2021
Earlier work this paper cites.
Pitfalls in machine learning research: Reexamining the development cycle
S. Biderman and W. J. Scheirer. 2021 · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
A. Birhane, V. U. Prabhu and E. Kahembwe. 2021 · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. v. d. Driessche, J.-B. Lespiau et al. 2021 · 2021
Earlier work this paper cites.
When is memorization of irrelevant training data necessary for high-accuracy learning?
G. Brown, M. Bun, V. Feldman, A. Smith and K. Talwar. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda et al. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek et al. 2021 · 2021
Earlier work this paper cites.
A human being wrote this law review article: Gpt-3 and the practice of law
A. B. Cyphert. 2021 · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
N. De Cao, W. Aziz and I. Titov. 2021 · 2021
Earlier work this paper cites.
M. Dehghani, Y. Tay, A. A. Gritsenko, Z. Zhao, N. Houlsby, F. Diaz, D. Metzler and O. Vinyals. 2021 · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
J. Dodge, M. Sap, A. Marasović, W. Agnew, G. Ilharco, D. Groeneveld, M. Mitchell and M. Gardner. 2021 · 2021
Earlier work this paper cites.
Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding
N. Dziri, A. Madotto, O. Zaiane and A. J. Bose. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai et al. 2021 · 2021
Earlier work this paper cites.
Augmenting transformers with KNN-based composite memory for dialog
A. Fan, C. Gardent, C. Braud and A. Bordes. 2021 · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph and N. Shazeer. 2021 · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu et al. 2021 · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Naumann, J. Gao et al. 2021 · 2021
Earlier work this paper cites.
The hardware lottery
S. Hooker. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang and W. Chen. 2021 · 2021
Earlier work this paper cites.
Accelerated sparse neural training: A provable and efficient method to find n:m transposable masks
I. Hubara, B. Chmiel, M. Island, R. Banner, J. Naor and D. Soudry. 2021 · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates et al. 2021 · 2021
Earlier work this paper cites.
Causal Effect Inference for Structured Treatments
J. Kaddour, Y. Zhu, Q. Liu, M. J. Kusner and R. Silva. 2021 · 2021
Earlier work this paper cites.
nmT5 - is parallel data still relevant for pre-training massively multilingual language models?
M. Kale, A. Siddhant, R. Al-Rfou, L. Xue, N. Constant and M. Johnson. 2021 · 2021
Earlier work this paper cites.
Z. Kenton, T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik and G. Irving. 2021 · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in nlp
D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad et al. 2021 · 2021
Earlier work this paper cites.
Target classification in the 14th round of the critical assessment of protein structure prediction (casp14)
L. N. Kinch, R. D. Schaeffer, A. Kryshtafovych and N. V. Grishin. 2021 · 2021
Earlier work this paper cites.
Considering the possibilities and pitfalls of generative pre-trained transformer 3 (gpt-3) in healthcare delivery
D. M. Korngiebel and S. D. Mooney. 2021 · 2021
Earlier work this paper cites.
GeDi: Generative discriminator guided sequence generation
B. Krause, A. D. Gotmare, B. McCann, N. S. Keskar, S. Joty, R. Socher and N. F. Rajani. 2021 · 2021
Earlier work this paper cites.
GitHub Copilot AI Is Leaking Functional API Keys
A. Kulkarni. 2021 · 2021
Earlier work this paper cites.
Deduplicating training data makes language models better
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch and N. Carlini. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou and N. Constant. 2021 · 2021
Earlier work this paper cites.
Base layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal and L. Zettlemoyer. 2021 · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang. 2021 · 2021
Earlier work this paper cites.
Towards understanding and mitigating social biases in language models
P. P. Liang, C. Wu, L.-P. Morency and R. Salakhutdinov. 2021 · 2021
Earlier work this paper cites.
Jurassic-1: Technical details and evaluation
O. Lieber, O. Sharir, B. Lenz and Y. Shoham. 2021 · 2021
Earlier work this paper cites.
Synthesizing natural language to visualization (nl2vis) benchmarks from nl2sql benchmarks
Y. Luo, N. Tang, G. Li, C. Chai, W. Li and X. Qin. 2021 · 2021
Earlier work this paper cites.
Luna: Linear unified nested attention
X. Ma, X. Kong, S. Wang, C. Zhou, J. May, H. Ma and L. Zettlemoyer. 2021 · 2021
Earlier work this paper cites.
Accelerating sparse deep neural networks
A. Mishra, J. A. Latorre, J. Pool, D. Stosic, D. Stosic, G. Venkatesh, C. Yu and P. Micikevicius. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain et al. 2021 · 2021
Earlier work this paper cites.
Quasi-oracle estimation of heterogeneous treatment effects
X. Nie and S. Wager. 2021 · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz et al. 2021 · 2021
Earlier work this paper cites.
Probing toxic content in large pre-trained language models
N. Ousidhoum, X. Zhao, T. Fang, Y. Song and D.-Y. Yeung. 2021 · 2021
Earlier work this paper cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
A. Paullada, I. D. Raji, E. M. Bender, E. Denton and A. Hanna. 2021 · 2021
Earlier work this paper cites.
A MARVS analysis of two Chinese near-synonymous verbs of jumping based on Chinese corpora
Y. Peng. 2021 · 2021
Earlier work this paper cites.
Smoothing and shrinking the sparse Seq2Seq search space
B. Peters and A. F. T. Martins. 2021 · 2021
Earlier work this paper cites.
On releasing annotator-level labels and information in datasets
V. Prabhakaran, A. Mostafazadeh Davani and M. Diaz. 2021 · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. A. Smith and M. Lewis. 2021 · 2021
Earlier work this paper cites.
Overview and discussion of the competition on legal information Extraction/Entailment (COLIEE) 2021
J. Rabelo, R. Goebel, M.-Y. Kim, Y. Kano, M. Yoshioka and K. Satoh. 2022 · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson et al. 2021 · 2021
Earlier work this paper cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith and Y. He. 2021 · 2021
Earlier work this paper cites.
Ai and the everything in the whole wide world benchmark
I. D. Raji, E. M. Bender, A. Paullada, E. Denton and A. Hanna. 2021 · 2021
Earlier work this paper cites.
{ \{ ZeRO-Offload } \} : Democratizing { \{ Billion-Scale } \} model training
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li and Y. He. 2021 · 2021
Earlier work this paper cites.
Hash layers for large sparse models
S. Roller, S. Sukhbaatar, A. Szlam and J. Weston. 2021 · 2021
Earlier work this paper cites.
Human-compatible artificial intelligence
S. Russell. 2021 · 2021
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
T. Schick and H. Schütze. 2021 · 2021
Earlier work this paper cites.
Efficient attention: Attention with linear complexities
Z. Shen, M. Zhang, H. Zhao, S. Yi and H. Li. 2021 · 2021
Earlier work this paper cites.
Generative language modeling for antibody design
R. W. Shuai, J. A. Ruffolo and J. J. Gray. 2021 · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
K. Shuster, S. Poff, M. Chen, D. Kiela and J. Weston. 2021 · 2021
Cited alongside, same era.
Process for adapting language models to society (palms) with values-targeted datasets
I. Solaiman and C. Dennison. 2021 · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
J. Su, Y. Lu, S. Pan, B. Wen and Y. Liu. 2021 · 2021
Cited alongside, same era.
1-bit adam: Communication efficient large-scale training with adam’s convergence speed
H. Tang, S. Gan, A. A. Awan, S. Rajbhandari, C. Li, X. Lian, J. Liu, C. Zhang et al. 2021 · 2021
Cited alongside, same era.
Synthesizer: Rethinking self-attention for transformer models
Y. Tay, D. Bahri, D. Metzler, D.-C. Juan, Z. Zhao and C. Zheng. 2021 · 2021
Cited alongside, same era.
Limits for Learning with Language Models
N. Asher, S. Bhar, A. Chaturvedi, J. Hunter and S. Paul. 2023 · 2023
Closest in time.
Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum
J. W. Ayers, A. Poliak, M. Dredze, E. C. Leas, Z. Zhu, J. B. Kelley, D. J. Faix, A. M. Goodman et al. 2023 · 2023
Closest in time.
Eliciting latent predictions from transformers with the tuned lens
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman and J. Steinhardt. 2023 · 2023
Closest in time.
[…] we aren’t running out of text data any time soon. ml researchers massively underestimate how much text is out there
S. R. Biderman. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding emails and drafting responses–an approach using gpt-3
J. Thiergart, S. Huber and T. Übellacker. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
B. Wang and A. Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Want to reduce labeling cost? gpt-3 can help
S. Wang, Y. Liu, Y. Xu, C. Zhu and M. Zeng. 2021 · 2021
Cited alongside, same era.
Comment section personalization: Algorithmic, interface, and interaction design
Y. Wang. 2021 · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese et al. 2021 · 2021
Cited alongside, same era.
On Hallucination and Predictive Uncertainty in Conditional Language Generation
Y. Xiao and W. Y. Wang. 2021 · 2021
Cited alongside, same era.
Gspmd: general and scalable parallelization for ml computation graphs
Y. Xu, H. Lee, D. Chen, B. Hechtman, Y. Huang, R. Joshi, M. Krikun, D. Lepikhin et al. 2021 · 2021
Cited alongside, same era.
A. Blair-Stanek, N. Holzenberger and B. Van Durme. 2023 · 2023
Closest in time.
Gpt as knowledge worker: A zero-shot evaluation of (ai) cpa capabilities
J. Bommarito, M. Bommarito, D. M. Katz and J. Katz. 2023 · 2023
Closest in time.
A Categorical Archive of ChatGPT Failures
A. Borji. 2023 · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee et al. 2023 · 2023
Closest in time.
Explore, establish, exploit: Red teaming language models from scratch
S. Casper, J. Lin, J. Kwon, G. Culp and D. Hadfield-Menell. 2023 · 2023
Closest in time.
A Survey on Evaluation of Large Language Models
Y. Chang, X. Wang, J. Wang, Y. Wu, K. Zhu, H. Chen, L. Yang, X. Yi et al. 2023 · 2023
Closest in time.
L. Cheng, X. Li and L. Bing. 2023 · 2023
Closest in time.
Chatgpt goes to law school
J. H. Choi, K. E. Hickman, A. Monahan and D. Schwarcz. 2023 · 2023
Closest in time.
Undetectable Watermarks for Language Models
M. Christ, S. Gunn and O. Zamir. 2023 · 2023
Closest in time.
Missing model details (tweet)
H. W. Chung. 2023 · 2023
Closest in time.
Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining
H. W. Chung, X. Garcia, A. Roberts, Y. Tay, O. Firat, S. Narang and N. Constant. 2023 · 2023
Closest in time.
LM vs LM: Detecting Factual Errors via Cross Examination
R. Cohen, M. Hamri, M. Geva and A. Globerson. 2023 · 2023
Closest in time.
Redpajama: An open source recipe to reproduce llama training dataset
T. Computer. 2023 · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim and A. Garriga-Alonso. 2023 · 2023
Closest in time.
Chataug: Leveraging chatgpt for text data augmentation
H. Dai, Z. Liu, W. Liao, X. Huang, Z. Wu, L. Zhao, W. Liu, N. Liu et al. 2023 · 2023
Closest in time.
The nucleotide transformer: Building and evaluating robust foundation models for human genomics
H. Dalla-Torre, L. Gonzalez, J. Mendoza Revilla, N. Lopez Carranza, A. Henryk Grywaczewski, F. Oteri, C. Dallago, E. Trop et al. 2023 · 2023
Closest in time.
Hungry Hungry Hippos: Towards Language Modeling with State Space Models
T. Dao, D. Y. Fu, K. K. Saab, A. W. Thomas, A. Rudra and C. Ré. 2023 · 2023
Closest in time.
SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
L. Del Corro, A. Del Giorno, S. Agarwal, B. Yu, A. Awadallah and S. Mukherjee. 2023 · 2023
Closest in time.
How ready are pre-trained abstractive models and llms for legal case judgement summarization?
A. Deroy, K. Ghosh and S. Ghosh. 2023 · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
A. Deshpande, V. Murahari, T. Rajpurohit, A. Kalyan and K. Narasimhan. 2023 · 2023
Closest in time.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
N. Dey, G. Gosal, Zhiming, Chen, H. Khachane, W. Marshall, R. Pathria, M. Tom et al. 2023 · 2023
Closest in time.
Longnet: Scaling transformers to 1,000,000,000 tokens
J. Ding, S. Ma, L. Dong, X. Zhang, S. Huang, W. Wang and F. Wei. 2023 · 2023
Closest in time.
Questioning the survey responses of large language models
R. Dominguez-Olmedo, M. Hardt and C. Mendler-Dünner. 2023 · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson et al. 2023 · 2023
Closest in time.
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum and I. Mordatch. 2023 · 2023
Closest in time.
Analysis of large-language model versus human performance for genetics questions
D. Duong and B. D. Solomon. 2023 · 2023
Closest in time.
Faith and Fate: Limits of Transformers on Compositionality
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, P. West, C. Bhagavatula et al. 2023 · 2023
Closest in time.
On the Impossible Safety of Large AI Models
E.-M. El-Mhamdi, S. Farhadkhani, R. Guerraoui, N. Gupta, L.-N. Hoang, R. Pinot, S. Rouault and J. Stephan. 2023 · 2023
Closest in time.
Gpts are gpts: An early look at the labor market impact potential of large language models
T. Eloundou, S. Manning, P. Mishkin and D. Rock. 2023 · 2023
Closest in time.
Reward modeling for mitigating toxicity in transformer-based language models
F. Faal, K. Schmitt and J. Y. Yu. 2023 · 2023
Closest in time.
M. Fathi, J. Pilault, P.-L. Bacon, C. Pal, O. Firat and R. Goroshin. 2023 · 2023
Closest in time.
Should chatgpt be biased? challenges and risks of bias in large language models
E. Ferrara. 2023 · 2023
Closest in time.
What’s going on with the open llm leaderboard?
C. Fourrier, N. Habib, J. Launay and T. Wolf. 2023 · 2023
Closest in time.
Massive language models can be accurately pruned in one-shot
E. Frantar and D. Alistarh. 2023 · 2023
Closest in time.
Resolving code review comments with ml
A. Frömmgen and L. Kharatyan. 2023 · 2023
Closest in time.
Gptscore: Evaluate as you desire
J. Fu, S.-K. Ng, Z. Jiang and P. Liu. 2023 · 2023
Closest in time.
T. Fujii, K. Shibata, A. Yamaguchi, T. Morishita and Y. Sogawa. 2023 · 2023
Closest in time.
Understanding social reasoning in language models with language models
K. Gandhi, J.-P. Fränken, T. Gerstenbrg and N. D. Goodman. 2023 · 2023
Closest in time.
Is chatgpt a good causal reasoner? a comprehensive evaluation
J. Gao, X. Ding, B. Qin and T. Liu. 2023 · 2023
Closest in time.
Critic: Large language models can self-correct with tool-interactive critiquing
Z. Gou, Z. Shao, Y. Gong, Y. Shen, Y. Yang, N. Duan and W. Chen. 2023 · 2023
Closest in time.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz and M. Fritz. 2023 · 2023
Closest in time.
Susceptibility to influence of large language models
L. D. Griffin, B. Kleinberg, M. Mozes, K. T. Mai, M. Vau, M. Caldwell and A. Marvor-Parker. 2023 · 2023
Closest in time.
Y. Gu, S. Zhang, N. Usuyama, Y. Woldesenbet, C. Wong, P. Sanapathi, M. Wei, N. Valluri et al. 2023 · 2023
Closest in time.
The false promise of imitating proprietary llms
A. Gudibande, E. Wallace, C. Snell, X. Geng, H. Liu, P. Abbeel, S. Levine and D. Song. 2023 · 2023
Closest in time.
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann et al. 2023 · 2023
Closest in time.
Probing Quantifier Comprehension in Large Language Models
A. Gupta. 2023 · 2023
Closest in time.
Artificial muses: Generative artificial intelligence chatbots have risen to human-level creativity
J. Haase and P. H. P. Hanel. 2023 · 2023
Closest in time.
A theory of emergent in-context learning as implicit structure induction
M. Hahn and N. Goyal. 2023 · 2023
Closest in time.
Blind judgement: Agent-based supreme court modelling with gpt
S. Hamilton. 2023 · 2023
Closest in time.
In-context learning of large language models explained as kernel regression
C. Han, Z. Wang, H. Zhao and H. Ji. 2023 · 2023
Closest in time.
Large language models can be used to effectively scale spear phishing campaigns
J. Hazell. 2023 · 2023
Closest in time.
Efficient evolution of human antibodies from general protein language models
B. L. Hie, V. R. Shanker, D. Xu, T. U. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak et al. 2023 · 2023
Closest in time.
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
J. Hoelscher-Obermaier, J. Persson, E. Kran, I. Konstas and F. Barez. 2023 · 2023
Closest in time.
Large language models as simulated economic agents: What can we learn from homo silicus?
J. J. Horton. 2023 · 2023
Closest in time.
Bytes Are All You Need: Transformers Operating Directly On File Bytes
M. Horton, S. Mehta, A. Farhadi and M. Rastegari. 2023 · 2023
Closest in time.
What’s ahead for bard: More global, more visual, more integrated
S. Hsiao. 2023 · 2023
Closest in time.
Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Z. Hu, Y. Lan, L. Wang, W. Xu, E.-P. Lim, R. K.-W. Lee, L. Bing and S. Poria. 2023 · 2023
Closest in time.
Towards Reasoning in Large Language Models: A Survey
J. Huang and K. C.-C. Chang. 2023 · 2023
Closest in time.
Transformer-patcher: One mistake worth one neuron
Z. Huang, Y. Shen, X. Zhang, J. Zhou, W. Rong and Z. Xiong. 2023 · 2023
Closest in time.
Huggingchat v0.3.0
HuggingFace. 2023 · 2023
Closest in time.
Chatgpt by openai: The end of litigation lawyers?
K. Y. Iu and V. M.-Y. Wong. 2023 · 2023
Closest in time.
A. Jacovi, A. Caciularu, O. Goldman and Y. Goldberg. 2023 · 2023
Closest in time.
Bring your own data! self-supervised evaluation for large language models
N. Jain, K. Saifullah, Y. Wen, J. Kirchenbauer, M. Shu, A. Saha, M. Goldblum, J. Geiping et al. 2023 · 2023
Closest in time.
Exploring the Benefits of Training Expert Language Models over Instruction Tuning
J. Jang, S. Kim, S. Ye, D. Kim, L. Logeswaran, M. Lee, K. Lee and M. Seo. 2023 · 2023
Closest in time.
Esmfold hallucinates native-like protein sequences
J. R. Jeliazkov, D. del Alamo and J. D. Karpiak. 2023 · 2023
Closest in time.
Survey of Hallucination in Natural Language Generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang et al. 2023 · 2023
Closest in time.
Can large language models infer causation from correlation?
Z. Jin, J. Liu, Z. Lyu, S. Poff, M. Sachan, R. Mihalcea, M. Diab and B. Schölkopf. 2023 · 2023
Closest in time.
The MiniPile Challenge for Data-Efficient Language Models
J. Kaddour. 2023 · 2023
Closest in time.
No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models
J. Kaddour, O. Key, P. Nawrot, P. Minervini and M. J. Kusner. 2023 · 2023
Closest in time.
Tokenization issues (tweet)
A. Karpathy. 2023 · 2023
Closest in time.
Gpt-4 passes the bar exam
D. M. Katz, M. J. Bommarito, S. Gao and P. Arredondo. 2023 · 2023
Closest in time.
The impact of positional encoding on length generalization in transformers
A. Kazemnejad, I. Padhi, K. N. Ramamurthy, P. Das and S. Reddy. 2023 · 2023
Closest in time.
Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
O. Khattab, K. Santhanam, X. L. Li, D. Hall, P. Liang, C. Potts and M. Zaharia. 2023 · 2023
Closest in time.
Big little transformer decoder
S. Kim, K. Mangalam, J. Malik, M. W. Mahoney, A. Gholami and K. Keutzer. 2023 · 2023
Closest in time.
Chatgpt: Jack of all trades, master of none
J. Kocoń, I. Cichecki, O. Kaszyca, M. Kochanek, D. Szydło, J. Baran, J. Bielaniewicz, M. Gruza et al. 2023 · 2023
Closest in time.
Openassistant conversations–democratizing large language model alignment
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc et al. 2023 · 2023
Closest in time.
Pretraining language models with human preferences
T. Korbak, K. Shi, A. Chen, R. Bhalerao, C. L. Buckley, J. Phang, S. R. Bowman and E. Perez. 2023 · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
M. Kosinski. 2023 · 2023
Closest in time.
Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
K. Krishna, Y. Song, M. Karpinska, J. Wieting and M. Iyyer. 2023 · 2023
Closest in time.
vllm: Easy, fast, and cheap llm serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. Yu, J. Gonzalez, H. Zhang et al. 2023 · 2023
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality
E. Kıcıman, R. Ness, A. Sharma and C. Tan. 2023 · 2023
Closest in time.
Awesome-Prompt-Engineering
P. Lab. 2023 · 2023
Closest in time.
Passive learning of active causal strategies in agents and language models
A. K. Lampinen, S. C. Chan, I. Dasgupta, A. J. Nam and J. X. Wang. 2023 · 2023
Closest in time.
Do we still need clinical language models?
E. Lehman, E. Hernandez, D. Mahajan, J. Wulff, M. J. Smith, Z. Ziegler, D. Nadler, P. Szolovits et al. 2023 · 2023
Closest in time.
The diagnostic and triage accuracy of the gpt-3 artificial intelligence model
D. M. Levine, R. Tuwani, B. Kompa, A. Varma, S. G. Finlayson, A. Mehrotra and A. Beam. 2023 · 2023
Closest in time.
L. Lian, B. Li, A. Yala and T. Darrell. 2023 · 2023
Closest in time.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence and A. Zeng. 2023 · 2023
Closest in time.
Y.-T. Lin and Y.-N. Chen. 2023 · 2023
Closest in time.
ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing
R. Liu and N. B. Shah. 2023 · 2023
Closest in time.
S. Liu and Z. Wang. 2023 · 2023
Closest in time.
Analyzing Leakage of Personally Identifiable Information in Language Models
N. Lukas, A. Salem, R. Sim, S. Tople, L. Wutschitz and S. Zanella-Béguelin. 2023 · 2023
Closest in time.
Spawrious: A benchmark for fine control of spurious correlation biases
A. Lynch, G. J. Dovonon, J. Kaddour and R. Silva. 2023 · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri et al. 2023 · 2023
Closest in time.
Large language models generate functional protein sequences across diverse families
A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos Jr, C. Xiong et al. 2023 · 2023
Closest in time.
Training Models to Generate, Recognize, and Reframe Unhelpful Thoughts
M. Maddela, M. Ung, J. Xu, A. Madotto, H. Foran and Y.-L. Boureau. 2023 · 2023
Closest in time.
Memorization Capacity of Multi-Head Attention in Transformers
S. Mahdavi, R. Liao and C. Thrampoulidis. 2023 · 2023
Closest in time.
Fine-Tuning Language Models with Just Forward Passes
S. Malladi, T. Gao, E. Nichani, A. Damian, J. D. Lee, D. Chen and S. Arora. 2023 · 2023
Closest in time.
Large sequence models for software development activities
P. Maniatis and D. Tarlow. 2023 · 2023
Closest in time.
Inverse Scaling: When Bigger Isn’t Better
I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland et al. 2023 · 2023
Closest in time.
Mass-editing memory in a transformer
K. Meng, A. S. Sharma, A. J. Andonian, Y. Belinkov and D. Bau. 2023 · 2023
Closest in time.
Augmented language models: a survey
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick et al. 2023 · 2023
Closest in time.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer et al. 2023 · 2023
Closest in time.
DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning and C. Finn. 2023 · 2023
Closest in time.
Towards agile text classifiers for everyone
M. Mozes, J. Hoffmann, K. Tomanek, M. Kouate, N. Thain, A. Yuan, T. Bolukbasi and L. Dixon. 2023 · 2023
Closest in time.
Orca: Progressive learning from complex explanation traces of gpt-4
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi and A. Awadallah. 2023 · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Lieberum, J. Smith and J. Steinhardt. 2023 · 2023
Closest in time.
Transformers in healthcare: A survey
S. Nerella, S. Bandyopadhyay, J. Zhang, M. Contreras, S. Siegel, A. Bumin, B. Silva, J. Sena et al. 2023 · 2023
Closest in time.
Capabilities of gpt-4 on medical challenge problems
H. Nori, N. King, S. M. McKinney, D. Carignan and E. Horvitz. 2023 · 2023
Closest in time.
K. Nottingham, P. Ammanabrolu, A. Suhr, Y. Choi, H. Hajishirzi, S. Singh and R. Fox. 2023 · 2023
Closest in time.
Chatgpt goes to operating room: Evaluating gpt-4 performance and its potential in surgical education and training in the era of large language models
N. Oh, G.-S. Choi and W. Y. Lee. 2023 · 2023
Closest in time.
Chatgpt: Optimizing language models for dialogue
OpenAI. 2022 · 2023
Closest in time.
Faster causal attention over large sequences through sparse flash attention
M. Pagliardini, D. Paliotta, M. Jaggi and F. Fleuret. 2023 · 2023
Closest in time.
What in-context learning "learns" in-context: Disentangling task recognition and task learning
J. Pan, T. Gao, H. Chen and D. Chen. 2023 · 2023
Closest in time.
Art: Automatic multi-step reasoning and tool-use for large language models
B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer and M. T. Ribeiro. 2023 · 2023
Closest in time.
Bidirectional language models are also few-shot learners
A. Patel, B. Li, M. S. Rasooli, N. Constant, C. Raffel and C. Callison-Burch. 2023 · 2023
Closest in time.
Ai psychometrics: Using psychometric inventories to obtain psychological profiles of large language models
M. Pellert, C. M. Lechner, C. Wagner, B. Rammstedt and M. Strohmaier. 2023 · 2023
Closest in time.
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei et al. 2023 · 2023
Closest in time.
Language model tokenizers introduce unfairness between languages
A. Petrov, E. La Malfa, P. H. Torr and A. Bibi. 2023 · 2023
Closest in time.
Chatgpt, professor of law
T. Pettinato Oltz. 2023 · 2023
Closest in time.
An important next step on our ai journey
S. Pichai. 2023 · 2023
Closest in time.
Hyena Hierarchy: Towards Larger Convolutional Language Models
M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon et al. 2023 · 2023
Closest in time.
Measuring and Narrowing the Compositionality Gap in Language Models
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith and M. Lewis. 2023 · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning and C. Finn. 2023 · 2023
Closest in time.
ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
P. P. Ray. 2023 · 2023
Closest in time.
Pangu- Sigma: Towards trillion parameter language model with sparse heterogeneous computing
X. Ren, P. Zhou, X. Meng, X. Huang, Y. Wang, W. Wang, P. Li, X. Zhang et al. 2023 · 2023
Closest in time.
Language Modelling with Pixels
P. Rust, J. F. Lotz, E. Bugliarello, E. Salesky, M. de Lhoneux and D. Elliott. 2023 · 2023
Closest in time.
Can AI-Generated Text be Reliably Detected?
V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang and S. Feizi. 2023 · 2023
Closest in time.
Personality traits in large language models
M. Safdari, G. Serapio-García, C. Crepy, S. Fitz, P. Romero, L. Sun, M. Abdulhai, A. Faust et al. 2023 · 2023
Closest in time.
In-context impersonation reveals large language models’ strengths and biases
L. Salewski, S. Alaniz, I. Rio-Torto, E. Schulz and Z. Akata. 2023 · 2023
Closest in time.
Stay on topic with Classifier-Free Guidance
G. Sanchez, H. Fan, A. Spangher, E. Levi, P. S. Ammanamanchi and S. Biderman. 2023 · 2023
Closest in time.
Understanding the effectiveness of early weight averaging for training large language models
S. Sanyal, J. Kaddour, A. Kumar and S. Sanghavi. 2023 · 2023
Closest in time.
Explaining legal concepts with augmented large language models (gpt-4)
J. Savelka, K. D. Ashley, M. A. Gray, H. Westermann and H. Xu. 2023 · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda and S. Koyejo. 2023 · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda and T. Scialom. 2023 · 2023
Closest in time.
High-throughput generative inference of large language models with a single gpu
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré et al. 2023 · 2023
Closest in time.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal et al. 2023 · 2023
Closest in time.
Exploring the robustness of large language models for solving programming problems
A. Shirafuji, Y. Watanobe, T. Ito, M. Morishita, Y. Nakamura, Y. Oda and J. Suzuki. 2023 · 2023
Closest in time.
The curse of recursion: Training on generated data makes models forget
I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot and R. Anderson. 2023 · 2023
Closest in time.
S. Sia and K. Duh. 2023 · 2023
Closest in time.
Towards expert-level medical question answering with large language models
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, L. Hou, K. Clark, S. Pfohl et al. 2023 · 2023
Closest in time.
Emergent deception and emergent optimization
J. Steinhardt. 2023 · 2023
Closest in time.
A simple and effective pruning approach for large language models
M. Sun, Z. Liu, A. Bair and J. Z. Kolter. 2023 · 2023
Closest in time.
A short survey of viewing large language models in legal aspect
Z. Sun. 2023 · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
D. Surís, S. Menon and C. Vondrick. 2023 · 2023
Closest in time.
Piling on to the pile-on (sorry - it’s always easy to criticize), here’s a rant about benchmarks for LLMs that are used to back claims of "stronger" or "better" models. Let’s start with a tour through GPT-3’s Appendix G… 1/8
Susan Zhang [@suchenzang]. 2023 · 2023
Closest in time.
Schema-learning and rebinding as mechanisms of in-context learning and emergence
S. Swaminathan, A. Dedieu, R. V. Raju, M. Shanahan, M. Lazaro-Gredilla and D. George. 2023 · 2023
Closest in time.
Evaluating large language models on medical evidence summarization
L. Tang, Z. Sun, B. Idnay, J. G. Nestor, A. Soroush, P. A. Elias, Z. Xu, Y. Ding et al. 2023b · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang and T. B. Hashimoto. 2023 · 2023
Closest in time.
Better language models of code through self-improvement
H. Q. To, N. D. Bui, J. Guo and T. N. Nguyen. 2023 · 2023
Closest in time.
AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks
A. Tornede, D. Deng, T. Eimer, J. Giovanelli, A. Mohan, T. Ruhkopf, S. Segel, D. Theodorakopoulos et al. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal et al. 2023 · 2023
Closest in time.
Survey of protein sequence embedding models
C. Tran, S. Khadkikar and A. Porollo. 2023 · 2023
Closest in time.
Holistic evaluation of langauge models results page
S. University. 2023 · 2023
Closest in time.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
K. Valmeekam, A. Olmo, S. Sreedharan and S. Kambhampati. 2023 · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
S. Vemprala, R. Bonatti, A. Bucker and A. Kapoor. 2023 · 2023
Closest in time.
Pubmed gpt: A domain- specific large language model for biomedical text
A. Venigalla, J. Frankle and M. Carbin. 2022 · 2023
Closest in time.
Fairpy: A toolkit for evaluation of social biases and their mitigation in large language models
H. Viswanath and T. Zhang. 2023 · 2023
Closest in time.
Go smol or go home
H. d. Vries. 2023 · 2023
Closest in time.
Jailbroken: How Does LLM Safety Training Fail?
A. Wei, N. Haghtalab and J. Steinhardt. 2023 · 2023
Closest in time.
Causal parrots: Large language models may talk causality but are not causal
M. Willig, M. ZEČEVIĆ, D. S. Dhami and K. Kersting. 2023 · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Y. Wolf, N. Wies, Y. Levine and A. Shashua. 2023 · 2023
Closest in time.
M. Wornow, Y. Xu, R. Thapa, B. Patel, E. Steinberg, S. Fleming, M. A. Pfeffer, J. Fries et al. 2023 · 2023
Closest in time.
When geometric deep learning meets pretrained protein language models
F. Wu, D. Radev and J. Xu. 2023a · 2023
Closest in time.
L. Yan, L. Sha, L. Zhao, Y. Li, R. Martinez-Maldonado, G. Chen, X. Li, Y. Jin et al. 2023 · 2023
Closest in time.
Statler: State-maintaining language models for embodied reasoning
T. Yoneda, J. Fang, P. Li, H. Zhang, T. Jiang, S. Lin, B. Picker, D. Yunis et al. 2023 · 2023
Closest in time.
Robust Natural Language Watermarking through Invariant Features
K. Yoo, W. Ahn, J. Jang and N. Kwak. 2023 · 2023
Closest in time.
Megabyte: Predicting million-byte sequences with multiscale transformers
L. Yu, D. Simig, C. Flaherty, A. Aghajanyan, L. Zettlemoyer and M. Lewis. 2023 · 2023
Closest in time.
Chatdoctor: A medical chat model fine-tuned on llama model using medical domain knowledge
L. Yunxiang, L. Zihan, Z. Kai, D. Ruilong and Z. You. 2023 · 2023
Closest in time.
[…] that’s an unhelpful order of magnitude difference in how large of a model you should be training in order to be considered “compute optimal”
S. Zhang. 2023 · 2023
Closest in time.
Secrets of RLHF in Large Language Models Part I: PPO
R. Zheng, S. Dou, S. Gao, W. Shen, B. Wang, Y. Liu, S. Jin, Q. Liu et al. 2023 · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen et al. 2023 · 2023
Closest in time.
A survey on efficient training of transformers
B. Zhuang, J. Liu, Z. Pan, H. He, Y. Weng and C. Shen. 2023 · 2023
Closest in time.