Fetching the paper…
Reading the bibliography…
The rapid advancement of artificial intelligence, particularly with the development of Large Language Models (LLMs) built on the transformer architecture, has redefined the capabilities of natural language processing.
“Multi-task deep neural networks for natural language understanding”
Xiaodong Liu, Pengcheng He, Weizhu Chen and Jianfeng Gao · 1901
Earlier work this paper cites.
“Parameter-Efficient Transfer Learning for NLP”, 2019
Neil Houlsby et al · 1902
Earlier work this paper cites.
“Single-cell trajectories reconstruction, exploration and mapping of omics data with STREAM”
H. Chen · 1903
Earlier work this paper cites.
“Generating Long Sequences with Sparse Transformers”
Rewon Child, Scott Gray, Alec Radford and Ilya Sutskever · 1904
Earlier work this paper cites.
“Did the model understand the question?”
Pramod Mudrakarta, Ankur Taly, Mukund Sundararajan and Kedar Dhamdhere · 1906
Earlier work this paper cites.
“Fast Transformer Decoding: One Write-Head is All You Need”
Noam Shazeer · 1911
Earlier work this paper cites.
“WWW’18 Open Challenge: Financial Opinion Mining and Question Answering”
Macedo Maia, Siegfried Handschuh and Andr“’e Freitas · 1942
Earlier work this paper cites.
“More is Different: Broken Symmetry and the Nature of the Hierarchical Structure of Science”
Philip. Anderson · 1972
Earlier work this paper cites.
“Stock movement prediction from tweets and historical prices”
Yumo Xu and Shay Cohen · 1979
Earlier work this paper cites.
“Machine Learning”
Tom. Mitchell · 1997
Earlier work this paper cites.
“Statistical Learning Theory”
Vladimir Vapnik · 1998
Earlier work this paper cites.
“Transductive inference for text classification using support vector machines”
Thorsten Joachims · 1999
Earlier work this paper cites.
“Conditional random fields: Probabilistic models for segmenting and labeling sequence data”
John. Lafferty, Andrew McCallum and Fernando C.. Pereira · 2001
Earlier work this paper cites.
“Towards a Human-like Open-Domain Chatbot”, 2020
Daniel Adiwardana et al · 2001
Earlier work this paper cites.
“A Neural Probabilistic Language Model”
Yoshua Bengio, Rejean Ducharme, Pascal Vincent and Christian Janvin · 2003
Earlier work this paper cites.
“VAL: Automatic plan validation, continuous effects and mixed initiative planning using PDDL”
R. Howey, D. Long and M. Fox · 2004
Earlier work this paper cites.
“Conceptnet–a practical commonsense reasoning tool-kit”
Hugo Liu and Push Singh · 2004
Earlier work this paper cites.
“Learning with unlabeled data and its application to image retrieval”
Dengyong Zhou et al · 2004
Earlier work this paper cites.
“Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks”, 2020
Suchin Gururangan et al · 2004
Earlier work this paper cites.
“Semi-supervised Learning Literature Survey”
Xiaojin Zhu · 2005
Earlier work this paper cites.
“Language Models Are Few-Shot Learners”, 2020
Tom. Brown et al · 2005
Earlier work this paper cites.
“Manifold regularization: A geometric framework for learning from labeled and unlabeled examples”
Mikhail Belkin, Partha Niyogi and Vikas Sindhwani · 2006
Earlier work this paper cites.
“One-Shot Learning of Object Categories”
Li Fei-Fei, Rob Fergus and Pietro Perona · 2006
Earlier work this paper cites.
“A tutorial on planning graph based reachability heuristics”
D. Bryce and S. Kambhampati · 2007
Earlier work this paper cites.
“Semi-supervised Learning”
Olivier Chapelle, Bernhard Scholkopf and Alexander Zien · 2009
Earlier work this paper cites.
“The Path to Personalized Medicine”
M.. Hamburg and F.. Collins · 2010
Earlier work this paper cites.
“Rectified Linear Units Improve Restricted Boltzmann Machines”
Vinod Nair and Geoffrey. Hinton · 2010
Earlier work this paper cites.
“Deep Sparse Rectifier Neural Networks”
Xavier Glorot, Antoine Bordes and Yoshua Bengio · 2011
Earlier work this paper cites.
“Thinking, Fast and Slow”
Daniel Kahneman · 2011
Earlier work this paper cites.
“Long Range Arena: A Benchmark for Efficient Transformers”
Yi Tay et al · 2011
Earlier work this paper cites.
“Pseudo-Label: The Simple and Efficient Semi-supervised Learning Method for Deep Neural Networks”
Dong-Hyun Lee · 2013
Earlier work this paper cites.
“Rectifier nonlinearities improve neural network acoustic models”
Andrew. Maas, Awni. Hannun and Andrew. Ng · 2013
Earlier work this paper cites.
“Distributed Representations of Words and Phrases and Their Compositionality”
Tomas Mikolov et al · 2013
Earlier work this paper cites.
“Efficient Estimation of Word Representations in Vector Space”
Tomas Mikolov, Kai Chen, Greg Corrado and Jeffrey Dean · 2013
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate”
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2014
Earlier work this paper cites.
“Good debt or bad debt: Detecting semantic orientations in economic texts”
Pekka Malo et al · 2014
Earlier work this paper cites.
“Domain adaption of named entity recognition to support credit risk assessment”
Julio Cesar Alvarado, Karin Verspoor and Timothy Baldwin · 2015
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift”
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
“Aligning books and movies: Towards story-like visual explanations by watching movies and reading books”
Yukun Zhu et al · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey. Hinton · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Gaussian Error Linear Units (GELUs)”
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
“Gaussian Error Linear Units (GELUs)”
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
“Convolutional Neural Networks Using Logarithmic Data Representation”
Daisuke Miyashita, Edward. Lee and Boris Murmann · 2016
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Regularization with stochastic transformations and perturbations for deep semi-supervised learning”
Mehdi Sajjadi, Mehran Javanmardi and Tolga Tasdizen · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units”
Rico Sennrich, Barry Haddow and Alexandra Birch · 2016
Earlier work this paper cites.
“Google’s Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation”
Yonghui Wu et al · 2016
Earlier work this paper cites.
“Massive exploration of neural machine translation architectures”
Denny Britz, Anna Goldie, Minh-Thang Luong and Quoc. Le · 2017
Earlier work this paper cites.
“Deep reinforcement learning from human preferences”
Paul. Christiano et al · 2017
Earlier work this paper cites.
“Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations”
Itay Hubara et al · 2017
Earlier work this paper cites.
“Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference”, 2017
Benoit Jacob et al · 2017
Earlier work this paper cites.
“Searching for activation functions”
Prajit Ramachandran, Barret Zoph and Quoc. Le · 2017
Earlier work this paper cites.
“Transfer Learning for Sequence Tagging with Hierarchical Recurrent Networks”, 2017
Zhilin Yang, Ruslan Salakhutdinov and William. Cohen · 2017
Earlier work this paper cites.
“Big Data and Machine Learning in Health Care”
Andrew. Beam and Isaac. Kohane · 2018
Earlier work this paper cites.
“Deep learning and algorithmic trading”
Hans Buehler, Lukas Gonon, Josef Teichmann and Ben Wood · 2018
Earlier work this paper cites.
“Universal Language Model Fine-tuning for Text Classification”, 2018
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing”
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
“Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks”
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
“Deep Contextualized Word Representations”, arXiv preprint, 2018
Matthew. Peters et al · 2018
Earlier work this paper cites.
“Improving Language Understanding by Generative Pre-training” Available online, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“Self-attention with relative position representations”
Peter Shaw, Jakob Uszkoreit and Ashish Vaswani · 2018
Earlier work this paper cites.
“Deep EHR: A survey of recent advances in deep learning techniques for electronic health record (EHR) analysis”
Benjamin Shickel · 2018
Earlier work this paper cites.
“What makes reading comprehension questions easier?”
Saku Sugawara, Kentaro Inui, Satoshi Sekine and Akiko Aizawa · 2018
Earlier work this paper cites.
“A Simple Method for Commonsense Reasoning”
Trieu. Trinh and Quoc. Le · 2018
Earlier work this paper cites.
“GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”
Alex Wang et al · 2018
Earlier work this paper cites.
“Attention? Attention!”
Lilian Weng · 2018
Earlier work this paper cites.
“Hybrid deep sequential modeling for social text-driven stock prediction”
Huizhe Wu, Wei Zhang, Weiwei Shen and Jun Wang · 2018
Earlier work this paper cites.
“HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering”
Zhilin Yang et al · 2018
Earlier work this paper cites.
“Publicly available clinical BERT embeddings”
Emily Alsentzer · 2019
Earlier work this paper cites.
“Adaptive Input Representations for Neural Language Modeling” OpenReview.net
Alexei Baevski and Michael Auli · 2019
Earlier work this paper cites.
“Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Unified Language Model Pre-training for Natural Language Understanding and Generation”
Li Dong et al · 2019
Earlier work this paper cites.
“OpenWebText Corpus”, http://Skylion007.github.io/OpenWebTextCorpus , 2019
Aaron Gokaslan, Ellie Pavlick and Stefanie Tellex · 2019
Earlier work this paper cites.
“PubMedQA: A Dataset for Biomedical Research Question Answering”
Qingyu Jin et al · 2019
Earlier work this paper cites.
“The Personalization of Conversational Agents in Health Care: Systematic Review”
A. Kocaballi et al · 2019
Earlier work this paper cites.
“Natural Questions: A Benchmark for Question Answering Research”
Tom Kwiatkowski et al · 2019
Earlier work this paper cites.
“RoBERTa: A Robustly Optimized BERT Pretraining Approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks”, 2019
Jason Phang, Thibault F“’evry and Samuel. Bowman · 2019
Earlier work this paper cites.
“Language Models Are Unsupervised Multitask Learners”, OpenAI Blog, 2019
Alec Radford et al · 2019
Earlier work this paper cites.
“Transfer Learning in Natural Language Processing”
Sebastian Ruder, Matthew. Peters, Swabha Swayamdipta and Thomas Wolf · 2019
Earlier work this paper cites.
“DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”
Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf · 2019
Earlier work this paper cites.
“Improving fraud detection in financial services through deep learning”
Timothy Smith and Manish Kumar · 2019
Earlier work this paper cites.
“Energy and Policy Considerations for Deep Learning in NLP”
Emma Strubell, Ananya Ganesh and Andrew McCallum · 2019
Earlier work this paper cites.
“CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge”, 2019
Alon Talmor, Jonathan Herzig, Nicholas Lourie and Jonathan Berant · 2019
Earlier work this paper cites.
“Defending Against Neural Fake News” NeurIPS 2019, December 8-14
Rowan Zellers et al · 2019
Earlier work this paper cites.
“Root Mean Square Layer Normalization”
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
“Deep learning enables rapid identification of potent DDR1 kinase inhibitors”
Alex Zhavoronkov · 2019
Earlier work this paper cites.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“The Pushshift Reddit Dataset” ICWSM 2020, Held Virtually
Jason Baumgartner et al · 2020
Earlier work this paper cites.
“Scaling Laws for Autoregressive Generative Modeling”
Tom Henighan et al · 2020
Earlier work this paper cites.
“The Curious Case of Neural Text Degeneration” OpenReview.net
Ari Holtzman et al · 2020
Earlier work this paper cites.
“Ethical considerations for AI in finance”
Michael Jones, Ravinder Barn, Mohammed Karim and Jason R Nurse · 2020
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”
Jared Kaplan et al · 2020
Earlier work this paper cites.
“BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”
Mike Lewis et al · 2020
Earlier work this paper cites.
“MAEC: A Multimodal Aligned Earnings Conference Call Dataset for Financial Risk Prediction”
Jiazheng Li, Linyi Yang, Barry Smyth and Ruihai Dong · 2020
Earlier work this paper cites.
“Natural language processing in risk management and compliance”
Jin Li, Scott Spangler and Yue Yu · 2020
Earlier work this paper cites.
“Understanding the difficulty of training transformers”
Lizi Liu et al · 2020
Earlier work this paper cites.
“The financial document causality detection shared task (fincausal 2020)”
Dominique Mariko, Hanna Akl and Estelle Labidurie · 2020
Earlier work this paper cites.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”
Colin Raffel et al · 2020
Earlier work this paper cites.
“GLU Variants Improve Transformer”
Noam Shazeer · 2020
Earlier work this paper cites.
“On Layer Normalization in the Transformer Architecture”
Ruibo Xiong et al · 2020
Earlier work this paper cites.
“Big Bird: Transformers for Longer Sequences”
Manzil Zaheer et al · 2020
Earlier work this paper cites.
“Muppet: Massive multi-task representations with pre-finetuning”
Anna Aghajanyan et al · 2021
Earlier work this paper cites.
“A General Language Assistant as a Laboratory for Alignment”
Amanda Askell et al · 2021
Earlier work this paper cites.
“Program synthesis with large language models”
James Austin et al · 2021
Earlier work this paper cites.
“On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”
Emily Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
“Extracting training data from large language models”
Nicholas Carlini et al · 2021
Earlier work this paper cites.
“Evaluating Large Language Models Trained on Code”, arXiv preprint arXiv:2107.03374, 2021
Mark Chen et al · 2021
Earlier work this paper cites.
“CogView: Mastering Text-to-Image Generation via Transformers”
Ming Ding et al · 2021
Earlier work this paper cites.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2021
Earlier work this paper cites.
“The Pile: An 800GB Dataset of Diverse Text for Language Modeling”
Leo Gao et al · 2021
Earlier work this paper cites.
“Multilingual and cross-lingual intent detection from spoken data”
Daniela Gerz et al · 2021
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”, 2021
Edward. Hu et al · 2021
Earlier work this paper cites.
“Alignment of language agents”
Z. Kenton et al · 2021
Earlier work this paper cites.
Michael. Krell, Mario Kosec, Santiago. Perez and Andrew Fitzgibbon · 2021
Earlier work this paper cites.
“Why machine reading comprehension models learn shortcuts?”
Yuxuan Lai et al · 2021
Earlier work this paper cites.
“The Power of Scale for Parameter-Efficient Prompt Tuning”, 2021
Brian Lester, Rami Al-Rfou and Noah Constant · 2021
Earlier work this paper cites.
“Prefix-Tuning: Optimizing Continuous Prompts for Generation”, 2021
Xiang Li and Percy Liang · 2021
Earlier work this paper cites.
“A survey on deep learning in medical image analysis”
Zhi Li, Qiang Zhang and Qi Dou · 2021
Earlier work this paper cites.
“Jurassic-1: Technical details and evaluation”
Or Lieber, Or Sharir, Barak Lenz and Yoav Shoham · 2021
Earlier work this paper cites.
Pengfei Liu et al · 2021
Earlier work this paper cites.
“A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers”, 2021
Shen-Yun Miao, Chao-Chun Liang and Keh-Yih Su · 2021
Earlier work this paper cites.
“WebGPT: Browser-assisted Question-Answering with Human Feedback”
R. Nakano et al · 2021
Earlier work this paper cites.
“Do Transformer Modifications Transfer Across Implementations and Applications?”
Sharan Narang et al · 2021
Earlier work this paper cites.
“GPT3-toPlan: Extracting Plans from Text using GPT-3”
A. Olmo, S. Sreedharan and S. Kambhampati · 2021
Earlier work this paper cites.
“Enhancing customer service through AI-driven virtual assistants in the banking sector”
Arpan Pal, Aniruddha Kundu and Rajdeep Chakraborty · 2021
Earlier work this paper cites.
Baolin Peng, Xiang Li and Percy Liang · 2021
Earlier work this paper cites.
“Learning how to ask: Querying LMs with mixtures of soft prompts”
Guanghui Qin and Jason Eisner · 2021
Earlier work this paper cites.
“Learning Transferable Visual Models From Natural Language Supervision”, 2021
Alec Radford et al · 2021
Earlier work this paper cites.
“Scaling language models: Methods, analysis & insights from training Gopher”
Jack. Rae et al · 2021
Cited alongside, same era.
“Zero-Shot Text-to-Image Generation”, 2021
Aditya Ramesh et al · 2021
Cited alongside, same era.
“Impact of news on the commodity market: Dataset and results”
Ankur Sinha and Tanmay Khandait · 2021
Cited alongside, same era.
“RoFormer: Enhanced Transformer with Rotary Position Embedding”
Jianlin Su et al · 2021
Cited alongside, same era.
“Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models”
Alex Tamkin, Singh Trisha, Davide Giovanardi and Noah Goodman · 2021
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%*ChatGPT Quality” [Online]. Available: https://vicuna.lmsys.org , 2023
W.-L. Chiang et al · 2023
Later among the works it cites.
“ChatGPT goes to law school” Accessed: 2024-02-14, Available at SSRN, 2023
Jinho. Choi, Kristin. Hickman, Andrew Monahan and Daniel Schwarcz · 2023
Later among the works it cites.
“A Survey on In-context Learning”, 2023
Qingxiu Dong et al · 2023
Later among the works it cites.
“Faith and fate: Limits of transformers on compositionality”
Nouha Dziri et al · 2023
Later among the works it cites.
“A Closer Look at Large Language Models: Emergent Abilities” Accessed: 2023-07-14, https://www.notion.so/yaofu/A-Closer-Look-at-Large-Language-Models-Emergent-Abilities-493876b55df5479d80686f68a1abd72f , 2023
Yao Fu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Frozen in Time: Temporal Contextualization for In-Context Learning”
Katerina Tsimpoukelli, Giorgos Karamanolakis, Athanasios Katsamanis and Petros Maragos · 2021
Cited alongside, same era.
“Milvus: A Purpose-Built Vector Data Management System”
J. Wang et al · 2021
Cited alongside, same era.
Weihua Zeng et al · 2021
Cited alongside, same era.
“Medical image analysis with artificial intelligence”
Jun Zhang · 2021
Cited alongside, same era.
“Calibrate Before Use: Improving Few-shot Performance of Language Models”
Zihao Zhao et al · 2021
Cited alongside, same era.
“Global Table Extractor (GTE): A Framework for Joint Table Identification and Cell Structure Recognition Using Visual Context”
Xinyi Zheng et al · 2021
Cited alongside, same era.
“Trade the event: Corporate events detection for news-based event-driven trading”
Zhihan Zhou, Liqian Ma and Han Liu · 2021
Cited alongside, same era.
“What Can Transformers Learn In-Context? A Case Study of Simple Function Classes”, 2023
Shivam Garg, Dimitris Tsipras, Percy Liang and Gregory Valiant · 2023
Later among the works it cites.
“Large language models are not abstract reasoners”
Guillaume Gendron, Qiaozi Bao, Michael Witbrock and Gill Dobbie · 2023
Later among the works it cites.
“Pre-training to Learn in Context” arXiv preprint arXiv:2305.09137, 2023
Yuxian Gu, Li Dong, Furu Wei and Minlie Huang · 2023
Later among the works it cites.
“Leveraging pre-trained large language models to construct and utilize world models for model-based task planning”
Lin Guan, Karthik Valmeekam, Shashank Sreedharan and Subbarao Kambhampati · 2023
Later among the works it cites.
“How close is ChatGPT to human experts? Comparison corpus, evaluation, and detection”
Binbin Guo et al · 2023
Later among the works it cites.
“A Theory of Emergent In-Context Learning as Implicit Structure Induction”, 2023
Michael Hahn and Navin Goyal · 2023
Later among the works it cites.
“Using ChatGPT to Conduct a Literature Review”
Micha Haman and Marcin Skolnik · 2023
Later among the works it cites.
“Reasoning with language model is planning with world model”
S. Hao et al · 2023
Later among the works it cites.
“ChatGPT as Your Personal Data Scientist”
Md Hassan, Richard. Knipper and Shakked K.. Santu · 2023
Later among the works it cites.
“LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models”, 2023
Zhiqiang Hu et al · 2023
Later among the works it cites.
Shaohan Huang et al · 2023
Later among the works it cites.
“Artificial Hallucinations in ChatGPT: Implications in Scientific Writing” Available on PubMed
S… Hussam · 2023
Later among the works it cites.
“Advanced RAG Techniques: An Illustrated Overview” Accessed: 2024-12-24
I. Ilin · 2023
Later among the works it cites.
Yuxuan Ji et al · 2023
Later among the works it cites.
“MultiFin: A Dataset for Multilingual Financial NLP”
Rasmus Jrgensen et al · 2023
Later among the works it cites.
“On the role of large language models in planning”
Subbarao Kambhampati, Karthik Valmeekam, Marcos Marquez and Lin Guan · 2023
Later among the works it cites.
G. Kim et al · 2023
Later among the works it cites.
“Large Language Models are Zero-Shot Reasoners”, 2023
Takeshi Kojima et al · 2023
Later among the works it cites.
“OpenAssistant Conversations–Democratizing Large Language Model Alignment”
Andreas Kopf et al · 2023
Later among the works it cites.
“Theory of Mind May Have Spontaneously Emerged in Large Language Models”
M. Kosinski · 2023
Later among the works it cites.
“Understanding In-Context Learning”, 2023
Stanford Lab · 2023
Later among the works it cites.
“StockEmotions: Discover Investor Emotions for Financial Sentiment Analysis and Multivariate Time Series”
Jean Lee, Hoyoul Youn, Josiah Poon and Soyeon Han · 2023
Later among the works it cites.
“Contextual Prompting for In-Context Learning”
Mukai Li et al · 2023
Later among the works it cites.
X. Li et al · 2023
Later among the works it cites.
Xianzhi Li et al · 2023
Later among the works it cites.
“Finding Supporting Examples for In-Context Learning” arXiv preprint arXiv:2302.13539, 2023
Xiaonan Li and Xipeng Qiu · 2023
Later among the works it cites.
“MoT: Memory-of-Thought Enables ChatGPT to Self-Improve”, 2023
Xiaonan Li and Xipeng Qiu · 2023
Later among the works it cites.
“Making Large Language Models Better Reasoners with Step-Aware Verifier”, 2023
Yifei Li et al · 2023
Later among the works it cites.
“Transformers as Algorithms: Generalization and Stability in In-context Learning”, 2023
Yingcong Li, M. Ildiz, Dimitris Papailiopoulos and Samet Oymak · 2023
Later among the works it cites.
“LLM+P: Empowering Large Language Models with Optimal Planning Proficiency”, 2023
Bo Liu et al · 2023
Later among the works it cites.
“Reviewergpt? An Exploratory Study on Using Large Language Models for Paper Reviewing”
R. Liu and N.. Shah · 2023
Later among the works it cites.
Scott Longpre et al · 2023
Later among the works it cites.
“The FLAN Collection: Designing Data and Methods for Effective Instruction Tuning”
Shayne Longpre et al · 2023
Later among the works it cites.
“Multimodal procedural planning via dual text-image prompting”
Y. Lu et al · 2023
Later among the works it cites.
“Faithful chain-of-thought reasoning”
Q. Lyu et al · 2023
Later among the works it cites.
“Query rewriting for retrieval-augmented large language models”
X. Ma et al · 2023
Later among the works it cites.
“Query Rewriting for Retrieval-Augmented Large Language Models”, 2023
Xinbei Ma et al · 2023
Later among the works it cites.
K. Malinka et al · 2023
Later among the works it cites.
“Do You Know English Grammar Better Than ChatGPT?”, 2023
Lev Maximov · 2023
Later among the works it cites.
R.. McCoy et al · 2023
Later among the works it cites.
“Transformers learn in-context by gradient descent”, 2023
Johannes von Oswald et al · 2023
Later among the works it cites.
“What In-context Learning “Learns”
J. Pan, T. Gao, H. Chen and D. Chen · 2023
Later among the works it cites.
“Generative Agents: Interactive Simulacra of Human Behavior”, 2023
Joon Park et al · 2023
Later among the works it cites.
“Can ChatGPT be used to generate scientific hypotheses?”, 2023
Yang Park et al · 2023
Later among the works it cites.
Gerardo Penedo et al · 2023
Later among the works it cites.
“RWKV: Reinventing RNNs for the Transformer Era”
Bin Peng et al · 2023
Later among the works it cites.
Hugging Face
“Perplexity - Transformers” Accessed: 2024-04-06, 2023 · 2023
Later among the works it cites.
“Hyena hierarchy: Towards larger convolutional language models”
Michael Poli et al · 2023
Later among the works it cites.
“GPT-4: A Large-Scale Generative Pre-trained Transformer”
Alec Radford et al · 2023
Later among the works it cites.
“Understanding Encoder and Decoder”, 2023
Sebastian Raschka · 2023
Later among the works it cites.
“Are Emergent Abilities of Large Language Models a Mirage?”, 2023
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo · 2023
Later among the works it cites.
“Toolformer: Language models can teach themselves to use tools”
Timo Schick et al · 2023
Later among the works it cites.
“Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy”
Z. Shao et al · 2023
Later among the works it cites.
“Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface”
Y. Shen et al · 2023
Later among the works it cites.
“Reflexion: Language agents with verbal reinforcement learning”, 2023
N. Shinn et al · 2023
Later among the works it cites.
“ChatGPT: A Study on Its Utility for Ubiquitous Software Engineering Tasks”
G. Sridhara, R.. G. and S. Mazumdar · 2023
Later among the works it cites.
“Adaplanner: Adaptive planning from feedback with language models”
H. Sun et al · 2023
Later among the works it cites.
“Automatic Code Summarization via ChatGPT: How Far Are We?”
W. Sun et al · 2023
Later among the works it cites.
“Retentive Network: A Successor to Transformer for Large Language Models”
Yi Sun et al · 2023
Later among the works it cites.
“Stanford ALPACA: An Instruction-Following LLaMA Model”, https://github.com/tatsu-lab/stanford-alpaca , 2023
Rohan Taori et al · 2023
Later among the works it cites.
“UL2: Unifying Language Learning Paradigms”, 2023
Yi Tay et al · 2023
Later among the works it cites.
“LLaMA 2: Open Foundation and Fine-Tuned Chat Models”
Hugo Touvron et al · 2023
Later among the works it cites.
“LLaMA: Open and Efficient Foundation Language Models”, 2023
Hugo Touvron et al · 2023
Later among the works it cites.
“Large language models fail on trivial alterations to theory-of-mind tasks”
Tomer Ullman · 2023
Later among the works it cites.
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev and Ali Ghodsi · 2023
Later among the works it cites.
“On the planning abilities of large language models: A critical investigation”
Karthik Valmeekam, Marcos Marquez, Shashank Sreedharan and Subbarao Kambhampati · 2023
Later among the works it cites.
“Attention Is All You Need” v7, 2023
Ashish Vaswani et al · 2023
Later among the works it cites.
URL: https://vllm.ai/
“vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention”, Available online, 2023 · 2023
Later among the works it cites.
“Efficient prompting via dynamic in-context learning”
Chunshu Wang, Yuchen Jiang, Ryan Cotterell and Mrinmaya Sachan · 2023
Later among the works it cites.
“Voyager: An Open-Ended Embodied Agent with Large Language Models”, 2023
Guanzhi Wang et al · 2023
Later among the works it cites.
L. Wang et al · 2023
Later among the works it cites.
“FinGPT: Instruction Tuning Benchmark for Open-Source Large Language Models in Financial Datasets”
Neng Wang, Hongyang Yang and Christina Wang · 2023
Later among the works it cites.
“Images Speak in Images: A Generalist Painter for In-Context Visual Learning”
Xinlong Wang et al · 2023
Later among the works it cites.
“SegGPT: Segmenting Everything in Context”
Xinlong Wang et al · 2023
Later among the works it cites.
Xinyi Wang, Wanrong Zhu and William Wang · 2023
Later among the works it cites.
“In-Context Learning Unlocked for Diffusion Models” arXiv preprint arXiv:2305.01115, 2023
Zhendong Wang et al · 2023
Later among the works it cites.
Zihao Wang et al · 2023
Later among the works it cites.
“Larger language models do in-context learning differently”, 2023
Jerry Wei et al · 2023
Later among the works it cites.
“Symbol tuning improves in-context learning in language models”, 2023
Jerry Wei et al · 2023
Later among the works it cites.
“The Learnability of In-context Learning”
N. Wies, Y. Levine and A. Shashua · 2023
Later among the works it cites.
“Bayesian Inference”, 2023
Wikipedia · 2023
Later among the works it cites.
“BLOOM: A 176B-Parameter Open-Access Multilingual Language Model”, 2023
BigScience Workshop · 2023
Later among the works it cites.
“BloombergGPT: A Large Language Model for Finance”, 2023
Shijie Wu et al · 2023
Later among the works it cites.
“OpenICL: An Open-Source Framework for In-context Learning”, 2023
Zhenyu Wu et al · 2023
Later among the works it cites.
“Conversational Automated Program Repair”
C.. Xia and L. Zhang · 2023
Later among the works it cites.
“Pixiu: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance”
Q. Xie et al · 2023
Later among the works it cites.
“kNN Prompting: Learning Beyond the Context with Nearest Neighbor Inference” 2023a
Benfeng Xu et al · 2023
Later among the works it cites.
“WizardLM: Empowering Large Language Models to Follow Complex Instructions”, 2023
Can Xu et al · 2023
Later among the works it cites.
“Small Models are Valuable Plug-ins for Large Language Models” arXiv preprint arXiv:2305.08848, 2023
Canwen Xu et al · 2023
Later among the works it cites.
“Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data”
Chen Xu, Dong Guo, Nan Duan and Julian McAuley · 2023
Later among the works it cites.
“InvestLM: A Large Language Model for Investment Using Financial Domain Instruction Tuning”
Yi Yang, Yixuan Tang and Kar Tam · 2023
Later among the works it cites.
“Tree of thoughts: Deliberate problem solving with large language models”
S. Yao et al · 2023
Later among the works it cites.
“A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models”, 2023
Junjie Ye et al · 2023
Later among the works it cites.
“Making retrieval-augmented language models robust to irrelevant context”
O. Yoran, T. Wolfson, O. Ram and J. Berant · 2023
Later among the works it cites.
“One small step for generative AI, one giant leap for AGI: A complete survey on ChatGPT in AIGC era”
C. Zhang et al · 2023
Later among the works it cites.
“AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning”, 2023
Qingru Zhang et al · 2023
Later among the works it cites.
“A Survey of Large Language Models”
Wayne Zhao et al · 2023
Later among the works it cites.
“Take a step back: Evoking reasoning via abstraction in large language models”
H.. Zheng et al · 2023
Later among the works it cites.
“CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Evaluations on HumanEval-X”
Qinkai Zheng et al · 2023
Later among the works it cites.
“MemoryBank: Enhancing Large Language Models with Long-Term Memory”, 2023
Wanjun Zhong et al · 2023
Later among the works it cites.
“Large language models are human-level prompt engineers”
Y. Zhou et al · 2023
Later among the works it cites.
Aaron et al · 2024
Later among the works it cites.
“GPT-4 Technical Report”, 2024
Josh et al · 2024
Later among the works it cites.
“Claude 3 Model Card” Accessed: 2024-12-24
Anthropic · 2024
Later among the works it cites.
“Gemma: Google introduces new state-of-the-art open models”, Google AI Blog, 2024
Jeanine Banks and Tris Warkentin · 2024
Later among the works it cites.
Ning Bian et al · 2024
Later among the works it cites.
“On the Relation between Sensitivity and Accuracy in In-context Learning”, 2024
Yanda Chen et al · 2024
Later among the works it cites.
“Retrieval-Augmented Generation for Large Language Models: A Survey”, 2024
Yunfan Gao et al · 2024
Later among the works it cites.
“Robust planning with LLMmodulo framework: Case study in travel planning”
A. Gundawar et al · 2024
Later among the works it cites.
“Self-planning Code Generation with Large Language Models”, 2024
Xue Jiang et al · 2024
Later among the works it cites.
“Can large language models reason and plan?”
Subbarao Kambhampati · 2024
Later among the works it cites.
“LLMs Can’t Plan, But Can Help Planning in LLM-Modulo Frameworks”, 2024
Subbarao Kambhampati et al · 2024
Later among the works it cites.
“A Survey of Large Language Models in Finance (FinLLMs)”, 2024
Jean Lee, Nicholas Stevens, Soyeon Han and Minseok Song · 2024
Later among the works it cites.
“LMStudio” Accessed: 2024-07-26, 2024
LMStudio · 2024
Later among the works it cites.
“OpenAI’s O3 model aced a test of AI reasoning – but it’s still not AGI” Accessed: 2024-06-09, 2024
New Scientist · 2024
Later among the works it cites.
“Learning to Reason with LLMs” Accessed: 2024-06-24, 2024
OpenAI · 2024
Later among the works it cites.
“Code Llama: Open Foundation Models for Code”, 2024
Baptiste Rozi“‘ere et al · 2024
Later among the works it cites.
“Gemma: Open Models Based on Gemini Research and Technology”, 2024
Gemma Team et al · 2024
Later among the works it cites.
M. Verma, S. Bhambri and S. Kambhampati · 2024
Later among the works it cites.
Kevin Wang et al · 2024
Later among the works it cites.
“Bridging the preference gap between retrievers and LLMs”
Z. Wang et al · 2024
Later among the works it cites.
“The Llama 3 Herd of Models” Accessed: 2024-07-25, https://ai.meta.com/research/publications/the-llama-3-herd-of-models/
Meta AI · 2024
Later among the works it cites.
“BigQuery Dataset” Accessed: 2024-04-14, https://cloud.google.com/bigquery?hl=zh-cn
2024
Later among the works it cites.
“Common Crawl” Accessed: 2024-04-15, https://commoncrawl.org/
2024
Later among the works it cites.
“Project Gutenberg” Accessed: 2024-04-14, https://www.gutenberg.org/
2024
Later among the works it cites.
“Wikipedia” Accessed: 2024-04-14, https://en.wikipedia.org/wiki/Main_Page
2024
Later among the works it cites.