Fetching the paper…
Reading the bibliography…
The rapid advancement of Large Language Models (LLMs) has revolutionized various sectors by automating routine tasks, marking a step toward the realization of Artificial General Intelligence (AGI).
Well-read students learn better: The impact of student initialization on knowledge distillation
Turc, I., Chang, M., Lee, K., and Toutanova, K · 1908
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
Knowledge acquisition and explanation for multi-attribute decision making
Bohanec, M. and Rajkovic, V · 1988
Earlier work this paper cites.
International application of a new probability algorithm for the diagnosis of coronary artery disease
Detrano, R. C., Jánosi, A., Steinbrunn, W., Pfisterer, M. E., Schmid, J.-J., Sandhu, S., Guppy, K., Lee, S., and Froelicher, V · 1989
Earlier work this paper cites.
Nuclear feature extraction for breast tumor diagnosis
Street, W. N., Wolberg, W. H., and Mangasarian, O. L · 1993
Earlier work this paper cites.
Comparative analysis of statistical pattern recognition methods in high dimensional settings
Aeberhard, S., Coomans, D., and de Vel, O. Y · 1994
Earlier work this paper cites.
Modeling wine preferences by data mining from physicochemical properties
Cortez, P., Cerdeira, A. L., Almeida, F., Matos, T., and Reis, J · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Adapting visual category models to new domains
Saenko, K., Kulis, B., Fritz, M., and Darrell, T · 2010
Earlier work this paper cites.
A data-driven approach to predict the success of bank telemarketing
Moro, S., Cortez, P., and Rita, P · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Hypernetworks
Ha, D., Dai, A. M., and Le, Q. V · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Deep visual domain adaptation: A survey
Wang, M. and Deng, W · 2018
Earlier work this paper cites.
Classification of rice varieties using artificial intelligence methods
Cınar, I. and Koklu, M · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Hypergan: A generative model for diverse, performant neural networks
Ratzlaff, N. and Li, F · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Specializing smaller language models towards multi-step reasoning
Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T · 2023
Later among the works it cites.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Giorno, A. D., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T., and Li, Y · 2023
Later among the works it cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Hsieh, C., Li, C., Yeh, C., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C., and Pfister, T · 2023
Later among the works it cites.
HINT: hypernetwork instruction tuning for efficient zero- and few-shot generalisation
Ivison, H., Bhagia, A., Wang, Y., Hajishirzi, H., and Peters, M. E · 2023
Later among the works it cites.
DUET: A tuning-free device-cloud collaborative parameters generation framework for efficient device model generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multiclass classification of dry beans using computer vision and machine learning techniques
Koklu, M. and Özkan, I. A · 2020
Cited alongside, same era.
Distributionally robust neural networks
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2020
Cited alongside, same era.
The iris data set: In search of the source of virginica
Unwin, A. and Kleinman, K · 2021
Cited alongside, same era.
Generalizing from a few examples: A survey on few-shot learning
Wang, Y., Yao, Q., Kwok, J. T., and Ni, L. M · 2021
Cited alongside, same era.
Hyperstyle: Stylegan inversion with hypernetworks for real image editing
Alaluf, Y., Tov, O., Mokady, R., Gal, R., and Bermano, A · 2022
Cited alongside, same era.
Hyperinverter: Improving stylegan inversion via hypernetwork
Dinh, T. M., Tran, A. T., Nguyen, R., and Hua, B · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Lv, Z., Zhang, W., Zhang, S., Kuang, K., Wang, F., Wang, Y., Chen, Z., Shen, T., Yang, H., Ooi, B. C., and Wu, F · 2023
Later among the works it cites.
A comprehensive overview of large language models
Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Barnes, N., and Mian, A · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., et al · 2023
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face, 2023
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Later among the works it cites.
Beyond memorization: Violating privacy via inference with large language models
Staab, R., Vero, M., Balunovic, M., and Vechev, M. T · 2023
Later among the works it cites.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Yang, J., Jin, H., Tang, R., Han, X., Feng, Q., Jiang, H., Yin, B., and Hu, X · 2023
Later among the works it cites.
A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly
Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, E., and Zhang, Y · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J., and Wen, J · 2023
Later among the works it cites.
A survey of resource-efficient llm and multimodal foundation models, 2024
Xu, M., Yin, W., Cai, D., Yi, R., Xu, D., Wang, Q., Wu, B., Zhao, Y., Yang, C., Wang, S., Zhang, Q., Lu, Z., Zhang, L., Wang, S., Li, Y., Liu, Y., Jin, X., and Liu, X · 2024
Closest in time.