Fetching the paper…
Reading the bibliography…
The temperature parameter plays a profound role during training and/or inference with large foundation models (LFMs) such as large language models (LLMs) and CLIP models.
Information theory and statistical mechanics
Jaynes, E. T · 1957
Earlier work this paper cites.
On general minimax theorems
Sion, M · 1958
Earlier work this paper cites.
The need for biases in learning generalizations
Mitchell, T. M · 1980
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Variational analysis , volume 317
Rockafellar, R. T. and Wets, R. J.-B · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Toward controlled generation of text
Hu, Z., Yang, Z., Liang, X., Salakhutdinov, R., and Xing, E. P · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Softmax q-distribution estimation for structured prediction: A theoretical interpretation for raml
Ma, X., Yin, P., Liu, J., Neubig, G., and Hovy, E · 2017
Earlier work this paper cites.
Determining the optimal temperature parameter for softmax function in reinforcement learning
He, Y.-L., Zhang, X.-L., Ao, W., and Huang, J. Z · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Zhang, X., Yu, F. X., Karaman, S., Zhang, W., and Chang, S · 2018
Earlier work this paper cites.
Adaptive temperature tuning for mellowmax in deep reinforcement learning
Kim, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Cited alongside, same era.
Subjqa: a dataset for subjectivity and review comprehension
Bjerva, J., Bhutani, N., Golshan, B., Tan, W.-C., and Augenstein, I · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Cyclip: Cyclic contrastive language-image pretraining
Goel, S., Bansal, H., Bhatia, S., Rossi, R., Vinay, V., and Grover, A · 2022
Later among the works it cites.
Liu, J., Liu, B., Li, H., and Liu, Y · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions, 2022
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Later among the works it cites.
A systematic evaluation of large language models of code
Xu, F. F., Alon, U., Neubig, G., and Hellendoorn, V. J · 2022
Later among the works it cites.
Provable stochastic optimization for global contrastive learning: Small batch does not harm performance
Yuan, Z., Wu, Y., Qiu, Z.-H., Du, X., Zhang, L., Zhou, D., and Yang, T · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Contextual temperature for language modeling
Wang, P.-H., Hsieh, S.-I., Chang, S.-C., Chen, Y.-T., Pan, J.-Y., Wei, W., and Juan, D.-C · 2020
Cited alongside, same era.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch, 9 2023
Andonian, A., Anthony, Q., Biderman, S., Black, S., Gali, P., Gao, L., Hallahan, E., Levy-Kramer, J., Leahy, C., Nestler, L., Parker, K., Pieler, M., Phang, J., Purohit, S., Schoelkopf, H., Stander, D., Songz, T., Tigges, C., Thérien, B., Wang, P., and Weinbach, S · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Later among the works it cites.
Temperature-scaled large language models for lean proofstep prediction
Gloeckle, F., Roziere, B., Hayat, A., and Synnaeve, G · 2023
Later among the works it cites.
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al · 2023
Later among the works it cites.
Sample-dependent adaptive temperature scaling for improved calibration
Joy, T., Pinto, F., Lim, S.-N., Torr, P. H., and Dokania, P. K · 2023
Later among the works it cites.
Temperature schedules for self-supervised contrastive methods on long-tail data
Kukleva, A., Böhle, M., Schiele, B., Kuehne, H., and Rupprecht, C · 2023
Later among the works it cites.
Dystress: Dynamically scaled temperature in self-supervised contrastive learning
Manna, S., Chattopadhyay, S., Dey, R., Bhattacharya, S., and Pal, U · 2023
Later among the works it cites.
Stochastic constrained dro with a complexity independent of sample size
Qi, Q., Lyu, J., sik Chan, K., Bai, E. W., and Yang, T · 2023
Later among the works it cites.
Not all semantics are created equal: Contrastive self-supervised learning with automatic temperature individualization
Qiu, Z.-H., Hu, Q., Yuan, Z., Zhou, D., Zhang, L., and Yang, T · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Label distributionally robust losses for multi-class classification: Consistency, robustness and adaptivity
Zhu, D., Ying, Y., and Yang, T · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2024
Closest in time.