Fetching the paper…
Reading the bibliography…
For industrial control, developing high-performance controllers with few samples and low technical debt is appealing.
Dynamic programming
Bellman, R. (1966) · 1966
Earlier work this paper cites.
Tutorial overview of model predictive control
Rawlings, J. B. (2000) · 2000
Earlier work this paper cites.
Fuzzy control of hvac systems optimized by genetic algorithms
Alcalá, R., Benítez, J. M., Casillas, J., Cordón, O., and Pérez, R. (2003) · 2003
Earlier work this paper cites.
Dynamic model of an hvac system for control analysis
Tashtoush, B., Molhim, M., and Al-Rousan, M. (2005) · 2005
Earlier work this paper cites.
Outliers: The story of success
Gladwell, M. (2008) · 2008
Earlier work this paper cites.
Application of an intelligent pid control in heating ventilating and air-conditioning system
Wang, J., Zhang, C., and Jing, Y. (2008) · 2008
Earlier work this paper cites.
Design and application of handheld auto-tuning pid instrument used in hvac
Liu, J., Cai, W.-J., and Zhang, G.-Q. (2009) · 2009
Earlier work this paper cites.
Model predictive control of thermal energy storage in building cooling systems
Ma, Y., Borrelli, F., Hencey, B., Packard, A., and Bortoff, S. (2009) · 2009
Earlier work this paper cites.
A fuzzy logic based efficient energy saving approach for domestic heating systems
Villar, J. R., de la Cal, E., and Sedano, J. (2009) · 2009
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. (2020) · 2010
Earlier work this paper cites.
Smart grid controller for optimizing hvac energy consumption
Al-Ali, A., Tubaiz, N. A., Al-Radaideh, A., Al-Dmour, J. A., and Murugan, L. (2012) · 2012
Earlier work this paper cites.
Predictive control for energy efficient buildings with thermal storage: Modeling, stimulation, and experiments
Ma, Y., Kelman, A., Daly, A., and Borrelli, F. (2012) · 2012
Earlier work this paper cites.
A review of intelligent control techniques in hvac systems
Mirinejad, H., Welch, K. C., and Spicer, L. (2012) · 2012
Earlier work this paper cites.
API design for machine learning software: experiences from the scikit-learn project
Buitinck, L., Louppe, G., Blondel, M., Pedregosa, F., Mueller, A., Grisel, O., Niculae, V., Prettenhofer, P., Gramfort, A., Grobler, J., Layton, R., VanderPlas, J., Joly, A., Holt, B., and Varoquaux, G. (2013) · 2013
Earlier work this paper cites.
An efficient design of genetic algorithm based adaptive fuzzy logic controller for multivariable control of hvac systems
Khan, M. W., Choudhry, M. A., and Zeeshan, M. (2013) · 2013
Earlier work this paper cites.
Theory and applications of hvac control systems–a review of model predictive control (mpc)
Afram, A. and Janabi-Sharifi, F. (2014) · 2014
Earlier work this paper cites.
Hvac control methods-a review
Belic, F., Hocenski, Z., and Sliskovic, D. (2015) · 2015
Earlier work this paper cites.
Reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., Sunehag, P., and Coppin, B. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Making contextual decisions with low technical debt
Agarwal, A., Bird, S., Cozowicz, M., Hoang, L., Langford, J., Lee, S., Li, J., Melamed, D., Oshri, G., Ribas, O., et al. (2016) · 2016
Earlier work this paper cites.
Model predictive control: theory, computation, and design
Rawlings, J. B., Mayne, D. Q., and Diehl, M. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Deep reinforcement learning for building hvac control
Wei, T., Wang, Y., and Zhu, Q. (2017) · 2017
Earlier work this paper cites.
Modeling techniques used in building hvac control systems: A review
Afroz, Z., Shafiullah, G., Urmee, T., and Higgins, G. (2018) · 2018
Earlier work this paper cites.
Universal sentence encoder
Cer, D., Yang, Y., yi Kong, S., Hua, N., Limtiaco, N., John, R. S., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Sung, Y.-H., Strope, B., and Kurzweil, R. (2018) · 2018
Earlier work this paper cites.
Babyai: A platform to study the sample efficiency of grounded language learning
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y. (2018) · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning, fast and slow
Botvinick, M., Ritter, S., Wang, J. X., Kurth-Nelson, Z., Blundell, C., and Hassabis, D. (2019) · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D., Goo, W., Nagarajan, P., and Niekum, S. (2019) · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Cited alongside, same era.
Reinforcement learning for whole-building hvac control and demand response
Azuatalam, D., Lee, W.-L., de Nijs, F., and Liebman, A. (2020) · 2020
Cited alongside, same era.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
Brown, D. S., Goo, W., and Niekum, S. (2020) · 2020
Cited alongside, same era.
Rlbench: The robot learning benchmark & learning environment
Chatgpt: Optimizing language models for dialogue
OpenAI (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. (2022) · 2022
Later among the works it cites.
Black-box tuning for language-model-as-a-service
Sun, T., Shao, Y., Qian, H., Huang, X., and Qiu, X. (2022) · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J. (2020) · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021) · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N. (2021) · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P. (2021) · 2021
Cited alongside, same era.
What makes good in-context examples for gpt- 3 3 ?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., and Chen, W. (2021) · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022) · 2022
Later among the works it cites.
Claude 2
Antropic (2023) · 2023
Closest in time.
Open llm leaderboard
Beeching, E., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. (2023) · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. (2023) · 2023
Closest in time.
Grounding large language models in interactive environments with online reinforcement learning
Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y. (2023) · 2023
Closest in time.
Learning universal policies via text-guided video generation
Dai, Y., Yang, M., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P. (2023) · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. (2023) · 2023
Closest in time.
Guiding pretraining in reinforcement learning with large language models
Du, Y., Watkins, O., Wang, Z., Colas, C., Darrell, T., Abbeel, P., Gupta, A., and Andreas, J. (2023) · 2023
Closest in time.
The false promise of imitating proprietary llms
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D. (2023) · 2023
Closest in time.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z. (2023) · 2023
Closest in time.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D. (2023) · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
An important next step on our ai journey
Pichai, S. (2023) · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A. (2023) · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. (2023) · 2023
Closest in time.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
Yuan, H., Zhang, C., Wang, H., Xie, F., Cai, P., Dong, H., and Lu, Z. (2023) · 2023
Closest in time.
Towards generalizable reinforcement learning for trade execution
Zhang, C., Duan, Y., Chen, X., Chen, J., Li, J., and Zhao, L. (2023) · 2023
Closest in time.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al. (2023) · 2023
Closest in time.
Chatbot arena
Zheng, L., Sheng, Y., Chiang, W.-L., Li, D., Li, Z., Lin, Z., Wu, Z., Zhuang, S., Zhuang, Y., Zhang, H., Stoica, I., Gonzalez, J. E., and Xing, E. P. (2023) · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N. (2023) · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., et al. (2023) · 2023
Closest in time.
Zhu, X., Chen, Y., Tian, H., Tao, C., Su, W., Yang, C., Huang, G., Li, B., Lu, L., Wang, X., et al. (2023) · 2023
Closest in time.
A distributed predictive control approach to building temperature regulation
Ma, Y., Anderson, G., and Borrelli, F. (2011) · 2094
Closest in time.