Fetching the paper…
Reading the bibliography…
We present Language-Image Value learning (LIV), a unified objective for vision-language representation and reward learning from action-free videos with text annotations.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation, and machine learning , volume 133
Rubinstein, R. Y. and Kroese, D. P · 2004
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense, 2017
Goyal, R., Kahou, S. E., Michalski, V., Materzyńska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., and Memisevic, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Model predictive path integral control: From theory to parallel computation
Williams, G., Aldrich, A., and Theodorou, E. A · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G · 2018
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
Lynch, C. and Sermanet, P · 2020
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
Stepputtis, S., Campbell, J., Phielipp, M., Lee, S., Baral, C., and Ben Amor, H · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards
Goyal, P., Niekum, S., and Mooney, R · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Bc-z: Zero-shot task generalization with robotic imitation learning
Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C · 2022
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai
Khandelwal, A., Weihs, L., Mottaghi, R., and Kembhavi, A · 2022
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G · 2022
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P · 2022
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
LeCun, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rrl: Resnet as representation for reinforcement learning
Shah, R. and Kumar, V · 2021
Cited alongside, same era.
Multi-task reinforcement learning with context-based representations
Sodhani, S., Zhang, A., and Pineau, J · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Cited alongside, same era.
Can foundation models perform zero-shot task specification for robot manipulation?
Cui, Y., Niekum, S., Gupta, A., Kumar, V., and Rajeswaran, A · 2022
Cited alongside, same era.
Dong, X., Bao, J., Zhang, T., Chen, D., Gu, S., Zhang, W., Yuan, L., Chen, D., Wen, F., and Yu, N · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Cited alongside, same era.
Lee, Y., Chen, A. S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C · 2022
Later among the works it cites.
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Later among the works it cites.
Instruction-following agents with jointly pre-trained vision-language models
Liu, H., Lee, L., Lee, K., and Abbeel, P · 2022
Later among the works it cites.
Interactive language: Talking to robots in real time
Lynch, C., Wahid, A., Tompson, J., Ding, T., Betker, J., Baruch, R., Armstrong, T., and Florence, P · 2022
Later among the works it cites.
Zero-shot reward specification via grounded natural language
Mahmoudieh, P., Pathak, D., and Darrell, T · 2022
Later among the works it cites.
The unsurprising effectiveness of pre-trained vision models for control
Parisi, S., Rajeswaran, A., Purushwalkam, S., and Gupta, A · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Later among the works it cites.
Masked visual pre-training for motor control
Xiao, T., Radosavovic, I., Darrell, T., and Malik, J · 2022
Later among the works it cites.
Language-driven representation learning for robotics
Karamcheti, S., Nair, S., Chen, A. S., Kollar, T., Finn, C., Sadigh, D., and Liang, P · 2023
Closest in time.
Where are we in the search for an artificial visual cortex for embodied intelligence?
Majumdar, A., Yadav, K., Arnaud, S., Ma, Y. J., Chen, C., Silwal, S., Jain, A., Berges, V.-P., Abbeel, P., Malik, J., et al · 2023
Closest in time.