Fetching the paper…
Reading the bibliography…
Online shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Estimating reactions and recommending products with generative models of reviews
J. Ni, Z. C. Lipton, S. Vikram, and J. McAuley · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
M. Post · 2018
Earlier work this paper cites.
Opentag: Open attribute value extraction from product profiles
G. Zheng, S. Mukherjee, X. L. Dong, and F. Li · 2018
Earlier work this paper cites.
Amazonqa: A review-based question answering task
M. Gupta, N. Kulkarni, R. Chanda, A. Rayasam, and Z. C. Lipton · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
J. Ni, J. Li, and J. McAuley · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
WINOGRANDE: an adversarial winograd schema challenge at scale, 2019
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2019
Earlier work this paper cites.
Session-based recommendation with graph neural networks
S. Wu, Y. Tang, Y. Zhu, L. Wang, X. Xie, and T. Tan · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Learning to extract attribute value from product via question answering: A multi-task approach
Q. Wang, L. Yang, B. Kanagal, S. Sanghai, D. Sivakumar, B. Shu, Z. Yu, and J. Elsas · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Earlier work this paper cites.
Tacc: A full-stack cloud computing infrastructure for machine learning tasks
K. Xu, X. Wan, H. Wang, Z. Ren, X. Liao, D. Sun, C. Zeng, and K. Chen · 2021
Cited alongside, same era.
A survey on multi-task learning
Y. Zhang and Q. Yang · 2021
Cited alongside, same era.
Retrieval-augmented multilingual keyphrase generation with retriever-generator iterative training
Y. Gao, Q. Yin, Z. Li, R. Meng, T. Zhao, B. Yin, I. King, and M. Lyu · 2022
Cited alongside, same era.
Short text pre-training with extended token classification for e-commerce query understanding
H. Jiang, T. Cao, Z. Li, C. Luo, X. Tang, Q. Yin, D. Zhang, R. Goutam, and B. Yin · 2022
Cited alongside, same era.
Query rewriting in taobao search
S. Li, F. Lv, T. Jin, G. Li, Y. Zheng, T. Zhuang, Q. Liu, X. Zeng, J. Kwok, and Q. Ma · 2022
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Later among the works it cites.
Amazon-m2: A multilingual multi-locale shopping session dataset for recommendation and text generation
W. Jin, H. Mao, Z. Li, H. Jiang, C. Luo, H. Wen, H. Han, H. Lu, Z. Wang, R. Li, et al · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier, et al · 2023
Later among the works it cites.
Text is all you need: Learning language representations for sequential recommendation
J. Li, M. Wang, J. Li, J. Fu, X. Shen, J. Shang, and J. McAuley · 2023
Later among the works it cites.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Truthfulqa: Measuring how models mimic human falsehoods, 2022
S. Lin, J. Hilton, and O. Evans · 2022
Cited alongside, same era.
Query attribute recommendation at amazon search
C. Luo, W. Headden, N. Avudaiappan, H. Jiang, T. Cao, Q. Yin, Y. Gao, Z. Li, R. Goutam, H. Zhang, et al · 2022
Cited alongside, same era.
Shopping queries dataset: A large-scale esci benchmark for improving product search
C. K. Reddy, L. Màrquez, F. Valero, N. Rao, H. Zaragoza, S. Bandyopadhyay, A. Biswas, A. Xing, and K. Subbian · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Symbolic knowledge distillation: from general language models to commonsense models
P. West, C. Bhagavatula, J. Hessel, J. Hwang, L. Jiang, R. Le Bras, X. Lu, S. Welleck, and Y. Choi · 2022
Cited alongside, same era.
Some practice for improving the search results of e-commerce
F. Wu, Y. Liu, R. Gazo, B. Bedrich, and X. Qu · 2022
Cited alongside, same era.
Pyabsa: Open framework for aspect-based sentiment analysis, 2022
H. Yang and K. Li · 2022
Cited alongside, same era.
Later among the works it cites.
Enhancing user intent capture in session-based recommendation with attribute patterns
X. Liu, Z. Li, Y. Gao, J. Yang, T. Cao, Z. Wang, B. Yin, and Y. Song · 2023
Later among the works it cites.
Code llama: Open foundation models for code
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Gpt-ner: Named entity recognition via large language models
S. Wang, X. Sun, X. Li, R. Ouyang, F. Wu, T. Zhang, J. Li, and G. Wang · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, et al · 2023
Later among the works it cites.
Introducing the next generation of Claude, 3 2024
Anthropic · 2024
Closest in time.
Bridging language and items for retrieval and recommendation
Y. Hou, J. Li, Z. He, A. Yan, X. Chen, and J. McAuley · 2024
Closest in time.
Large language models are zero-shot rankers for recommender systems
Y. Hou, J. Zhang, Z. Lin, H. Lu, R. Xie, J. McAuley, and W. X. Zhao · 2024
Closest in time.
Ecomgpt: Instruction-tuning large language models with chain-of-task tasks for e-commerce
Y. Li, S. Ma, X. Wang, S. Huang, C. Jiang, H.-T. Zheng, P. Xie, F. Huang, and Y. Jiang · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date, 4 2024
Meta-AI-Research · 2024
Closest in time.
B. Peng, X. Ling, Z. Chen, H. Sun, and X. Ning · 2024
Closest in time.
Llmrec: Large language models with graph augmentation for recommendation
W. Wei, X. Ren, J. Tang, Q. Wang, L. Su, S. Cheng, J. Wang, D. Yin, and C. Huang · 2024
Closest in time.
Can large language models transform computational social science?
C. Ziems, W. Held, O. Shaikh, J. Chen, Z. Zhang, and D. Yang · 2024
Closest in time.