Fetching the paper…
Reading the bibliography…
Transformer structure has achieved great success in multiple applied machine learning communities, such as natural language processing (NLP), computer vision (CV) and information retrieval (IR).
“Long short-term memory”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“The probabilistic relevance framework: BM25 and beyond”
Stephen Robertson and Hugo Zaragoza · 2009
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality”
Tomas Mikolov et al · 2013
Earlier work this paper cites.
“Efficient index-based snippet generation”
Hannah Bast and Marjan Celikik · 2014
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Parallelizing Linear Recurrent Neural Nets Over Sequence Length”
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
“Deep Contextualized Word Representations”
Matthew. Peters et al · 2018
Earlier work this paper cites.
“GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”
Alex Wang et al · 2018
Earlier work this paper cites.
“Asking clarifying questions in open-domain information-seeking conversations”
Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani and W Croft · 2019
Earlier work this paper cites.
“Deeper text understanding for IR with contextual neural language modeling”
Zhuyun Dai and Jamie Callan · 2019
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“Longformer: The long-document transformer”
Iz Beltagy, Matthew Peters and Arman Cohan · 2020
Earlier work this paper cites.
“Overview of the TREC 2019 deep learning track”
Nick Craswell et al · 2020
Earlier work this paper cites.
“Hippo: Recurrent memory with optimal polynomial projections”
Albert Gu et al · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel et al · 2020
Earlier work this paper cites.
“Overview of the TREC 2020 deep learning track. CoRR abs/2102.07662 (2021)”
Nick Craswell, Bhaskar Mitra, Emine Yilmaz and Daniel Campos · 2021
Earlier work this paper cites.
“Rethink training of BERT rerankers in multi-stage retrieval pipeline”
Luyu Gao, Zhuyun Dai and Jamie Callan · 2021
Cited alongside, same era.
“Efficiently Modeling Long Sequences with Structured State Spaces”
Albert Gu, Karan Goel and Christopher Re · 2021
Cited alongside, same era.
“Combining recurrent, convolutional, and continuous-time models with linear state space layers”
Albert Gu et al · 2021
Cited alongside, same era.
“Intra-document cascading: learning to select passages for neural document ranking”
Sebastian Hofstätter et al · 2021
Cited alongside, same era.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward Hu et al · 2021
Cited alongside, same era.
“Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations”
Jimmy Lin et al · 2021
“Towards explainable search results: a listwise explanation generator”
Puxuan Yu, Razieh Rahimi and James Allan · 2022
Later among the works it cites.
“Opt: Open pre-trained transformer language models”
Susan Zhang et al · 2022
Later among the works it cites.
“Pythia: A suite for analyzing large language models across training and scaling”
Stella Biderman et al · 2023
Later among the works it cites.
“Mamba: Linear-time sequence modeling with selective state spaces”
Albert Gu and Tri Dao · 2023
Later among the works it cites.
“PARADE: Passage Representation Aggregation forDocument Reranking”
Canjia Li et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Understanding the effectiveness of reviews in e-commerce top-n recommendation”
Zhichao Xu, Hansi Zeng and Qingyao Ai · 2021
Cited alongside, same era.
“A zero attentive relevance matching network for review modeling in recommendation system”
Hansi Zeng, Zhichao Xu and Qingyao Ai · 2021
Cited alongside, same era.
Leonid Boytsov et al · 2022
Cited alongside, same era.
“Scaling instruction-finetuned language models”
Hyung Chung et al · 2022
Cited alongside, same era.
“Diagonal state spaces are as effective as structured state spaces”
Ankit Gupta, Albert Gu and Jonathan Berant · 2022
Cited alongside, same era.
“On the parameterization and initialization of diagonal state space models”
Albert Gu, Karan Goel, Ankit Gupta and Christopher Ré · 2022
Cited alongside, same era.
Xueguang Ma et al · 2023
Later among the works it cites.
“RWKV: Reinventing RNNs for the Transformer Era”
Bo Peng et al · 2023
Later among the works it cites.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Later among the works it cites.
“An in-depth investigation of user response simulation for conversational search”
Zhenduo Wang, Zhichao Xu, Qingyao Ai and Vivek Srikumar · 2023
Later among the works it cites.
“Reward-free Policy Imitation Learning for Conversational Search”
Zhenduo Wang, Zhichao Xu and Qingyao Ai · 2023
Later among the works it cites.
“A lightweight constrained generation alternative for query-focused summarization”
Zhichao Xu and Daniel Cohen · 2023
Later among the works it cites.
“Counterfactual Editing for Search Result Explanation”
Zhichao Xu et al · 2023
Later among the works it cites.
“Context-aware Decoding Reduces Hallucination in Query-focused Summarization”
Zhichao Xu · 2023
Later among the works it cites.
“Rankt5: Fine-tuning t5 for text ranking with ranking losses”
Honglei Zhuang et al · 2023
Later among the works it cites.
“FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”
Tri Dao · 2024
Closest in time.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2024
Closest in time.
“Multi-dimensional Evaluation of Empathetic Dialog Responses”
Zhichao Xu and Jiepu Jiang · 2024
Closest in time.