Fetching the paper…
Reading the bibliography…
We present the Mixture-of-Tunable-Experts (MoTE), a method that extends the Mixture-of-Experts architecture of Large Language Models (LLMs).
“A Perspective on Judgment and Choice: Mapping Bounded Rationality”
Daniel Kahneman · 2003
Earlier work this paper cites.
“Visualizing data using t-SNE.”
Laurens Van and Geoffrey Hinton · 2008
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Glam: Efficient scaling of language models with mixture-of-experts”
Nan Du et al · 2022
Earlier work this paper cites.
“A Review of Sparse Expert Models in Deep Learning”, 2022
William Fedus, Jeff Dean and Barret Zoph · 2022
Earlier work this paper cites.
“Towards Monosemanticity: Decomposing Language Models With Dictionary Learning”
Trenton Bricken et al · 2023
Earlier work this paper cites.
“Efficient Memory Management for Large Language Model Serving with PagedAttention”
Woosuk Kwon et al · 2023
Cited alongside, same era.
“Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, 2023
Lianmin Zheng et al · 2023
Cited alongside, same era.
“Scaling and evaluating sparse autoencoders”, 2024
Leo Gao et al · 2024
Cited alongside, same era.
“Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet”
Adly Templeton et al · 2024
Cited alongside, same era.
“Cheaper, Better, Faster, Stronger — Mistral AI”, 2024
Mistral team · 2024
Cited alongside, same era.
“DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models”, 2024
“DeepSeek-V3 Technical Report”, 2024
DeepSeek-AI et al · 2024
Later among the works it cites.
Albert. Jiang et al · 2024
Later among the works it cites.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, 2025
DeepSeek-AI et al · 2025
Closest in time.
“CCP Sensitive Prompts (Huggingface Dataset)” Accessed: 2025-02-12, 2025
Promptfoo · 2025
Closest in time.
“1,156 Questions Censored by DeepSeek — promptfoo”, 2025
Deepseekjanuary · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Damai Dai et al · 2024
Cited alongside, same era.
We intend to release our modifications to vLLM in the future
Cited in the paper.