This repo contains a demo implementation of the paper, Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading.
Large Language Models (LLMs) have gained immense success in revolutionizing various applications, including content generation, search and recommendation, and AI-assisted operation. To reduce high training costs, Mixture-of-Experts (MoE) architecture has become a popular backbone for modern LLMs. However, despite the benefits, serving MoE-based LLMs experience severe memory inefficiency due to sparsely activated experts. Recent studies propose to offload inactive experts from GPU memory to CPU memory to improve the serving efficiency of MoE models. However, they either incur high inference latency or high model memory footprints due to coarse-grained designs. To tame the latency-memory trade-off in MoE serving, we present FineMoE, a fine-grained expert offloading system for MoE serving that achieves low inference latency with memory efficiency. We design FineMoE to extract fine-grained expert selection patterns from MoE models and semantic hints from input prompts to efficiently guide expert prefetching, caching, and offloading decisions. FineMoE is prototyped on top of HuggingFace Transformers and deployed on a six-GPU testbed. Experiments with open-source MoE models and real-world workloads show that FineMoE reduces inference latency by 47% and improves expert hit rate by 39% over state-of-the-art solutions.
FineMoE is built on top of MoE-Infinity. We thank the MoE-Infinity team for their codebase!
We describe how to build and run this demo.
- Operating system: Linux x86-64 (tested on Ubuntu)
- Python: 3.10–3.12
- Build: C++17 compiler
- GPU: one or more CUDA GPUs; RTX 3090 with 24 GB memory supported
- CPU: >= 8 cores
- Host memory: 192 GB recommended for the checkpoint and pinned expert weights
- Disk: >= 160 GB for the original checkpoint and prepared weights
- Network: required to download dependencies and the checkpoint
Run the commands below from the repository root.
- Download the GitHub repo.
git clone https://github.com/IntelliSys-Lab/FineMoE-EuroSys26
cd FineMoE-EuroSys26- Install uv and the locked CUDA dependencies.
./setup.sh- Select the GPUs in
config_common.pyand review the settings indemo/configs.
devices = ["cuda:0"]For two GPUs, set devices = ["cuda:0", "cuda:1"]. Indices refer to GPUs visible to the process.
- Prepare the model and gather data for the demo.
uv run --locked python -m demo.prepare_data- Process the collected data.
uv run --locked python -m demo.process_data- Execute the demo.
uv run --locked python -m demo.evalThe demo writes entropy and heatmap CSV files and serving metrics under demo/results. Plot the results with Matplotlib:
uv run --locked python -m demo.plot_entropyFigures are saved under demo/figures.
Experiment settings are in demo/configs.
This demo serves Qwen3.5-35B-A3B (text only) on a small sample of lmsys-chat-1m.
The dataset sample is in demo/states/lmsys-chat-1m~eval_prompts.json.