Skip to content

About

No description, website, or topics provided.

Resources

Stars

18 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

FineMoE-EuroSys26


This repo contains a demo implementation of the paper, Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading.

Large Language Models (LLMs) have gained immense success in revolutionizing various applications, including content generation, search and recommendation, and AI-assisted operation. To reduce high training costs, Mixture-of-Experts (MoE) architecture has become a popular backbone for modern LLMs. However, despite the benefits, serving MoE-based LLMs experience severe memory inefficiency due to sparsely activated experts. Recent studies propose to offload inactive experts from GPU memory to CPU memory to improve the serving efficiency of MoE models. However, they either incur high inference latency or high model memory footprints due to coarse-grained designs. To tame the latency-memory trade-off in MoE serving, we present FineMoE, a fine-grained expert offloading system for MoE serving that achieves low inference latency with memory efficiency. We design FineMoE to extract fine-grained expert selection patterns from MoE models and semantic hints from input prompts to efficiently guide expert prefetching, caching, and offloading decisions. FineMoE is prototyped on top of HuggingFace Transformers and deployed on a six-GPU testbed. Experiments with open-source MoE models and real-world workloads show that FineMoE reduces inference latency by 47% and improves expert hit rate by 39% over state-of-the-art solutions.


FineMoE is built on top of MoE-Infinity. We thank the MoE-Infinity team for their codebase!

We describe how to build and run this demo.

General Hardware Prerequisite

  • Operating system: Linux x86-64 (tested on Ubuntu)
  • Python: 3.10–3.12
  • Build: C++17 compiler
  • GPU: one or more CUDA GPUs; RTX 3090 with 24 GB memory supported
  • CPU: >= 8 cores
  • Host memory: 192 GB recommended for the checkpoint and pinned expert weights
  • Disk: >= 160 GB for the original checkpoint and prepared weights
  • Network: required to download dependencies and the checkpoint

Demo Instructions

Run the commands below from the repository root.

  1. Download the GitHub repo.
git clone https://github.com/IntelliSys-Lab/FineMoE-EuroSys26
cd FineMoE-EuroSys26

  1. Install uv and the locked CUDA dependencies.
./setup.sh

  1. Select the GPUs in config_common.py and review the settings in demo/configs.
devices = ["cuda:0"]

For two GPUs, set devices = ["cuda:0", "cuda:1"]. Indices refer to GPUs visible to the process.

  1. Prepare the model and gather data for the demo.
uv run --locked python -m demo.prepare_data

  1. Process the collected data.
uv run --locked python -m demo.process_data

  1. Execute the demo.
uv run --locked python -m demo.eval

Results and Figures

The demo writes entropy and heatmap CSV files and serving metrics under demo/results. Plot the results with Matplotlib:

uv run --locked python -m demo.plot_entropy

Figures are saved under demo/figures.

Experimental Settings and Workloads

Experiment settings are in demo/configs. This demo serves Qwen3.5-35B-A3B (text only) on a small sample of lmsys-chat-1m. The dataset sample is in demo/states/lmsys-chat-1m~eval_prompts.json.

About

No description, website, or topics provided.

Resources

Stars

18 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages