MobileMem is a comprehensive benchmarking framework for evaluating on-device memory systems in realistic mobile environments.
- 📑 Table of Contents
- 🔥 News
- 🎯 Applications
- 📊 Dataset Structure
- 🗂️ Project Structure
- 🚀 Getting Started
- 🔍 Analyzing Failures with MemTrace
- 🚩 Citation
- 2026-09-30 — We released a structured Chinese version of the dataset and benchmark.
- 2026-08-18 — We have officially released our technical report.
- 2026-08-01 — We publicly release the MobileMem dataset.
- 2026-05-16 — We launch the English version of the dataset and benchmark.
- 2026-05-03 — We launch the Chinese version of the dataset and benchmark.
MobileMem is built from multiple heterogeneous sources to enable comprehensive on-device memory modeling.
MobileMem contains three complementary tracks:
| Track | Modality | Description |
|---|---|---|
| text | Text | Long-horizon user–assistant conversations and structured mobile-app events for evaluating textual memory systems. |
| omni | Text and images | Multimodal mobile interactions with screenshots and photos. |
| struct | Structured data | Simulated structured data from on-device applications. |
The dataset is available for download at HuggingFace.
The repository is organized into three main tracks. Track-specific details remain in each subdirectory.
MobileMem/
│
├── text/ # 📖 Text Track
│ ├── README.md # Track-specific guide and dataset download
│ ├── keme/ # 🛠️ KEME synthesis pipeline code
│ └── eval/ # ⚙️ Evaluation scripts for text track
│
├── omni/ # 🖼️ Omni Track
│ ├── README.md # Track-specific guide and dataset download
│ ├── src/ # 🛠️ Data construction pipeline code
│ └── eval/ # ⚙️ Evaluation scripts for omni track
│
├── struct/ # 🧩 Struct Track
│ ├── README.md # Track-specific guide and dataset download
│ ├── construct/ # 🛠️ Example generation and review (optional)
│ └── eval/ # ⚙️ Evaluation for struct track
│
└── README.md # This fileMobileMem offers three benchmark tracks. Choose the path that fits your needs and navigate to the corresponding resources.
For an interactive visualization of the MobileMem data, visit the Dataset Explorer branch.
The textual benchmark for evaluating memory systems on long-term, knowledge-intensive mobile agent trajectories.
| Section | Description | Quick Link |
|---|---|---|
| 📥 Data Access | Download the synthesized KEME trajectories and QA pairs from HuggingFace. | Link |
| ⚙️ How to Evaluate | Detailed evaluation guide for reproducing leaderboard results is available in the MemBase repository. | Link |
| 🛠️ Data Construction | Reproduce the KEME synthesis pipeline from scratch. | Link |
The multimodal benchmark for evaluating on-device memory with realistic mobile images and dialogues.
| Section | Description | Quick Link |
|---|---|---|
| 📥 Data Access | Download the MobileMem-Omni dataset, including images and dialogues. | Link |
| ⚙️ How to Evaluate | Detailed evaluation guide for reproducing leaderboard results is available in the MemBase repository. | Link |
| 🛠️ Data Construction | Rebuild the entire MobileMem-Omni dataset with the provided pipeline. | Link |
Simulated structured data from on-device applications.
| Section | Description | Quick Link |
|---|---|---|
| 📥 Data Access | Download Struct cases and structured app records from HuggingFace. | Link |
| ⚙️ How to Evaluate | Evaluate agent JSONL traces against Struct cases. | Link |
| 🛠️ Data Construction | Example generation and review Skills. | Link |
We recommend using MemTrace to perform an in-depth error analysis. MemTrace helps you visualize and diagnose where and why your memory system fails, making it easier to pinpoint areas for improvement. For an example of how to use MemTrace, please refer to the tutorial in MemBase.
If this work or datasets is helpful, please kindly cite as this:
@techreport{mobilemem,
title={MobileMem: Learning from a Year of Mobile Experiences},
author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
year={2026},
institution={OPPO and OpenKG},
eprint={2608.13606},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.13606},
}