Demo: Coming Soon~
Download the datasets:
- PanoVQA:
Google Drive Download Link(removed due to storage limitations) - PanoVQA-mini:
Download Link
[2026.09.26 | Update] PanoVQA is now available on Hugging Face. After downloading, pls unzip the files manually.
Models' checkpoints are available at https://huggingface.co/wakinghours/PLM/tree/main
Please organize your workspace as follows. Ensure all datasets are placed under their corresponding directories before running the code:
Workspace/
├── PLM/ # (Your code directory)
│ ├── plm-finetune/ # Training and inference scripts
│ ├── plm-utils/ # Helper utilities
│ ├── eval_benchmark/ # Evaluation codes
│ └── ...
└── data/ # (Datasets you just downloaded)
├── PanoVQA/
│ ├── BlendPASS/
│ ├── DeepAccident/
│ └── NuScenes/
├── PanoVQA_mini/
│ ├── BlendPASS/
│ ├── DeepAccident/
│ └── NuScenes/
└── Panorama/
└── images/
We recommend using Miniconda to manage your virtual environment.
1. Create and activate a new Conda environment:
conda create -n plm python=3.11 -y
conda activate plm2. Install core dependencies: We recommend installing the required dependencies with the following commands to match the Qwen-VL ecosystem:
pip install transformers==4.51.3 accelerate3. Install the vision processing toolkit:
We offer a toolkit to help you handle various types of visual inputs more conveniently. We highly recommend using the [decord] feature for faster video loading:
pip install qwen-vl-utils[decord]4. Install Flash-Attention 2 (Recommended): For better acceleration and memory savings, especially in multi-image and video scenarios, please install the latest version of Flash Attention 2:
pip install -U flash-attn --no-build-isolation- Configure your data paths in:
./plm-finetune/plm/data/__init__.py- Navigate to the fine-tuning directory and run your training script:
cd plm-finetune
sh scripts/sft_3b.sh # Or run your respective scriptRun inference to generate predictions:
cd eval_benchmark
python eval.script.mini.adaptor.py \
--model_path "../plm-finetune/output/adaptor/Pano_adaptor_former_finetune(adapter,llm,mlp)" \
--save_path "outputs/inferences/sft3b.json"Evaluate the generated predictions using an OpenAI API key:
python get_gpt_score.py \
--input outputs/inferences/sft3b.json \
--output outputs/gpt_score/sft3b.json \
-k "YOUR_OPENAI_API_KEY"Checkpoints can be found in Huggingface
For full transparency and to help you reproduce our results, you can access our comprehensive training logs and GPT evaluation outputs here:
If you find our work helpful, please consider citing:
@InProceedings{fan2026PanoVQA_PLM,
author = {Fan, Weijia and Liu, Ruiping and Wei, Jiale and Chen, Yufan and Zheng, Junwei and Zeng, Zichao and Zhang, Jiaming and Li, Qiufu and Shen, Linlin and Stiefelhagen, Rainer},
title = {More than the Sum: Panorama-Language Models for Adverse Omni-Scenes},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {30874-30884}
}- Update paper
- Update training code
- Update evaluating code
- Update training and evaluation logs
- Release the checkpoints
- Update demo
- Update faster version of PSA
- Submit to HuggingFace (thanks to Niels for the advice)
This work was supported by the Shenzhen University Overseas Exchange Scholarship, which supported my living expenses in Karlsruhe. Thanks to SZU! Huge thanks to KIT and fellows, I had a quite nice experience there.
This work is based on the Qwen-VL repository. Huge thanks to the contributors for their efforts in the community!
