📄 Paper | 🌐 Project Page | 📘 Dataset and Benchmarks
Official repository for the SAFIRE.
Authors: Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer.
- [21 August 2026] SAFIRE is accepted to EMNLP 2026 as Findings! Project page is now live.
SAFIRE (Safety-Aware Fire-smoke Reasoning Evaluation) is the first large-scale benchmark designed to evaluate Multimodal Large Language Models (MLLMs) in safety-critical fire and smoke scenarios.
Current benchmarks often fail to distinguish between critical fires (e.g., Residential scenario fires) and benign ones (e.g., campfires), or misinterpret water vapor, haze, cloud as smoke. SAFIRE is designed to test models' specific context-aware reasoning capabilities on fire and smoke.
SAFIRE addresses this by providing:
- 83K high-quality images across 20 real-world scenarios.
- 193K Multiple-Choice VQA (MCVQA) pairs.
- 10 Reasoning Dimensions:
| 🔢 Target Counting | 🏷️ Classification |
| 🧠 General Reasoning | 🔥 Fire/Smoke Intention |
| 😨 Emotional Response | 📖 Linguistic Polysemy |
| 👁️ Fire/smoke Attributes | 📐 Spatial Correlation |
| 📍 Position Identification | 🚫 Human Presence |
The dataset is categorized into 5 Groups and 20 Scenarios:
| Group | Scenarios |
|---|---|
| 🌿 Natural Phenomena | Grassland, Volcano, Forest, Meteor |
| 🏭 Industrial Operations | Aerospace assets, Flare stack, Metal forging |
| Residential fire, Explosion, Vehicle fire | |
| 🎆 Recreational Activities | Firework, SkyLantern, Barbecue, Campfire, Torch |
| 🕯️ Civil Controlled Scenes | Waste disposal, Gas stove, Incense burning, Candle, Smoking |
-
Clone the repository:
git clone https://github.com/RISys-Lab/SAFIRE.git cd SAFIRE -
Create a virtual environment:
conda create -n safire python=3.10 conda activate safire
-
Install dependencies: We utilize
vLLMfor efficient MLLM inference.pip install -r requirements.txt
[!NOTE] Different hardware configurations may require specific vLLM installation steps or arguments. Please refer to the official vLLM documentation for detailed instructions tailored to your hardware.
The SAFIRE benchmark consists of multiple datasets hosted on Hugging Face. The evaluation script automatically downloads the requested subset/split.
-
RISys-Lab/SAFIRE_193K_mcvqa (193K Multiple Choice QA)
- Subset:
mcqa - Split:
test
- Subset:
-
RISys-Lab/SAFIRE_83K (83K Images with Captions and Scenario Categories)
- Subset:
data - Split:
train
- Subset:
To evaluate MLLMs on SAFIRE, use the safire/evaluate_mllm.py script. You can run it as a module:
python -m safire.evaluate_mllm --model <model_name> --batch_size <batch_size> --output_dir <output_dir>model_name: The name of the MLLM to evaluate. This should be a Hugging Face model name that can be loaded using vLLM (e.g.,Qwen/Qwen2.5-VL-7B-Instruct).batch_size: The batch size to use for inference (default:128).output_dir: The directory to save the output JSONL file (default:./outputs).- Note: The
safire.evaluate_mllmscript supports all parameters provided byvLLM. You can pass anyvLLMconfiguration arguments (e.g.,--tensor-parallel-size,--gpu-memory-utilization) directly.
If the model cannot fit on a single GPU, please use
--tensor-parallel-sizeto specify the tensor parallelism (see vLLM documentation for more details).
📈 Evaluation Output Format (Click to expand)
The evaluation script generates a JSON file named <model_name>_<timestamp>-results.json in the specified output_dir. This file allows for in-depth analysis of model performance.
Example Output:
{
"model": "Qwen/Qwen2.5-VL-7B-Instruct",
"timestamp": "20251212_153000",
"overall_accuracy": 58.6,
"total_samples": 224000,
"scenario_accuracy": {
"House Fire": {
"accuracy": 61.7,
"correct": ...,
"count": ...
},
...
}
}This output structure facilitates tracking performance metrics across different models and specific scenarios.
To conduct scenario-wise zero-shot inference, use safire/evaluate_mme_zero.py.
python safire/evaluate_mme_zero.py --model <model_name> --output "./result.xlsx" --dataset "RISys-Lab/SAFIRE_IMG" --data_dir "safire_11K" --split "train" To conduct scenario-wise few-shot inference, use safire/evaluate_mme_few.py.
python safire/evaluate_mme_few.py --model <model_name> --output "./result.xlsx" --few_shot_percent 0.03 --min_train_per_class 5 --min_images_per_class 10 --dataset "RISys-Lab/SAFIRE_IMG" --data_dir "safire_11K" --split "train"If you find SAFIRE useful in your research, please consider citing our paper:
@inproceedings{li2026safire,
title={SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs},
author={Li, Pengfei and Suryanto, Naufal and Zhang, Sicheng and Alsharid, Mohammad and Naseer, Muzammal},
booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing},
year={2026}
}