Skip to content

Latest commit

 

History

39 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SAFIRE SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

📄 Paper  |   🌐 Project Page  |   📘 Dataset and Benchmarks

Official repository for the SAFIRE.

Authors: Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer.


📋 Table of Contents


📢 News

  • [21 August 2026] SAFIRE is accepted to EMNLP 2026 as Findings! Project page is now live.

🔥 Overview

SAFIRE (Safety-Aware Fire-smoke Reasoning Evaluation) is the first large-scale benchmark designed to evaluate Multimodal Large Language Models (MLLMs) in safety-critical fire and smoke scenarios.

💡 Why Context Matters

Current benchmarks often fail to distinguish between critical fires (e.g., Residential scenario fires) and benign ones (e.g., campfires), or misinterpret water vapor, haze, cloud as smoke. SAFIRE is designed to test models' specific context-aware reasoning capabilities on fire and smoke.

SAFIRE Overview

SAFIRE addresses this by providing:

  • 83K high-quality images across 20 real-world scenarios.
  • 193K Multiple-Choice VQA (MCVQA) pairs.
  • 10 Reasoning Dimensions:
🔢 Target Counting 🏷️ Classification
🧠 General Reasoning 🔥 Fire/Smoke Intention
😨 Emotional Response 📖 Linguistic Polysemy
👁️ Fire/smoke Attributes 📐 Spatial Correlation
📍 Position Identification 🚫 Human Presence

📂 Dataset Statistics

The dataset is categorized into 5 Groups and 20 Scenarios:

Group Scenarios
🌿 Natural Phenomena Grassland, Volcano, Forest, Meteor
🏭 Industrial Operations Aerospace assets, Flare stack, Metal forging
⚠️ Accident Incidents Residential fire, Explosion, Vehicle fire
🎆 Recreational Activities Firework, SkyLantern, Barbecue, Campfire, Torch
🕯️ Civil Controlled Scenes Waste disposal, Gas stove, Incense burning, Candle, Smoking

🚀 Getting Started

1. Installation

  1. Clone the repository:

    git clone https://github.com/RISys-Lab/SAFIRE.git
    cd SAFIRE
  2. Create a virtual environment:

    conda create -n safire python=3.10
    conda activate safire
  3. Install dependencies: We utilize vLLM for efficient MLLM inference.

    pip install -r requirements.txt

    [!NOTE] Different hardware configurations may require specific vLLM installation steps or arguments. Please refer to the official vLLM documentation for detailed instructions tailored to your hardware.

📊 Dataset and Benchmarks

The SAFIRE benchmark consists of multiple datasets hosted on Hugging Face. The evaluation script automatically downloads the requested subset/split.

Available Datasets

🧪 Evaluation

1. MLLM Evaluation

To evaluate MLLMs on SAFIRE, use the safire/evaluate_mllm.py script. You can run it as a module:

python -m safire.evaluate_mllm --model <model_name> --batch_size <batch_size> --output_dir <output_dir>
  • model_name: The name of the MLLM to evaluate. This should be a Hugging Face model name that can be loaded using vLLM (e.g., Qwen/Qwen2.5-VL-7B-Instruct).
  • batch_size: The batch size to use for inference (default: 128).
  • output_dir: The directory to save the output JSONL file (default: ./outputs).
  • Note: The safire.evaluate_mllm script supports all parameters provided by vLLM. You can pass any vLLM configuration arguments (e.g., --tensor-parallel-size, --gpu-memory-utilization) directly.

If the model cannot fit on a single GPU, please use --tensor-parallel-size to specify the tensor parallelism (see vLLM documentation for more details).

📈 Evaluation Output Format (Click to expand)

The evaluation script generates a JSON file named <model_name>_<timestamp>-results.json in the specified output_dir. This file allows for in-depth analysis of model performance.

Example Output:

{
    "model": "Qwen/Qwen2.5-VL-7B-Instruct",
    "timestamp": "20251212_153000",
    "overall_accuracy": 58.6,
    "total_samples": 224000,
    "scenario_accuracy": {
        "House Fire": {
            "accuracy": 61.7,
            "correct": ...,
            "count": ...
        },
        ...
    }
}

This output structure facilitates tracking performance metrics across different models and specific scenarios.

2. Vision-Language Encoder Evaluation

To conduct scenario-wise zero-shot inference, use safire/evaluate_mme_zero.py.

python safire/evaluate_mme_zero.py --model <model_name> --output "./result.xlsx" --dataset "RISys-Lab/SAFIRE_IMG" --data_dir "safire_11K" --split "train" 

To conduct scenario-wise few-shot inference, use safire/evaluate_mme_few.py.

python safire/evaluate_mme_few.py --model <model_name> --output "./result.xlsx" --few_shot_percent 0.03 --min_train_per_class 5 --min_images_per_class 10 --dataset "RISys-Lab/SAFIRE_IMG" --data_dir "safire_11K" --split "train"

📝 Citation

If you find SAFIRE useful in your research, please consider citing our paper:

@inproceedings{li2026safire,
  title={SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs},
  author={Li, Pengfei and Suryanto, Naufal and Zhang, Sicheng and Alsharid, Mohammad and Naseer, Muzammal},
  booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing},
  year={2026}
}

About

[EMNLP'26-Findings] SAFIRE: A Safety-Critical Benchmark for Fire and Smoke Scene Reasoning in Multimodal LLMs

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages