All-Weather VLM
Enhancing Vision Language Models’ Robustness Under Adverse Imaging Conditions
* Equal contribution · † Equal advising
Demo on Real-World Driving Footage
The base VLM and ours answer the same question about the same captured frame.
Footage: the public Seeing Through Fog dataset (Bijelic et al., CVPR 2020), recorded in real adverse weather. Questions were annotated manually, with answers verified by a human who watched the full driving sequence. Ours is fine-tuned from Qwen2.5-VL-7B.
The Task
One degraded frame and one question in, one answer out. No restoration, and no hint about what the degradation is.
-
Input
A raw frame captured in real fog, snow or darkness
- Question “How many buses are visible in the image?”
- Answer “1” — the VLM must get it right from the degraded frame itself
Why Not Restore the Image First?
The obvious fix, restore the image first and then ask the VLM, does not work.
- The problemVLMs give wrong answers in adverse imaging conditions: rain, fog, snow, low light, etc.
- The intuitive fixRun an image-restoration model first, then ask the VLM about the restored image.
- What actually happensRestoration hallucinates: it erases or invents the evidence, and the answer is still wrong.
Q: How many giraffes are in the image?
Correct answer: 1 giraffe
Base VLM
Rainy image
0 giraffes
Incorrect
Restore, then VLM
Restored image
0 giraffes
Still wrong
Ours
Same rainy image
1 giraffe
Correct
Q: What is the highway speed limit?
Correct answer: 120 km/h
Base VLM
Original foggy image
100 km/h
Incorrect
Restore, then VLM
Restored image
100 km/h
Still wrong
Ours
Same foggy image
120 km/h
Correct
Restoration introduces hallucination: e.g., it removes the giraffe and rewrites the sign. So we skip restoration and teach the VLM to reason through the degradation itself.
Our Approach
Four training stages build degradation awareness into Qwen2.5-VL-7B, ending in a single inference pass.
-
01
Recognize degradation
Fine-tune the vision side to predict corruption type and severity.
-
02
Condition understanding
The language model learns to use explicit degradation tokens.
-
03
Reason with chain of thought
One pass: the model reasons about the degradation inside <think>, then answers.
-
04
Direct Preference Optimization
Clean-image response preferred, degraded-image response rejected.
<think> tags, then answer.
<think>Degradation type: fog, Severity: 5</think> A crane
Results
Gains over the base model, Qwen2.5-VL-7B: robust on real-world footage, with no clean-image penalty.
Real-world driving footage
73.0% vs. 63.5%
Real camera degradation (R-Bench)
56.1 vs. 52.2
Eight simulated degradations
61.0% vs. 57.0% on RealWorldQA
Clean images, too
70.7% vs. 69.0% on RealWorldQA
Acknowledgement
We owe special thanks to Yuancheng Xu. His sharp insights and generous advice made this project substantially better.
BibTeX
@inproceedings{wang2026allweathervlm,
title = {All-Weather {VLM}: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions},
author = {Wang, Tianfu and Xie, Mingyang and Cai, Haoming and Xiong, Tianyi and Wang, Xiyao and Fu, Dongdong and Su, Guan-Ming and Cascante-Bonilla, Paola and Metzler, Christopher},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026}
}