All-Weather VLM Enhancing Vision Language Models’ Robustness Under Adverse Imaging Conditions

Published at COLM 2026 ↗

1University of Maryland 2Dolby Laboratories Inc. 3Stony Brook University

* Equal contribution · † Equal advising

Demo on Real-World Driving Footage

The base VLM and ours answer the same question about the same captured frame.

Real-world capture

Footage: the public Seeing Through Fog dataset (Bijelic et al., CVPR 2020), recorded in real adverse weather. Questions were annotated manually, with answers verified by a human who watched the full driving sequence. Ours is fine-tuned from Qwen2.5-VL-7B.

The Task

One degraded frame and one question in, one answer out. No restoration, and no hint about what the degradation is.

  1. A night-time driving frame in fog and snow Input A raw frame captured in real fog, snow or darkness
  2. Question “How many buses are visible in the image?”
  3. Answer “1” — the VLM must get it right from the degraded frame itself

Why Not Restore the Image First?

The obvious fix, restore the image first and then ask the VLM, does not work.

  1. The problemVLMs give wrong answers in adverse imaging conditions: rain, fog, snow, low light, etc.
  2. The intuitive fixRun an image-restoration model first, then ask the VLM about the restored image.
  3. What actually happensRestoration hallucinates: it erases or invents the evidence, and the answer is still wrong.

Q: How many giraffes are in the image?

Correct answer: 1 giraffe

Base VLM

Rainy image

A giraffe behind heavy rain streaks.

0 giraffes

Incorrect

Restore, then VLM

Restored image

The restored image: the rain is gone, and so is the giraffe.

0 giraffes

Still wrong

Ours

Same rainy image

The same rainy image with the giraffe.

1 giraffe

Correct

Q: What is the highway speed limit?

Correct answer: 120 km/h

Base VLM

Original foggy image

A 120 km/h speed-limit sign in thick fog.

100 km/h

Incorrect

Restore, then VLM

Restored image

The restored image: the sign now reads 100.

100 km/h

Still wrong

Ours

Same foggy image

The same foggy image of the 120 km/h sign.

120 km/h

Correct

Restoration introduces hallucination: e.g., it removes the giraffe and rewrites the sign. So we skip restoration and teach the VLM to reason through the degradation itself.

Our Approach

Four training stages build degradation awareness into Qwen2.5-VL-7B, ending in a single inference pass.

  1. 01

    Recognize degradation

    Fine-tune the vision side to predict corruption type and severity.

    degradation tokensquestion tokensVLMFog, 5DegradationEncoder"What degradation andseverity are in the image?"
  2. 02

    Condition understanding

    The language model learns to use explicit degradation tokens.

    img tokensdegradationquestion tokensDegradation VLMA craneVisionEncoder"Degraded byfog, severity 5""What animalis in the image?"Two passes: predict degradation, then answer.
  3. 03

    Reason with chain of thought

    One pass: the model reasons about the degradation inside <think>, then answers.

    img tokensquestion tokensDegradation VLM<think>Fog,severity 5</think>A craneVisionEncoder"What animal is in the image?Think step by step about howthe image is degraded, inside<think> tags."
  4. 04

    Direct Preference Optimization

    Clean-image response preferred, degraded-image response rejected.

    cleandegraded"Describethis image"CoT model(stage 03)CoT model(stage 03)Paired preference data✓ preferredA tall white cranewith black markingson grassy ground.✗ rejectedA brown bird withblack markingsnear a grey hill.same prompt, clean vs. degradedDPO trainingDirect Preference OptimizationThe model always seesthe degraded image.
image tokens degradation tokens question tokens output tokens trained in this stage
Chain-of-thought (CoT) prompt Identify the degradation type and rate its severity (1–10) inside <think> tags, then answer. <think>Degradation type: fog, Severity: 5</think> A crane

Results

Gains over the base model, Qwen2.5-VL-7B: robust on real-world footage, with no clean-image penalty.

+9.5 pts

Real-world driving footage

73.0% vs. 63.5%

+3.9 pts

Real camera degradation (R-Bench)

56.1 vs. 52.2

+4.0 pts

Eight simulated degradations

61.0% vs. 57.0% on RealWorldQA

+1.7 pts

Clean images, too

70.7% vs. 69.0% on RealWorldQA

Evaluations on Simulated Degradations

Simulated degradations, but more diverse scenes (MMHal-Bench). Both models see only the degraded input.

Six MMHal-Bench examples with low light, haze, rain, motion blur, snow and fog added. The base model miscounts bicycles, misses a giraffe and misreads a book title; our model names the degradation and its severity, then answers correctly.

Acknowledgement

We owe special thanks to Yuancheng Xu. His sharp insights and generous advice made this project substantially better.

BibTeX

@inproceedings{wang2026allweathervlm,
  title     = {All-Weather {VLM}: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions},
  author    = {Wang, Tianfu and Xie, Mingyang and Cai, Haoming and Xiong, Tianyi and Wang, Xiyao and Fu, Dongdong and Su, Guan-Ming and Cascante-Bonilla, Paola and Metzler, Christopher},
  booktitle = {Conference on Language Modeling (COLM)},
  year      = {2026}
}
Figure Open original ↗