Context-Aware Evaluation of Autonomous Driving in CARLA with Vision-Language Models

Institut
Professur für autonome Fahrzeugsysteme (TUM-ED)
Typ
Masterarbeit /
Inhalt
 
Beschreibung

Motivation


Rule-based metrics are now the standard way to evaluate autonomous driving in both open-loop and closed-loop settings. Well-known examples include collision checks, lane-keeping, time-to-collision, jerk, and progress scores. They are transparent, reproducible, and physically grounded. But they don't understand the scene. A fixed rule can't tell a reckless lane departure apart from a sensible nudge around a parked delivery truck. It also can't tell hard braking caused by poor planning from hard braking for a child running onto the road. In long-tail scenarios, good driving is often situational. Here, rule-based metrics tend to penalize reasonable behavior too much, or they miss unsafe behavior that breaks no explicit rule.

Vision-Language Models (VLMs) provide the missing semantic scene understanding. Recent work such as NVIDIA's DriveJudge (2026) shows that a VLM can read the driving context and decide which rule-based checks apply. This combines contextual reasoning with the spatial precision of deterministic metrics. So far these approaches have mainly been studied on open-loop, real-world datasets. How well they transfer to closed-loop simulation, where scenarios can be controlled and reproduced, is still an open question.

At the AVS Lab, we have developed a metric-based Evaluation Framework in CARLA. It scores driving behavior across Safety, Comfort, and Efficiency. This thesis will extend it with a VLM-based evaluator. The main question is when, and how much, scene understanding improves driving evaluation over pure metrics.

Objectives


The goal of this thesis is to design, implement, and systematically analyze a VLM-augmented scenario evaluator for CARLA. It builds on our existing metric-based evaluation framework.

Planned work packages include:

  • Literature Review: Survey VLM-based driving evaluation paradigms: direct scoring, preference-based critics, and VLM-guided rule invocation (e.g., DriveJudge, DriveCritic). Also review how evaluation metrics themselves are evaluated.
  • Scenario Design: Create a set of CARLA scenarios (e.g., with ScenarioRunner or OpenSCENARIO) where context decides what counts as good driving. Examples: nudging around obstacles, yielding to emergency vehicles, leaving the lane to avoid a hazard, and interactions with vulnerable road users. Each scenario should have both reasonable and unreasonable variants of the driving behavior.
  • VLM Evaluator Implementation: Build a pipeline that turns CARLA recordings (camera views, BEV renderings, trajectories, metric outputs) into VLM inputs. Integrate open-source VLMs (e.g., Qwen-VL) into the existing framework.
  • Evaluation Paradigms: Implement and compare several ways of combining metrics and VLMs:
    (a) direct VLM scoring,
    (b) fusion of metric and VLM scores,
    (c) VLM-guided selection and weighting of existing metric rules.
  • Analysis & Ablations: Compare metric scores with VLM scores against human-annotated ground truth, for example with classification of driving quality (AUC/AP) and pairwise trajectory preference. Run ablations on input modality (front camera, multi-view, BEV, text-only telemetry), prompting strategy, model size, temporal context, and the effect of giving metric outputs to the VLM.
  • Optional Extension: Fine-tune the VLM (SFT and/or RL) on the generated scenario data.

Requirements

  • Strong interest in autonomous driving, simulation, and foundation models / VLMs.
  • Strong Python programming skills. Experience with PyTorch and Hugging Face is a plus.
  • Basic knowledge of machine learning and deep learning.
  • A structured and independent working style, and interest in careful experimental analysis.
  • Proficiency in German or English.
  • Prior experience with CARLA, ScenarioRunner, or VLM prompting/fine-tuning is helpful but not mandatory.

Start Date: Immediate / As soon as possible.

How to Apply: If interested, please submit your application via email, including your CV, current transcript of records, and a brief cover letter explaining why you would like to work on this subject.

Möglicher Beginn
sofort
Kontakt
Christoph Bank
christoph.banktum.de