World Models for 3D Scenario Generation from 2D Data
- Institute
- Lehrstuhl für Robotik, Künstliche Intelligenz und Echtzeitsysteme (TUM-CIT)
- Type
- Master's Thesis
- Content
- experimental
- Description
World models have seen a spike of interest in recent years. Essentially, they are a generali-
zation of the video reconstruction task, allowing for a generalizable prediction of how arbitrary
environments evolve over time. The same world model may be able to predict how a bird flies,
a car driver behaves and a lion jumps within a given time horizon.
With the advent of end-to-end learning in robotics, world models have seen an increased adop-
tion, especially in autonomous driving [1, 2, 3, 4]. Here, two main pillars are the prediction of the
current traffic scenario on board the vehicle, and the efficient simulation of traffic in order to train
an onboard deep learning network. State-of-the-art examples include, e.g., the camera-based
GAIA-3 [5] and GAIA-4 [6] from Wayve. The Waymo World Model [7] is specifically designed
to produce consistent environments across different sensors, i.e., the generated camera and
LiDAR output are congruent.We are especially interested in the synthetic generation of traffic scenarios that can be used to
train E2E models. World models usually have the disadvantage that the learned behavior might
be visually consistent but unrealistic, especially across multiple traffic participants in the same
scene, whereas classic approaches for testing motion planners usually use consistent, oftenti-
mes recorded, traffic data. Our chair has developed the largest and most-used benchmark and
software suite for motion planning, called CommonRoad [8]. However, our data mostly contains
2D trajectory scenarios. Now, we want to use world models to generate the synthetic camera
and LiDAR data associated with these scenarios, so that our consistent scenarios can be used
for vision-based E2E learning whilst still having the underlying 2D real-world data foundation.This thesis is the combination of a literature deep dive and practical work. The student should
first research how world models work conceptually and, subsequently, what promising, state-
of-the-art world models can be used as a starting point for our task. In the second phase, the
student shall develop a world model for generating 3D synthetic environment data from 2D
scenarios.- Requirements
If you are interested in the topic, please send an email to the contact information provided on
the right and attach a short CV, your current grade report, and your Bachelor’s grade report.
Required skills are a good understanding of deep learning, good Python3 coding skills, practical
knowledge of using Ubuntu, and an interest in applying novel learning algorithms and practical
deep learning.- Possible start
- sofort
- Contact
-
Tobias Mascetta, M.Sc.
Room: 5607.03.060 (MPI Building, Floor 3, Part 7)
tobias.mascettatum.de - Announcement
-