elib
DLR-Header
DLR-Logo -> http://www.dlr.de
DLR Portal Home | Imprint | Privacy Policy | Accessibility | Contact | Deutsch
Fontsize: [-] Text [+]

Text-to-3D Scene Generation: A Transformer-Driven Framework for Custom Virtual Environments

Kothari, Akshat (2025) Text-to-3D Scene Generation: A Transformer-Driven Framework for Custom Virtual Environments. Master's, Brandenburgisch-Technische Universität Cottbus-Senftenberg.

Full text not available from this repository.

Abstract

The rapid evolution of generative AI has transformed 3D content creation, yet current text-to-3D pipelines face a fundamental trade-off between inference speed and creative control. Assetretrieval approaches offer rapid scene assembly but remain constrained by predefined asset libraries, while neural rendering approaches achieve photorealistic quality but suffer from high computational costs and non-editable representations. This thesis bridges this gap by proposing an integrated pipeline that combines LLM-driven spatial planning with a transformer-based mesh generator and a hybrid texturing module. At the core of this system is a modified MeshGPT architecture optimized for fast inference, embedded within an agentic workflow that iteratively validates spatial layouts. The pipeline offers a flexible trade-off between speed and fidelity through a dual-mode texturing module, supporting both rapid UV mapping and diffusion-based synthesis. Experimental evaluation demonstrates that the proposed mesh transformer achieves a 29–39% inference speedup and a 32–44% improvement in perceptual quality compared to the baseline MeshGPT model. Endto-end evaluation confirms the system’s ability to generate valid scenes in 4–12 minutes with over 99% semantic precision. Furthermore, a mobility infrastructure case study validates the system’s practical utility, showing that its modular editing capabilities reduce modification time to 56% of the initial generation cost, thereby facilitating real-time participatory design and rapid prototyping.

Item URL in elib:https://elib.dlr.de/221891/
Document Type:Thesis (Master's)
Title:Text-to-3D Scene Generation: A Transformer-Driven Framework for Custom Virtual Environments
Authors:
AuthorsInstitution or Email of AuthorsAuthor's ORCID iDORCID Put Code
Kothari, Akshatakshat.kothari (at) dlr.deUNSPECIFIEDUNSPECIFIED
DLR Supervisors:
ContributionDLR SupervisorInstitution or E-MailDLR Supervisor's ORCID iD
Thesis advisorWeiss, Danieldaniel.weiss (at) dlr.dehttps://orcid.org/0000-0003-2851-1040
Date:27 November 2025
Open Access:No
Number of Pages:105
Status:Published
Keywords:AI mesh generation
Institution:Brandenburgisch-Technische Universität Cottbus-Senftenberg
Department:Faculty 1 - Institute for Computer Science
HGF - Research field:Aeronautics, Space and Transport
HGF - Program:Transport
HGF - Program Themes:Transport System
DLR - Research area:Transport
DLR - Program:V VS - Verkehrssystem
DLR - Research theme (Project):V - DiVe - Digital organisiertes Verkehrssystem
Location: Berlin-Adlershof
Institutes and Institutions:Institute of Transport Research > Transport Markets and Mobility Services
Deposited By: Galich, Dr. Anton
Deposited On:15 Jan 2026 21:37
Last Modified:15 Jan 2026 21:37

Repository Staff Only: item control page

Browse
Search
Help & Contact
Information
OpenAIRE Validator logo electronic library is running on EPrints 3.3.12
Website and database design: Copyright © German Aerospace Center (DLR). All rights reserved.