elib
DLR-Header
DLR-Logo -> http://www.dlr.de
DLR Portal Home | Impressum | Datenschutz | Barrierefreiheit | Kontakt | English
Schriftgröße: [-] Text [+]

Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems

Thirunavukkarasu, Arunachalam und Helms, Domenik und Bernal, Christian Ojeda und Mantowsky, Sven und Bukhari, Saqib und Bücs, Robert Lajos (2026) Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems. Advances in Artificial Intelligence and Machine Learning, Seiten 1-19. Shimur Publications. doi: 10.54364/AAIML.2026.64332. ISSN 2582-9793. (im Druck)

[img] PDF - Verlagsversion (veröffentlichte Fassung)
159kB

Offizielle URL: https://www.oajaiml.com/archive/model-compression-techniques-for-scalable-deployment-of-generative-ai-in-automotive-systems

Kurzfassung

The deployment of generative AI models within automotive systems is fundamentally constrained by the limited computational resources, memory capacity, and energy budgets of in-vehicle hardware. In a companion article, we examined the scalability challenges, current applications, and evolving E/E architectures that frame this problem. The present paper builds on that foundation by providing an in-depth survey of model compression and optimization techniques that enable the deployment of large neural networks on resource-constrained automotive platforms. We first present an overview of deployment strategies including edge computing, hardware selection, optimized inference frameworks, AI compilers, dedicated accelerators, and knowledge distillation. We then conduct a detailed review of three core compression methodologies: pruning (both with and without retraining), quantization (post-training and quantization-aware approaches for both vision and language models), and low-rank tensor decomposition (including Canonical Polyadic, Tucker, Tensor Train, and Tensor Ring methods, as well as Neural Architecture Search-based compression). The techniques were evaluated by analyzing the underlying principles, reviewing the current available methods, and discussing how the techniques can be used for particular situations in an automotive environment. By using a comparative analysis, trade-offs between compression ratio, speedup for inference, accuracy maintained, and hardware compatibility were all evaluated together so that some conclusions could be drawn. The results of this evaluation provide evidence that multiple compression techniques should be used in combination to achieve the best opportunity for deploying generative AI in real time within an automotive environment. We identify this as a open research direction for the automotive AI community.

elib-URL des Eintrags:https://elib.dlr.de/226346/
Dokumentart:Zeitschriftenbeitrag
Titel:Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems
Autoren:
AutorenInstitution oder E-Mail-AdresseAutoren-ORCID-iDORCID Put Code
Thirunavukkarasu, Arunachalamarunachalam.thirunavukkarasu (at) dlr.dehttps://orcid.org/0000-0003-0824-140XNICHT SPEZIFIZIERT
Helms, Domenikdomenik.helms (at) dlr.dehttps://orcid.org/0000-0001-7326-200XNICHT SPEZIFIZIERT
Bernal, Christian Ojedachristian.ojeda-bernal (at) valeo.comNICHT SPEZIFIZIERTNICHT SPEZIFIZIERT
Mantowsky, Svensven.mantowsky (at) zf.comNICHT SPEZIFIZIERTNICHT SPEZIFIZIERT
Bukhari, Saqibsaqib.bukhari (at) zf.comNICHT SPEZIFIZIERTNICHT SPEZIFIZIERT
Bücs, Robert Lajosrobert.buecs (at) aptiv.comNICHT SPEZIFIZIERTNICHT SPEZIFIZIERT
Datum:17 August 2026
Erschienen in:Advances in Artificial Intelligence and Machine Learning
Referierte Publikation:Ja
Open Access:Ja
Gold Open Access:Nein
In SCOPUS:Ja
In ISI Web of Science:Ja
DOI:10.54364/AAIML.2026.64332
Seitenbereich:Seiten 1-19
Herausgeber:
HerausgeberInstitution und/oder E-Mail-Adresse der HerausgeberHerausgeber-ORCID-iDORCID Put Code
Ralescu, Ance L.anca.ralescu (at) uc.eduNICHT SPEZIFIZIERTNICHT SPEZIFIZIERT
Verlag:Shimur Publications
ISSN:2582-9793
Status:im Druck
Stichwörter:Automotive AI, Generative AI, Knowledge Distillation, Low-Rank Approximation, Model Compression, Neural Network Pruning, Quantization
HGF - Forschungsbereich:Luftfahrt, Raumfahrt und Verkehr
HGF - Programm:Verkehr
HGF - Programmthema:Straßenverkehr
DLR - Schwerpunkt:Verkehr
DLR - Forschungsgebiet:V ST Straßenverkehr
DLR - Teilgebiet (Projekt, Vorhaben):V - Reliable AI
Standort: Oldenburg
Institute & Einrichtungen:Institut für Systems Engineering für zukünftige Mobilität > System Evolution and Operation
Hinterlegt von: Thirunavukkarasu, Arunachalam
Hinterlegt am:28 Aug 2026 10:43
Letzte Änderung:28 Aug 2026 11:36

Nur für Mitarbeiter des Archivs: Kontrollseite des Eintrags

Blättern
Suchen
Hilfe & Kontakt
Informationen
OpenAIRE Validator logo electronic library verwendet EPrints 3.3.12
Gestaltung Webseite und Datenbank: Copyright © Deutsches Zentrum für Luft- und Raumfahrt (DLR). Alle Rechte vorbehalten.