Thirunavukkarasu, Arunachalam und Helms, Domenik und Bernal, Christian Ojeda und Mantowsky, Sven und Bukhari, Saqib und Bücs, Robert Lajos (2026) Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems. Advances in Artificial Intelligence and Machine Learning, Seiten 1-19. Shimur Publications. doi: 10.54364/AAIML.2026.64332. ISSN 2582-9793. (im Druck)
|
PDF
- Verlagsversion (veröffentlichte Fassung)
159kB |
Offizielle URL: https://www.oajaiml.com/archive/model-compression-techniques-for-scalable-deployment-of-generative-ai-in-automotive-systems
Kurzfassung
The deployment of generative AI models within automotive systems is fundamentally constrained by the limited computational resources, memory capacity, and energy budgets of in-vehicle hardware. In a companion article, we examined the scalability challenges, current applications, and evolving E/E architectures that frame this problem. The present paper builds on that foundation by providing an in-depth survey of model compression and optimization techniques that enable the deployment of large neural networks on resource-constrained automotive platforms. We first present an overview of deployment strategies including edge computing, hardware selection, optimized inference frameworks, AI compilers, dedicated accelerators, and knowledge distillation. We then conduct a detailed review of three core compression methodologies: pruning (both with and without retraining), quantization (post-training and quantization-aware approaches for both vision and language models), and low-rank tensor decomposition (including Canonical Polyadic, Tucker, Tensor Train, and Tensor Ring methods, as well as Neural Architecture Search-based compression). The techniques were evaluated by analyzing the underlying principles, reviewing the current available methods, and discussing how the techniques can be used for particular situations in an automotive environment. By using a comparative analysis, trade-offs between compression ratio, speedup for inference, accuracy maintained, and hardware compatibility were all evaluated together so that some conclusions could be drawn. The results of this evaluation provide evidence that multiple compression techniques should be used in combination to achieve the best opportunity for deploying generative AI in real time within an automotive environment. We identify this as a open research direction for the automotive AI community.
| elib-URL des Eintrags: | https://elib.dlr.de/226346/ | ||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Dokumentart: | Zeitschriftenbeitrag | ||||||||||||||||||||||||||||
| Titel: | Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems | ||||||||||||||||||||||||||||
| Autoren: |
| ||||||||||||||||||||||||||||
| Datum: | 17 August 2026 | ||||||||||||||||||||||||||||
| Erschienen in: | Advances in Artificial Intelligence and Machine Learning | ||||||||||||||||||||||||||||
| Referierte Publikation: | Ja | ||||||||||||||||||||||||||||
| Open Access: | Ja | ||||||||||||||||||||||||||||
| Gold Open Access: | Nein | ||||||||||||||||||||||||||||
| In SCOPUS: | Ja | ||||||||||||||||||||||||||||
| In ISI Web of Science: | Ja | ||||||||||||||||||||||||||||
| DOI: | 10.54364/AAIML.2026.64332 | ||||||||||||||||||||||||||||
| Seitenbereich: | Seiten 1-19 | ||||||||||||||||||||||||||||
| Herausgeber: |
| ||||||||||||||||||||||||||||
| Verlag: | Shimur Publications | ||||||||||||||||||||||||||||
| ISSN: | 2582-9793 | ||||||||||||||||||||||||||||
| Status: | im Druck | ||||||||||||||||||||||||||||
| Stichwörter: | Automotive AI, Generative AI, Knowledge Distillation, Low-Rank Approximation, Model Compression, Neural Network Pruning, Quantization | ||||||||||||||||||||||||||||
| HGF - Forschungsbereich: | Luftfahrt, Raumfahrt und Verkehr | ||||||||||||||||||||||||||||
| HGF - Programm: | Verkehr | ||||||||||||||||||||||||||||
| HGF - Programmthema: | Straßenverkehr | ||||||||||||||||||||||||||||
| DLR - Schwerpunkt: | Verkehr | ||||||||||||||||||||||||||||
| DLR - Forschungsgebiet: | V ST Straßenverkehr | ||||||||||||||||||||||||||||
| DLR - Teilgebiet (Projekt, Vorhaben): | V - Reliable AI | ||||||||||||||||||||||||||||
| Standort: | Oldenburg | ||||||||||||||||||||||||||||
| Institute & Einrichtungen: | Institut für Systems Engineering für zukünftige Mobilität > System Evolution and Operation | ||||||||||||||||||||||||||||
| Hinterlegt von: | Thirunavukkarasu, Arunachalam | ||||||||||||||||||||||||||||
| Hinterlegt am: | 28 Aug 2026 10:43 | ||||||||||||||||||||||||||||
| Letzte Änderung: | 28 Aug 2026 11:36 |
Nur für Mitarbeiter des Archivs: Kontrollseite des Eintrags