elib
DLR-Header
DLR-Logo -> http://www.dlr.de
DLR Portal Home | Imprint | Privacy Policy | Contact | Deutsch
Fontsize: [-] Text [+]

BERTRAFFIC: BERT-BASED JOINT SPEAKER ROLE AND SPEAKER CHANGE DETECTION FOR AIR TRAFFIC CONTROL COMMUNICATIONS

Zuluaga-Gomez, Juan Pablo and Sarfjoo, Saeed Seyyed and Prasad, Amrutha and Nigmatulina, Iuliia and Motlicek, Petr and Ondřej, Karel and Ohneiser, Oliver and Helmke, Hartmut (2023) BERTRAFFIC: BERT-BASED JOINT SPEAKER ROLE AND SPEAKER CHANGE DETECTION FOR AIR TRAFFIC CONTROL COMMUNICATIONS. In: 2022 IEEE Spoken Language Technology Workshop, SLT 2022 - Proceedings. The 2022 IEEE Spoken Language Workshop Technology Workshop (SLT 2022), 2023-01-09 - 2023-01-12, Doha, Qatar. doi: 10.1109/SLT54892.2023.10022718. ISBN 979-835039690-4. ISSN 2639-5479.

[img] PDF
441kB

Abstract

Automatic speech recognition (ASR) allows transcribing the communications between air traffic controllers (ATCOs) and aircraft pilots. The transcriptions are used later to extract ATC named entities, e.g., aircraft callsigns. One common challenge is speech activity detection (SAD) and speaker diarization (SD). In the failure condition, two or more segments remain in the same recording, jeopardizing the overall performance. We propose a system that combines SAD and a BERT model to perform speaker change detection and speaker role detection (SRD) by chunking ASR transcripts, i.e., SD with a defined number of speakers together with SRD. The proposed model is evaluated on real-life public ATC databases. Our BERT SD model baseline reaches up to 10% and 20% token-based Jaccard error rate (JER) in public and private ATC databases. We also achieved relative improvements of 32% and 7.7% in JERs and SD error rat

Item URL in elib:https://elib.dlr.de/189419/
Document Type:Conference or Workshop Item (Speech)
Title:BERTRAFFIC: BERT-BASED JOINT SPEAKER ROLE AND SPEAKER CHANGE DETECTION FOR AIR TRAFFIC CONTROL COMMUNICATIONS
Authors:
AuthorsInstitution or Email of AuthorsAuthor's ORCID iDORCID Put Code
Zuluaga-Gomez, Juan PabloIdiap, EPFLUNSPECIFIEDUNSPECIFIED
Sarfjoo, Saeed SeyyedIdiapUNSPECIFIEDUNSPECIFIED
Prasad, AmruthaUNSPECIFIEDUNSPECIFIEDUNSPECIFIED
Nigmatulina, IuliiaIdiapUNSPECIFIEDUNSPECIFIED
Motlicek, PetrUNSPECIFIEDUNSPECIFIEDUNSPECIFIED
Ondřej, KarelBUT, Brno, Czech RepulicUNSPECIFIEDUNSPECIFIED
Ohneiser, OliverUNSPECIFIEDhttps://orcid.org/0000-0002-5411-691XUNSPECIFIED
Helmke, HartmutUNSPECIFIEDhttps://orcid.org/0000-0002-1939-0200UNSPECIFIED
Date:2023
Journal or Publication Title:2022 IEEE Spoken Language Technology Workshop, SLT 2022 - Proceedings
Refereed publication:Yes
Open Access:Yes
Gold Open Access:No
In SCOPUS:Yes
In ISI Web of Science:Yes
DOI:10.1109/SLT54892.2023.10022718
ISSN:2639-5479
ISBN:979-835039690-4
Status:Published
Keywords:Text-based speaker diarization, speaker change detection, speaker role detection, air traffic control communications, chunking
Event Title:The 2022 IEEE Spoken Language Workshop Technology Workshop (SLT 2022)
Event Location:Doha, Qatar
Event Type:international Conference
Event Start Date:9 January 2023
Event End Date:12 January 2023
HGF - Research field:Aeronautics, Space and Transport
HGF - Program:Aeronautics
HGF - Program Themes:Air Transportation and Impact
DLR - Research area:Aeronautics
DLR - Program:L AI - Air Transportation and Impact
DLR - Research theme (Project):L - Integrated Flight Guidance
Location: Braunschweig
Institutes and Institutions:Institute of Flight Guidance > Controller Assistance
Deposited By: Diederich, Kerstin
Deposited On:22 Feb 2023 10:02
Last Modified:24 Apr 2024 20:50

Repository Staff Only: item control page

Browse
Search
Help & Contact
Information
electronic library is running on EPrints 3.3.12
Website and database design: Copyright © German Aerospace Center (DLR). All rights reserved.