From Speech to Underwater Acoustics: A Transfer Learning Framework for Real-Time Passive Diver Detection Using Keyword Spotting Models

Authors

  • Osama Deeb Department of Telecommunications, Higher Institute for Applied Sciences and Technology, Damascus, Syria https://orcid.org/0009-0008-5915-1192
  • Saier Mahmoud Department of Electronic and Mechanical System, Higher Institute for Applied Sciences and Technology, Damascus, Syria
  • Louay Saleh Department of Electronic and Mechanical System, Higher Institute for Applied Sciences and Technology, Damascus, Syria
  • Assef Jafar Department of Informatics, Higher Institute for Applied Sciences and Technology, Damascus, Syria https://orcid.org/0000-0002-7868-8621
  • Oumayma Al Dakkak Department of Telecommunications, Higher Institute for Applied Sciences and Technology, Damascus, Syria https://orcid.org/0000-0002-8842-0979
  • Ibrahim Chouaib Department of Electronic and Mechanical System, Higher Institute for Applied Sciences and Technology, Damascus, Syria

Abstract

Passive acoustic detection of divers faces challenges such as low signal-to-noise ratios (SNRs), data scarcity, and latency in conventional methods. This paper proposes keyword spotting for diver detection (KWS-DD) – a transfer learning framework that repurposes speech-oriented KWS models for data-efficient diver detection. Diver inhalation signatures are treated as acoustic ‘keywords,’ enabling adaptation of the transformer-based HuBERT architecture (pre-trained on speech) to identify quasi-periodic respiratory events in underwater audio. The core innovation of this work lies in adapting the state-of-the-art speech model HuBERT for accurate diver detection via non-speech inhalation acoustics. This approach eliminates the need for accumulation during the respiratory cycle, enabling real-time detection using a minimal amount of domain-specific data (120 inhalation samples). The solution, deployed in diverse marine conditions, achieved 94.4% accuracy and 94.6% F1-score for inhalation sounds. This represents a more than 50% range extension over conventional methods, which proved unreliable at distances greater than 10 meters in low-SNR environments. The framework reduces false alarms caused by boat noise and generalizes to external datasets, validating cross-domain transferability. This work bridges AI-based speech processing and passive sonar signal processing, offering a resource-efficient solution for real-time underwater surveillance.

Downloads

Published

2026-07-22

How to Cite

Deeb, Osama, et al. “From Speech to Underwater Acoustics: A Transfer Learning Framework for Real-Time Passive Diver Detection Using Keyword Spotting Models”. Archives of Acoustics, July 2026, https://wydawnictwo.pan.pl/index.php/aa/article/view/2753.

Issue

Section

Articles

Similar Articles

You may also start an advanced similarity search for this article.