Skip to main navigation Skip to search Skip to main content

Interpretable Forensic Multi-Domain Signal Framework for Speech Stress Analysis Using Residual and Modulation Dynamics

Research output: Contribution to journalArticlepeer-review

Abstract

Speech-based stress analysis is relevant to forensic-oriented speech processing, security screening, and behavioral monitoring, yet its reliability is often limited by speaker variability, recording conditions, and acoustic mismatch. This study proposes an interpretable multi-domain signal processing framework that models stress-related speech variation through excitation dynamics, vocal tract characteristics, and temporal modulation patterns. The framework integrates source–filter decomposition, residual-domain analysis, harmonic structure analysis, modulation spectrum characterization, and prosodic variability into a unified representation. The SUSAS corpus is used as the primary dataset for supervised stress evaluation. RAVDESS and SAVEE are employed only as controlled arousal-related proxy datasets to examine the consistency of stress-related acoustic patterns, rather than as physiological stress ground truth. VoxCeleb is used exclusively for robustness and domain-variability analysis because it lacks stress labels. For probabilistic evidence assessment, Gaussian mixture models are adopted as the more interpretable density estimator, while normalizing flow is included as a flexible performance-oriented comparator for modeling non-Gaussian feature distributions. Evaluation incorporates likelihood ratio analysis, DET curves, EER, ablation studies, and robustness testing. The proposed framework achieves an EER of 5.8% in the primary supervised evaluation, showing competitive performance while preserving physically meaningful interpretation.

Original languageEnglish
Article number56
JournalSignals
Volume7
Issue number3
DOIs
Publication statusPublished - Jun 2026

Keywords

  • forensic-oriented speech processing
  • likelihood ratio analysis
  • modulation spectrum
  • residual signal analysis
  • source–filter model
  • speech stress analysis

Fingerprint

Dive into the research topics of 'Interpretable Forensic Multi-Domain Signal Framework for Speech Stress Analysis Using Residual and Modulation Dynamics'. Together they form a unique fingerprint.

Cite this