Signal Processing for Communication

The Signal Processing for Communication group focuses on establishing digital algorithms tackling a wide range of challenges arising in human communication.

Introduction

From the speaker perspective, challenges arise in the presence of speaking impairments. From the listener perspective, challenges arise in the presence of hearing impairments. From the acoustic environment perspective, challenges arise in the presence of undesired interferences such as noise and reverberation. 

With these challenges in mind, the high-level objectives of the group are to establish novel digital signal processing and machine learning algorithms for speech, audio, and multimodal signals to automatically detect speaking and hearing impairments, to provide speaking and hearing assistance, and to improve the communication experience in the presence of undesired interference. 

Alumni

AL-DABBOUSSI, Mustapha
ALTUN, Uğur
BISWAS, Sayantan
DUBOIS, Adrien
GARCIA SCHU PEIXOTO, Guilherme
GE, Chang
HOSSEINI KIVANANI, Nina
KALAAJI, Dana
MALEK, Mekki
MARCEL, Lucie Erine
MILONE, Matteo
SANTOS REVILLA, Andrea Elena
SHEIKH, Shakeel Ahmad
SYLA, Valmir

Ongoing projects

CHASPEEPRO

Oral verbal communication represents the main communication channel among humans. In most communication contexts, speakers must speak clearly and accurately in order to be intelligible. Intelligible speech can be disrupted in a variety of conditions of motor speech disorders (MSD). MSD in adults refers to a broad set of altered speech dimensions (articulation, speech rate, voice, prosody) in the course of several neurological diseases, which can dramatically impact patients’ communication. MSDs are due to disruption in the processes transforming a linguistic message intoarticulated speech, i.e., (a) the retrieval/encoding, contextualization and coordination of speech goals into a speech plan, (b) the preparation of motor programs with detailed neuromuscular specifications, and (c) the execution of these programs. Impairments atthese different stages have been associated with different MSDs, with apraxia of speech(AoS) associated with impairments at the first stage, i.e., the planning stage, and dysarthria associated with impairments at the programming or at the execution stage. Nonetheless, defining planning and programming stages, as well as distinguishing impairments at these two levels in terms of speech features and clinical differential diagnosis, is far from being clear-cut. This proposal builds on the successful outcomes of the Sinergia MoSpeeDi project (2017-2021, https://www.unige.ch/fapse/mospeedi/) led by the same multidisciplinary consortium. Thanks to the complementary expertise in speech and language pathology, psycholinguistics, neurology, phonetics, and speech engineering, we have collected an impressive database of MSD speech, have developed procedures sensitive enough to assess and classify mild and moderate MSD, and have obtained converging experimental evidences for the characterization of processes occurring at the planning and motor programming stages. This knowledge gained from carefully designed experiments and laboratory settings should now be expanded to speech production elicited in a more natural clinical setting. The distinction between speech planning and programming processes should also be further tackled to overcome the difficulty in defining and operationalizing processes at these two stages. With the overarching goal of understanding and modelling speech planning and programming and their related disorders, we will pursue our synergic approach based on the integration of methods and on the convergence of evidence obtained with experimentally induced speech behaviours, electrophysiological brain signals, and acoustic analyses of typical and impaired speech. Based on the results and expertise developed in the ongoing project to pinpoint speech planning and programming and to classify speakers and speech samples, in this project we propose to (a) develop assessment and classification methods applicable to realistic clinical constraints and needs, (b) build on the convergence of phonetic knowledge-based approaches and knowledge-free approaches, (c) enrich our set of acoustic descriptors in order to capture alterations at different scales of speech organization, and (d) complement acoustic-based characterization of speech planning, programming, and MSD classification with EEG signals. The outcomes of the project will rely on substantial data of disordered speech collected from over 180 French speaking participants with different types of MSDs including AoS and subtypes of dysarthria following stroke or neurodegenerative diseases. Results will be used to challenge current models of speech production which need to integrate data from MSD and will contribute to the development of speech assessment systems adapted to atypical speech and to the needs of clinical practice.

EVOICENET

The recent emergence of Artificial Intelligence (AI) methods and audio signal processing techniques open new perspectives on the use of voice to detect or monitor diseases. The vocal biomarker research field is booming but is facing, like other AI-driven fields, a reproducibility and generalisability crisis. To move this promising field of research forward, there is an urgent need to develop a common research framework in Europe, with standardization principles and definitions of good practices and guidelines to help it reach its full potential. A multidisciplinary approach is necessary to tackle this complex task impacting all stakeholders in healthcare virtually.

eVoiceNet aims to establish Europe as a leader in vocal biomarkers, create a collaborative and multidisciplinary network, and facilitate knowledge sharing, standardisation, and the development of privacy-aware solutions, by creating an international network of clinicians, AI experts, speech/voice processing specialists, voice pathologists, privacy/data protection experts, go-to-market specialists, venture capitalists and other providers of financial resources, patients organisations, policymakers as well as industrial partners and other stakeholders (end users, regulators, and other decision-makers), to overcome the major challenges and boost the integration of voice technologies into clinical practice.

PAUSE

Digital communication plays a vital role in our daily lives, occurring in many different forms including mobile communication, meeting through teleconferencing platforms such as Zoom, and using voice-activated agents such as Apple’s Siri or Amazon’s Alexa. Since we live in a noisy world, the recorded microphone signals in these digital communication applications are often contaminated by background noise. Background noise causes signal degradation, thereby impairing speech quality and intelligibility and decreasing the performance of many signal processing techniques required in several applications for a high fidelity communication experience. To deal with background noise, speech enhancement approaches aiming to recover the clean speech signal are indispensable. As such, wide range of speech enhancement approaches have been proposed in past decades, with e.g., the traditional statistical model-based approaches and the more recent deep learning-based approaches having shown strong enhancement performance. Statistical model-based approaches generally rely on a Bayesian framework and assume that the clean speech spectral coefficients follow a Gaussian or super-Gaussian statistical distribution. Deep learning-based approaches generally rely on a large amount of training data to learn a mapping from noisy signals to clean signals. Although these approaches have shown advantageous enhancement performance, they have been devised for scenarios where the target speakers are neurotypical speakers, i.e., speakers that do not exhibit any speech impairments. However, many pathological conditions such as hearing loss, head and neck cancers, or neurological disorders, disrupt the speech production mechanism, resulting in speech impairments across different dimensions. As a result, the statistical distribution of pathological speech differs from that of neurotypical speech and preliminary investigations show that state-of-the-art enhancement approaches can yield a considerably lower performance for pathological signals than for neurotypical signals. Although conditions resulting in pathological speech are widely prevalent, speech enhancement approaches specifically targeting pathological speech, e.g., through using appropriate statistical distributions or training data, have never been established. The PAuSE project aims at developing model-based and deep learning-based speech enhancement approaches that yield an advantageous performance for pathological speech. For model-based approaches, we will derive Minimum Mean Square Error and Maximum A Posteriori estimators exploiting appropriate statistical distributions for pathological speech signals. For deep learning-based approaches, we will target two different research directions to deal with the lack of extensive pathological speech training data. First, we will develop approaches which rely only on neurotypical training signals but exploit knowledge of pathology-specific acoustic features that impact enhancement performance. Such approaches will be based on feature-aware networks, where the feature information is directly embedded in the network or different networks are specialized to enhance signals with different feature profiles. Second, we will develop approaches that also exploit the typically scarcely available pathological training data. Such approaches will be based on pathology-aware networks, augmentation strategies specifically targeting pathological speech, and transfer learning and adaptation strategies. The conducted research will provide a better understanding of the deterioration of enhancement performance for pathological signals. Combining this knowledge with novel estimators and training frameworks will produce enhancement approaches with a strong performance for the growing group of pathological speakers.