Speech & Audio Processing

Speech and audio processing, spoken language understanding

Introduction

The research group focuses on speech and audio signal processing, leveraging advanced machine learning techniques to develop and analyze speech and speaker recognition systems, multilinguality, low-resourced scenarios, speech generation, and voice cloning technologies. Its expertise further encompasses detection of synthetic and manipulated audio (deepfakes), paralinguistic analysis of speech, sign language processing, and pathological speech processing, with applications in human–machine interaction,security, and healthcare. 

Led by Dr. Petr Motlicek and Dr. Mathew Magimai-Doss, the group tackles a range of research problems, with each leader focusing on specific areas of expertise, including speech analytics, voice intelligence, analysis of spoken conversations, paraliguistics in speech, analyses of         pathological speech, sign language processing, etc. 

Alumni

ABROL, Vinayak
ABUTALEBI, Hamid Reza
ADIGA, Aniruddha
AGRAWAL, Sanat
AICHINGER, Ida
AJMERA, Jitendra
ANDERSON, Charles
ANTONELLO, Niccolò
ARADILLA ZAPATA, Guillermo
ASAEI, Afsaneh
ATHINEOS, Marios
BABY, Deepak
BAHAADINI, Sara
BARAN, Eray Abdurrahman
BARBER, David
BAROUDI, Séverin
BAYYA, Yegna Narayana
BENZEGHIBA, Mohamed
BEZZAM, Eric
BHATTACHARJEE, Mrinmoy
BITTAR, Alexandre
BLOYET, Lucas
BORNET, Annie
BORNET, Lisa
BOSTAN, Emrah
BOURDAUD, Nicolas
BOURLARD, Hervé
BOZZO, Vincent
BRAUN, Rudolf
BUSH, Keith
BUTTFIELD, Anna
CAESAR, Holger
CAJEUX, Estelle
CALIZZANO, Rémi
CANDY, Romain
CAROFILIS VASCO, Roberto Andrés
CEREKOVIC, Aleksandra
CEVHER, Volkan
CHAVARRIAGA, Ricardo
CHEN, Haolin
CHENG, Octavian
CHU, Dong
COLLADO, Thierry
COPPIETERS DE GIBSON, Louise
CRITTIN, Frank
DE GRÈVE, Zacharie
DELEZE, Laura
DELEZE, Maxime
DE NEGUERUELA, Cristina
DEWAN, Akshat
DEY, Subhadeep
DIELMANN, Alfred
DIGHE, Pranay
DINES, John
DIONELIS, Nikolaos
DO, Cong-Thanh
DRYGAJLO, Andrzej
DUBAGUNTA, Pavankumar
DUFFNER, Stefan
ELBANNA, Gasser
ESPUÑA FONTCUBERTA, Aleix
FABIEN, Maël
FAJČÍK, Martin
FATEMI, Mitra
FAVRE, Sarah
FERNANDEZ MARTINEZ, Fernando
FERRAS FONT, Marc
FLYNN, Mike
FRITSCH, Julian
GALAN MOLES, Ferran
GANAPATHY, Sriram
GARAU, Giulia
GARCIA-MORAL, Ana Isabel
GARCIA SCHU PEIXOTO, Guilherme
GARIMELLA, Sri Venkata Surya Sivaramakrish
GARIPELLI, Gangadhar
GASIMOV, Huseyn
GAUDARD, Cédric
GEORGOPOULOS, Leonidas
GERAZOV, Branislav
GHOSH, Sucheta
GNINENKO, Nicolas
GOMEZ, Lucas
GOMEZ ALANIS, Alejandro
GRANDVALET, Yves
GRANGIER, David
GRIGNARD, Arnaud
GUENNEC, David
HAGEN, Astrid
HAJIBABAEI, Mahdi
HALPERN, Bence
HANNEMANN, Mirko
HE, Weipeng
HE, Mutian
HEMPTINNE, Coralie
HERMANSKY, Hynek
HIMAWAN, Ivan
HONNET, Pierre-Edouard
HOSSEINI KIVANANI, Nina
IKBAL, Shajith
IMSENG, David
IVANOVA, Maria
JAFARPOUR, Sina
JANBAKHSHI, Parvaneh
JANJAR, Youssef
JEANNINGROS, Loïc
JEROMSON, Garry
KABIL, Selen
KESHET, Joseph
KETABDAR, Hamed
KHODABAKHSHANDEH, Hamid
KHONGLAH, Banriskhem
KHOSRAVANI, Abbas
KIM, Samuel
KNIGHT, James
KNOX, Mary
KOBLER, Andreas
KORCHAGIN, Danil
KRSTULOVIC, Sacha
KULKARNI, Atharva
KUMATANI, Kenichi
LATHOUD, Guillaume
LAURENT, César
LAZARIDIS, Alexandros
LEBRET, Remi
LECORVE, Gwénolé
LEGRAND, Joel
LEW, Eileen Yi Lee
LI, Weifeng
LIANG, Hui
LINKE, Julian
LOUPI, Dimitra
LOVITT, Andrew
LUYET, Gil
MADIKERI, Srikanth
MAGANTI, Hari Krishna
MALEK, Mekki
MARELLI, François
MARIÉTHOZ, Johnny
MARTINS, Renato
MASSON, Olivier
MBANGA NDJOCK, Pierre
MCCOWAN, Iain
MCGREEVY, Michael
MEIER, Corentin
MENDOZA, Viviana
MESOT, Bertrand
MILLÁN, José del R.
MILLIUS, Loris
MISRA, Hemant
MOHAMMADI, Gelareh
MOORE, Darren
MORRIS, Andrew
MOSTAANI, Zohreh
MOULIN, François
MUCKENHIRN, Hannah
MURALIDHAR, Skanda
MUSCAT, Amanda
NA, Xingyu
NADERI, Maryam
NALLANTHIGHAL, Venkata Srikanth
NARAYANAN, Srinivas
NATUREL, Xavier
NJOYIM TCHOUBITH, Peguy
OLIVEIRA PINHEIRO, Pedro Henrique
OUALIL, Youssef
PARIDA, Shantipriya
PARRA GALLEGO, Luis Felipe
PARTHASARATHI, Sree Hari Krishnan
PEREGOUDOV, Artem
PERRIN, Xavier
PERRUCHOUD, Loise
PETER, Naoki
PICART, Benjamin
PINTO, Francisco
PINTO, Joel
PITON, Timothy
POCARD, Valentin
POEL, Mannes
POTARD, Blaise
PRASAD, Amrutha
PRASAD, Ravi
PRONOBIS, Marianna
PUROHIT, Tilak
QUER, Guillem
RAISSI, Tina
RANGAPPA, Pradeep
RASIPURAM, Ramya
RAZAVI, Marzieh
RES, Jakub
RIMOLDI, Emanuele
RUFAI, Amina
RUVOLO, Barbara
SAHA, Atreyee
SAHEER, Lakshmi
SAILOR, Hardik
SALAMIN, Chloé
SAMUI, Suman
SANCHEZ LARA, Alejandra
SANTOS REVILLA, Andrea Elena
SAPRU, Ashtosh
SARFJOO, Saeed
SARKAR, Eklavya
SCARINGELLA, Nicolas
SCHNELL, Bastian
SCHWERY, Theresa
SEBASTIAN, Jilt
SECUJSKI, Milan
SHAHNAWAZUDDIN, Syed
SHAKAS, Alexis
SHANKAR, Ravi
SHARMA, Pulkit
SHARMA, Shivam
SILAGHI, Marius
SINGH, Muskaan
SKOUMAS, Georgios
SOKOLOW, Alexandre
SOLDO, Serena
SOMPURA, Prasanna
SPANO, Kyra
SPUCCHES, Elodie
SRINIVASAMURTHY, Ajay
STEPHENSON, Todd
STERPU, George
SUN, Yang
ŠVIHROVÁ, Radoslava
SZASZAK, György
TAGHIZADEH, Mohammad Javad
TARIGOPULA, Neha
THEUX, Julien
THOMAS, Samuel
THORBECKE, Iuliia
TKACZUK, Jakub
TO, Cuong
TONG, Sibo
TORNAY, Sandrine
TOSIC, Tamara
TRIASTCYN, Aleksei
TRUSCELLO, Léonard
TRUTNEV, Alex
TÜSKE, Zoltan
TYAGI, Hemant
TYAGI, Vivek
ULDRY, Laurent
ULLAL, Vijay Nayak
ULLMANN, Raphael
VALENTE, Fabio
VANDERREYDT, Geoffroy
VAN DOORN, Gerwin
VAN KOMMER, Robert
VASQUEZ-CORREA, Juan Camilo
VELASCO, Jose
VERZAT, Augustin
VIDAL, Maxime
VIJAYASENAN, Deepu
VITEK, Radovan
VLASENKO, Bogdan
VOMSATTEL, Sarah
VYAS, Sargam
VYAS, Apoorv
WANG, Lei
WANG, Yang
WEBER, Katrin
WELLNER, Pierre
YADAV, Sarthak
YELLA, Sree Harsha
ZHAN, Qingran
ZHANG, Alice
ZULUAGA GOMEZ, Juan Pablo

Ongoing projects

AI2PUB

Artificial Intelligence (AI) has become a powerful and pervasive technology in recent years, influencing numerous aspects of our daily lives. We encounter AI through recommendation algorithms in online stores, voice-activated smartphone assistants, or the widespread use of technologies like ChatGPT. However, its rapid growth and integration into society raise complex questions and concerns among the public. Public opinions on AI vary widely; while some people are enthusiastic about its potential to revolutionize industries and enable breakthroughs in fields such as medicine, others fear that AI could lead to undesirable outcomes, such as a loss of human control or privacy. These views are often shaped by media narratives and the competing interests of different stakeholders, which play a significant role in influencing both public opinion and policy decisions.

Our team of scientists and communication experts aims to enhance the understanding of AI technologies among the Swiss people, with a particular focus on teenagers and female students, to create a positive societal impact. Building on the foundations of our previous project, NewsOnAI, we will expand beyond traditional media such as newspapers and employ diverse methods, including artistic performances and interactive exhibitions. We will design these activities to be highly interactive, encouraging active participation and dialogue. Activities will include themed theater plays that explore AI’s impact on everyday life, exhibitions where participants can interact with AI tools, and workshops specifically designed for teenagers and female students to discuss AI’s future role in society. Feedback collected will include real-time audience reactions, structured questionnaires, and focus group discussions, which will be analyzed to continuously refine and adapt our engagement strategies.

Our primary audience includes Swiss citizens interested in cultural activities, particularly teenagers who are keen to follow new trends. Additionally, we are committed to addressing gender aspects by designing content and activities that specifically appeal to female students. We aim to inspire and empower young women to take on more prominent roles in shaping the digital world, acknowledging that they have historically been underrepresented in these fields. As societal attention shifts toward greater inclusion, our project will contribute to fostering a more balanced and equitable digital future.

While many individuals in our target groups may lack in-depth technical knowledge of AI, they often encounter new AI products, companies, and social issues through various media channels, including newspapers and science fiction movies. As a result, they may be aware of recent developments but also susceptible to misunderstandings and controversies related to technologies such as ChatGPT, Elon Musk's brain-chip startup, and other emerging AI applications. It is crucial to recognize that media portrayal significantly influences public opinion on AI, both positively and negatively. Media creators, even if they are not experts in AI, often produce content that captures public attention, which high-profile figures, including entrepreneurs, CEOs, and politicians, may leverage to advance their agendas. This can sometimes lead to skewed public perceptions, whether intentionally or unintentionally. Given this landscape, it is essential for AI scientists to collaborate with media creators, providing evidence-based insights to ensure accurate and balanced information is shared with the public. Our project fosters such collaboration, ensuring that both the potential and limitations of AI are clearly communicated. By sharing our findings through diverse media outlets, we aim to reach a broad audience, extending beyond Switzerland. Furthermore, our proactive engagement efforts will foster dynamic, two-way communication between scientists and the public, using interactive methods in exhibitions and theater plays to engage teenagers and female students specifically. Analyzing the feedback from these initiatives will provide invaluable insights into public perspectives on emerging technologies. This understanding will guide scientists in pursuing research directions that effectively address societal concerns, demonstrating the tangible benefits of our project for both scientific advancement and societal well-being. We anticipate that our efforts will have a multiplying social impact over time, promoting informed public discourse and a deeper understanding of AI technologies.

ELOQUENCE

ELOQUENCE is focused on the research and development of innovative technologies for collaborative voice/chat bots. Voice assistant-powered dialogue engines have previously been deployed in a number of commercial and governmental technological pipelines, with a diverse level of complexity. In our concept, such a complexity can be understood as a problem of analysing unstructured dialogues. ELOQUENCE’s key objective is to better comprehend those unstructured dialogues and translate them into explainable, safe, knowledge-grounded, trustworthy and bias-controlled language models. We envision to develop a technology capable of learning by its own, by adapting from a very data-limited corpora to efficiently support most of the EU languages; from a sustainable computational framework to efficient and green-power architectures and, in essence, that may serve as a guidance for all European citizens whilst being respectful and showing the best of our European values, specifically supporting safety-critical applications by involving humans-in-the-loop.
Overall, ELOQUENCE’s project considers building on top and to improve of prior achievements in the domain of conversational agents, e.g. recently launched and public-domain Large Language Models (LLMs), such as chatGPT (e.g., more recent versions), or LaMDa most of them developed in non-EU countries. While including key industrial enterprises from Europe (i.e., Omilia, Telefonica, Synelixis), ELOQUENCE will validate the developed technology through (i) safety-critical scenarios with human-in-the-loop for security-critical applications (i.e., emergency services in call centres) and (ii) smart home assistants via information retrieval and fact-checking against an online knowledge base for lesser risky autonomous systems (i.e., home-assistants). ELOQUENCE will target the R&D of these novel conversational AI technologies in multilingual and multimodal environments and demonstrated in several pilots.

EUROCONTROL

The ASR system will predict controller commands to reduce the recognition engine’s search space resulting in reduced command recognition error rates and checks the recognition output for plausibility.

EVOICENET

The recent emergence of Artificial Intelligence (AI) methods and audio signal processing techniques open new perspectives on the use of voice to detect or monitor diseases. The vocal biomarker research field is booming but is facing, like other AI-driven fields, a reproducibility and generalisability crisis. To move this promising field of research forward, there is an urgent need to develop a common research framework in Europe, with standardization principles and definitions of good practices and guidelines to help it reach its full potential. A multidisciplinary approach is necessary to tackle this complex task impacting all stakeholders in healthcare virtually.

eVoiceNet aims to establish Europe as a leader in vocal biomarkers, create a collaborative and multidisciplinary network, and facilitate knowledge sharing, standardisation, and the development of privacy-aware solutions, by creating an international network of clinicians, AI experts, speech/voice processing specialists, voice pathologists, privacy/data protection experts, go-to-market specialists, venture capitalists and other providers of financial resources, patients organisations, policymakers as well as industrial partners and other stakeholders (end users, regulators, and other decision-makers), to overcome the major challenges and boost the integration of voice technologies into clinical practice.

Past projects

AI2PUB

Artificial Intelligence (AI) has become a powerful and pervasive technology in recent years, influencing numerous aspects of our daily lives. We encounter AI through recommendation algorithms in online stores, voice-activated smartphone assistants, or the widespread use of technologies like ChatGPT. However, its rapid growth and integration into society raise complex questions and concerns among the public. Public opinions on AI vary widely; while some people are enthusiastic about its potential to revolutionize industries and enable breakthroughs in fields such as medicine, others fear that AI could lead to undesirable outcomes, such as a loss of human control or privacy. These views are often shaped by media narratives and the competing interests of different stakeholders, which play a significant role in influencing both public opinion and policy decisions.

Our team of scientists and communication experts aims to enhance the understanding of AI technologies among the Swiss people, with a particular focus on teenagers and female students, to create a positive societal impact. Building on the foundations of our previous project, NewsOnAI, we will expand beyond traditional media such as newspapers and employ diverse methods, including artistic performances and interactive exhibitions. We will design these activities to be highly interactive, encouraging active participation and dialogue. Activities will include themed theater plays that explore AI’s impact on everyday life, exhibitions where participants can interact with AI tools, and workshops specifically designed for teenagers and female students to discuss AI’s future role in society. Feedback collected will include real-time audience reactions, structured questionnaires, and focus group discussions, which will be analyzed to continuously refine and adapt our engagement strategies.

Our primary audience includes Swiss citizens interested in cultural activities, particularly teenagers who are keen to follow new trends. Additionally, we are committed to addressing gender aspects by designing content and activities that specifically appeal to female students. We aim to inspire and empower young women to take on more prominent roles in shaping the digital world, acknowledging that they have historically been underrepresented in these fields. As societal attention shifts toward greater inclusion, our project will contribute to fostering a more balanced and equitable digital future.

While many individuals in our target groups may lack in-depth technical knowledge of AI, they often encounter new AI products, companies, and social issues through various media channels, including newspapers and science fiction movies. As a result, they may be aware of recent developments but also susceptible to misunderstandings and controversies related to technologies such as ChatGPT, Elon Musk's brain-chip startup, and other emerging AI applications. It is crucial to recognize that media portrayal significantly influences public opinion on AI, both positively and negatively. Media creators, even if they are not experts in AI, often produce content that captures public attention, which high-profile figures, including entrepreneurs, CEOs, and politicians, may leverage to advance their agendas. This can sometimes lead to skewed public perceptions, whether intentionally or unintentionally. Given this landscape, it is essential for AI scientists to collaborate with media creators, providing evidence-based insights to ensure accurate and balanced information is shared with the public. Our project fosters such collaboration, ensuring that both the potential and limitations of AI are clearly communicated. By sharing our findings through diverse media outlets, we aim to reach a broad audience, extending beyond Switzerland. Furthermore, our proactive engagement efforts will foster dynamic, two-way communication between scientists and the public, using interactive methods in exhibitions and theater plays to engage teenagers and female students specifically. Analyzing the feedback from these initiatives will provide invaluable insights into public perspectives on emerging technologies. This understanding will guide scientists in pursuing research directions that effectively address societal concerns, demonstrating the tangible benefits of our project for both scientific advancement and societal well-being. We anticipate that our efforts will have a multiplying social impact over time, promoting informed public discourse and a deeper understanding of AI technologies.

ELOQUENCE

ELOQUENCE is focused on the research and development of innovative technologies for collaborative voice/chat bots. Voice assistant-powered dialogue engines have previously been deployed in a number of commercial and governmental technological pipelines, with a diverse level of complexity. In our concept, such a complexity can be understood as a problem of analysing unstructured dialogues. ELOQUENCE’s key objective is to better comprehend those unstructured dialogues and translate them into explainable, safe, knowledge-grounded, trustworthy and bias-controlled language models. We envision to develop a technology capable of learning by its own, by adapting from a very data-limited corpora to efficiently support most of the EU languages; from a sustainable computational framework to efficient and green-power architectures and, in essence, that may serve as a guidance for all European citizens whilst being respectful and showing the best of our European values, specifically supporting safety-critical applications by involving humans-in-the-loop.
Overall, ELOQUENCE’s project considers building on top and to improve of prior achievements in the domain of conversational agents, e.g. recently launched and public-domain Large Language Models (LLMs), such as chatGPT (e.g., more recent versions), or LaMDa most of them developed in non-EU countries. While including key industrial enterprises from Europe (i.e., Omilia, Telefonica, Synelixis), ELOQUENCE will validate the developed technology through (i) safety-critical scenarios with human-in-the-loop for security-critical applications (i.e., emergency services in call centres) and (ii) smart home assistants via information retrieval and fact-checking against an online knowledge base for lesser risky autonomous systems (i.e., home-assistants). ELOQUENCE will target the R&D of these novel conversational AI technologies in multilingual and multimodal environments and demonstrated in several pilots.

EUROCONTROL

The ASR system will predict controller commands to reduce the recognition engine’s search space resulting in reduced command recognition error rates and checks the recognition output for plausibility.

EVOICENET

The recent emergence of Artificial Intelligence (AI) methods and audio signal processing techniques open new perspectives on the use of voice to detect or monitor diseases. The vocal biomarker research field is booming but is facing, like other AI-driven fields, a reproducibility and generalisability crisis. To move this promising field of research forward, there is an urgent need to develop a common research framework in Europe, with standardization principles and definitions of good practices and guidelines to help it reach its full potential. A multidisciplinary approach is necessary to tackle this complex task impacting all stakeholders in healthcare virtually.

eVoiceNet aims to establish Europe as a leader in vocal biomarkers, create a collaborative and multidisciplinary network, and facilitate knowledge sharing, standardisation, and the development of privacy-aware solutions, by creating an international network of clinicians, AI experts, speech/voice processing specialists, voice pathologists, privacy/data protection experts, go-to-market specialists, venture capitalists and other providers of financial resources, patients organisations, policymakers as well as industrial partners and other stakeholders (end users, regulators, and other decision-makers), to overcome the major challenges and boost the integration of voice technologies into clinical practice.

Latest publications

Reducing prompt sensitivity in LLM-based speech recognition through learnable projection
Burdisso Sergio, Villatoro-Tello Esaú, Kumar Shashi, Madikeri Srikanth, Carofilis Andrés, Rangappa Pradeep, E Manjunath K, Hacioğlu Kadri, Motlicek Petr, Stolcke Andreas
Icassp 2026
2026
Text-only adaptation in LLM-based ASR through text denoising
Burdisso Sergio, Villatoro-Tello Esaú, Carofilis Andrés, Kumar Shashi, Hacioğlu Kadri, Madikeri Srikanth, Rangappa Pradeep, E Manjunath K, Motlicek Petr, Venkatesan Shankar, Stolcke Andreas
ICASSP
2026
Autocrime - open multimodal platform for combating organized crime
Madikeri Srikanth, Motlicek Petr, Sanchez-Cortes Dairazalia, Rangappa Pradeep, Hughes Joshua, Tkaczuk Jacob, Sanchez Lara Alejandra, Khalil Driss, Rohdin Johan, Zhu Dawei, Krishnan Aravind, Klakow Dietrich, Ahmadi Zahra, Kovac Marek, Boboš Dominik, Kalogiros Costas, Alexopoulos Andreas, Marraud Denis
Forensic Science International: Digital Investigation
2025
Efficient data selection for domain adaptation of ASR using pseudo-labels and multi-stage filtering
Rangappa Pradeep, Carofilis Andrés, Prakash Jeena, Kumar Shashi, Burdisso Sergio, Madikeri Srikanth, Villatoro-Tello Esaú, Sharma Bidisha, Motlicek Petr, Hacioğlu Kadri, Venkatesan Shankar, Vyas Saurabh, Stolcke Andreas
Proc. Interspeech
2025