Non-parallel articulatory-to-acoustic conversion using multiview-based time warping

Abstract

In this paper, we propose a novel algorithm called multiview temporal alignment by dependence maximisation in the latent space (TRANSIENCE) for the alignment of time series consisting of sequences of feature vectors with different length and dimensionality of the feature vectors. The proposed algorithm, which is based on the theory of multiview learning, can be seen as an extension of the well-known dynamic time warping (DTW) algorithm but, as mentioned, it allows the sequences to have different dimensionalities. Our algorithm attempts to find an optimal temporal alignment between pairs of nonaligned sequences by first projecting their feature vectors into a common latent space where both views are maximally similar. To do this, powerful, nonlinear deep neural network (DNN) models are employed. Then, the resulting sequences of embedding vectors are aligned using DTW. Finally, the alignment paths obtained in the previous step are applied to the original sequences to align them. In the paper, we explore several variants of the algorithm that mainly differ in the way the DNNs are trained. We evaluated the proposed algorithm on a articulatory-to-acoustic (A2A) synthesis task involving the generation of audible speech from motion data captured from the lips and tongue of healthy speakers using a technique known as permanent magnet articulography (PMA). In this task, our algorithm is applied during the training stage to align pairs of nonaligned speech and PMA recordings that are later used to train DNNs able to synthesis speech from PMA data. Our results show the quality of speech generated in the nonaligned scenario is comparable to that obtained in the parallel scenario.

Metadata

Item Type:	Article
Authors/Creators:	Gonzalez-Lopez, J.A. Gomez-Alanis, A. Pérez-Córdoba, J.L. Green, P.D. https://orcid.org/0000-0001-9103-7287
Copyright, Publisher and Additional Information:	© 2022 The Authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Keywords:	deep learning; multiview learning; dynamic time warping; canonical correlation analysis; silent speech interface; latent embedding
Dates:	Accepted: 21 January 2022 Published (online): 23 January 2022 Published: 23 January 2022
Institution:	The University of Sheffield
Academic Units:	The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield)
Depositing User:	Symplectic Sheffield
Date Deposited:	16 Feb 2022 10:05
Last Modified:	16 Feb 2022 10:05
Status:	Published
Publisher:	MDPI AG
Refereed:	Yes
Identification Number:	10.3390/app12031167
Open Archives Initiative ID (OAI ID):	oai:eprints.whiterose.ac.uk:183667

Download

Published Version

Filename: applsci-12-01167-v2.pdf

Licence: CC-BY 4.0

CLICK TO DOWNLOAD

CORE (COnnecting REpositories)

Non-parallel articulatory-to-acoustic conversion using multiview-based time warping

Abstract

Metadata

Download

Published Version

Export

Statistics