Personalised text-to-speech (TTS) is a key dimension of identity expression, communicative autonomy and long-term acceptance in augmentative and alternative communication (AAC). Yet child voices remain under-represented in commercial and research TTS ecosystems, and this scarcity is amplified by the ethical, acoustic and organisational difficulty of recording high-quality speech from minors. In AAC contexts, the problem is not only one of voice quality, but also of safeguarding: cloud-based synthesis can expose highly sensitive communicative content generated by children in everyday school and home interactions. This paper presents a privacy-preserving recording-to-deployment protocol for creating paediatric AAC voices intended for offline execution on edge devices. The contribution is methodological rather than experimental. It integrates quantified recording targets, practical guidance for speaker selection, safeguarding and quality assurance, a dataset preparation pipeline, and deployment criteria for lightweight neural TTS models that can run locally on resource-constrained hardware. To address reviewer feedback, the paper positions the protocol explicitly within the state of the art on personalised voices, child voice synthesis and edge-capable neural TTS, and it defines a participatory evaluation plan for subsequent validation with AAC users, caregivers and professionals.
Privacy-Preserving Child Voice TTS for On-Device AAC: A Recording-to-Deployment Protocol
Pagliara, Silvio
;Gerazov, Branislav;Zanfardino, Francesco;Tatulli, Ilaria;Mura, Antonello
2026-01-01
Abstract
Personalised text-to-speech (TTS) is a key dimension of identity expression, communicative autonomy and long-term acceptance in augmentative and alternative communication (AAC). Yet child voices remain under-represented in commercial and research TTS ecosystems, and this scarcity is amplified by the ethical, acoustic and organisational difficulty of recording high-quality speech from minors. In AAC contexts, the problem is not only one of voice quality, but also of safeguarding: cloud-based synthesis can expose highly sensitive communicative content generated by children in everyday school and home interactions. This paper presents a privacy-preserving recording-to-deployment protocol for creating paediatric AAC voices intended for offline execution on edge devices. The contribution is methodological rather than experimental. It integrates quantified recording targets, practical guidance for speaker selection, safeguarding and quality assurance, a dataset preparation pipeline, and deployment criteria for lightweight neural TTS models that can run locally on resource-constrained hardware. To address reviewer feedback, the paper positions the protocol explicitly within the state of the art on personalised voices, child voice synthesis and edge-capable neural TTS, and it defines a participatory evaluation plan for subsequent validation with AAC users, caregivers and professionals.I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.



