INTERSPEECH 2026 · HANYANG UNIVERSITY

Word-level Emotional
Intensity Control in TTS
via Emotion Residual Vectors

Ji-Hyun Park1 · Nam-Seok Song2 · Joon-Hyuk Chang1,2,**

1 Department of Artificial Intelligence, Hanyang University

2 Department of Electronic Engineering, Hanyang University

Seoul, Republic of Korea · ** Corresponding author

THE RESEARCH

Abstract

Modern emotional text-to-speech (TTS) models typically rely on utterance-level conditioning. However, natural emotional speech requires fine-grained variations across words to achieve nuanced, human-like expressiveness. To address this, we introduce emotion residual vectors (ERVs), which capture neutral-to-emotional deviations in word-aligned self-supervised speech embeddings as an annotation-free cue for local prosody. To stabilize intensity control, we compress ERVs into a low-dimensional bottleneck space and train a RoBERTa-based predictor to infer the projected vectors from text and emotion. During inference, the predicted vectors are converted into additive hidden-state offsets and injected into a neutral-conditioned TTS model, enabling continuous word-level intensity control. Experimental results show that our method improves local emotional controllability while preserving naturalness compared to controllable emotional TTS baselines.

Emotional text-to-speech / Emotion intensity / Word-level emotion control

01 / AUDIO DEMONSTRATIONS

Utterance-level intensity

40 SAMPLES

All words share the same intensity value, α. Compare five levels within each emotion, from α = 0 to α = 2.

Open an emotion to explore its samples. Starting a new sample pauses the previous one.
SAMPLE TEXT

“How I hate this foul pool.”

Angry5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
Sad5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
Happy5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
Surprise5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
SAMPLE TEXT

“She is now choosing skirt to wear.”

Angry5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
Sad5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
Happy5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
Surprise5 intensity levels
α = 0
α = 0.5
α = 1
α = 1.5
α = 2
02 / LOCAL CONTROL

One word, a different emphasis.

20 SAMPLES

Listen to how emotional intensity shifts when a different word is targeted. The highlighted word identifies the target in each recording.

SAMPLE TEXT

“Must a name mean something”

Angry5 target words
Must

Must a name mean something

a

Must a name mean something

name

Must a name mean something

mean

Must a name mean something

something

Must a name mean something

Sad5 target words
Must

Must a name mean something

a

Must a name mean something

name

Must a name mean something

mean

Must a name mean something

something

Must a name mean something

Happy5 target words
Must

Must a name mean something

a

Must a name mean something

name

Must a name mean something

mean

Must a name mean something

something

Must a name mean something

Surprise5 target words
Must

Must a name mean something

a

Must a name mean something

name

Must a name mean something

mean

Must a name mean something

something

Must a name mean something

03 / PAPER EXAMPLE

Localized emphasis, side by side.

2 SAMPLES

Compare HED-TTS and our method in the happy condition, with the last two words targeted.

TARGET EMOTION · HAPPY

“How I hate this foul pool.”

BASELINE

HED-TTS

PROPOSED METHOD

Ours · ERV

Audio example discussed in Section 3.4 and Figure 3 of the paper.