Six units26 meetingsAug 27 – Dec 3
Schedule
Classes begin August 26 and end December 4. Fall Break, Election Day, and Thanksgiving remove three of our meetings; November 24 is held online.
The instructor reserves the right to adjust readings. The calendar dates are fixed.
Unit I
Foundations: what is being measured?
Emotion theory and its critics, then the corpora everything downstream is trained on.
01Thu Aug 27
Course introduction
Scope of the field, seminar mechanics, how to read and present a technical paper. Presentation sign-ups.
-
Affective Computing
Picard · MIT Media Lab Perceptual Computing TR 321, 1995 · skim
02Tue Sep 1
Can emotion be inferred from behavior?
-
Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements
Barrett, Adolphs, Marsella, Martinez & Pollak · Psychological Science in the Public Interest 20(1):1–68, 2019 · read the executive summary and §1–2 closely, skim the rest
03Thu Sep 3
Corpora I: acted and elicited emotion
-
IEMOCAP: Interactive Emotional Dyadic Motion Capture Database
Busso et al. · Language Resources and Evaluation 42(4):335–359, 2008
-
CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset
Cao et al. · IEEE Transactions on Affective Computing 5(4):377–390, 2014
04Tue Sep 8
Corpora II: naturalistic and in-the-wild data
-
Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech From Existing Podcast Recordings
Lotfian & Busso · IEEE Transactions on Affective Computing 10(4):471–483, 2019
-
AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild
Mollahosseini, Hasani & Mahoor · IEEE Transactions on Affective Computing 10(1):18–31, 2019
Unit II
Learned speech representations
What self-supervised models learn from raw audio, and how much of it is affect.
05Thu Sep 10
Self-supervised speech representation learning
-
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Baevski, Zhou, Mohamed & Auli · NeurIPS 2020
06Tue Sep 15
Masked prediction and full-stack pretraining
-
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Hsu et al. · IEEE/ACM TASLP 29:3451–3460, 2021
-
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Chen et al. · IEEE JSTSP 2022
07Thu Sep 17
What do these representations actually encode?
-
SUPERB: Speech Processing Universal PERformance Benchmark
Yang et al. · Interspeech 2021
08Tue Sep 22
Prompting and parameter-efficient adaptation
-
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, Al-Rfou & Constant · EMNLP 2021
-
Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning
Shi, Zhang & Gao · Interspeech 2025
Unit III
Audio-language models and emotion
General hearing abilities bolted onto language models — and what it takes to make them hear affect.
09Thu Sep 24
Audio-language models I
-
Pengi: An Audio Language Model for Audio Tasks
Deshmukh, Elizalde, Singh & Wang · NeurIPS 2023
-
SALMONN: Towards Generic Hearing Abilities for Large Language Models
Tang et al. · ICLR 2024
10Tue Sep 29
Audio-language models II
-
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Chu et al. · arXiv:2311.07919, 2023
-
Listen, Think, and Understand
Gong, Luo, Liu, Karlinsky & Glass · ICLR 2024
11Thu Oct 1
Making audio LLMs emotion-aware
-
Emotion-Aware Audio Large Language Models with Dual Cross-Attention and Context-Aware Instruction Tuning
Du, Lu, Zhou & Gao · Interspeech 2025
12Tue Oct 6
Structured prompting and low-resource affect
-
Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
Shi, Du, Hong & Gao · IEEE ICASSP 2026
-
Role-Guided Annotation and Prototype-Aligned Representation Learning for Historical Literature Sentiment Classification
Du, Shi, Myerston, Lu, Zhou & Gao · Findings of EMNLP 2025
Thu Oct 8
No class — Fall Break, October 8–11.
13Tue Oct 13
Adapting at test time
-
Tent: Fully Test-Time Adaptation by Entropy Minimization
Wang, Shelhamer, Liu, Olshausen & Darrell · ICLR 2021
-
EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
Shi, Du, Hong & Gao · IEEE ICASSP 2026
Unit IV
Generating emotional speech
Codecs, zero-shot synthesis, style control, and aligning generation with human preference.
14Thu Oct 15
Neural audio codecs
-
SoundStream: An End-to-End Neural Audio Codec
Zeghidour, Luebs, Omran, Skoglund & Tagliasacchi · IEEE/ACM TASLP 2021
-
High Fidelity Neural Audio Compression (EnCodec)
Défossez, Copet, Synnaeve & Adi · arXiv:2210.13438, 2022
15Tue Oct 20
Codecs that preserve affect
-
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
Shi, Du, Song, Hong, Zhang & Gao · Findings of ACL 2026
-
AudioLM: A Language Modeling Approach to Audio Generation
Borsos et al. · IEEE/ACM TASLP 2023
16Thu Oct 22
Zero-shot text-to-speech
-
Neural Codec Language Models Are Zero-Shot Text to Speech Synthesizers (VALL-E)
Wang et al. · arXiv:2301.02111, 2023
-
NaturalSpeech 2: Latent Diffusion Models Are Natural and Zero-Shot Speech and Singing Synthesizers
Shen et al. · ICLR 2024
17Tue Oct 27
Style and expressivity
-
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
Li, Han, Raghavan, Mischler & Mesgarani · NeurIPS 2023
Mon Oct 26
Last day to withdraw.
18Thu Oct 29
Aligning generative models with preferences
-
Direct Preference Optimization: Your Language Model Is Secretly a Reward Model
Rafailov, Sharma, Mitchell, Manning, Ermon & Finn · NeurIPS 2023
-
Diffusion Model Alignment Using Direct Preference Optimization
Wallace et al. · CVPR 2024
Tue Nov 3
No class — Election Day.
19Thu Nov 5
Preference optimization for emotional TTS
-
Emo-BPO: Emotion Bidirectional Preference Optimization for Diffusion-based Emotional TTS
Shi, Du, Song, Hong, Zhang & Gao · Interspeech 2026
-
Emotion-Aligned Generation in Diffusion Text-to-Speech Models via Preference-Guided Optimization
Shi, Du, He, Hong & Gao · IEEE ICASSP 2026
Unit V
Interpretability and robustness
What the features mean, and what happens to them outside the training distribution.
20Tue Nov 10
Opening the black box: interpretable emotion control
-
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Huben, Cunningham, Smith, Ewart & Sharkey · ICLR 2024
-
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
Du, Shi, Lu, Zhou & Gao · ICML 2026
21Thu Nov 12
Domain adaptation
-
Domain-Adversarial Training of Neural Networks
Ganin et al. · JMLR 17(59):1–35, 2016
-
Adversarial Discriminative Domain Adaptation
Tzeng, Hoffman, Saenko & Darrell · CVPR 2017
22Tue Nov 17
Distribution shift in deployed affect systems
-
A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks
Lee, Lee, Lee & Shin · NeurIPS 2018
-
E-ADDA: Unsupervised Adversarial Domain Adaptation Enhanced by a New Mahalanobis Distance Loss
Gao, Baucom, Gordon, Rose, Wang & Stankovic · IEEE SmartComp 2023
23Thu Nov 19
Affect sensing in the wild: health deployments
-
Emotion Recognition Robust to Indoor Environmental Distortions Using Out-of-Distribution Detection
Gao, Salekin, Gordon, Rose, Wang & Stankovic · ACM Transactions on Computing for Healthcare, 2021
-
Gait-Guard: Turn-aware Freezing of Gait Detection for Non-intrusive Intervention Systems
Koltermann, Clapham, Blackwell, Jung, Burnet, Gao, Shao et al. · IEEE/ACM CHASE 2024
24Tue Nov 24
Beyond speech: multimodal and physiological affect Online
November 23–24 are university-designated remote instruction days. This meeting is held online.
-
Multimodal Transformer for Unaligned Multimodal Language Sequences
Tsai, Bai, Liang, Kolter, Morency & Salakhutdinov · ACL 2019
-
Hierarchical Convolution Multibranch Transformer for EEG Signals
Mersa, Jiang, Gao, Li & Zhang · IEEE ICASSP 2026
Thu Nov 26
No class — Thanksgiving Break, November 25–29.
Unit VI
Project presentations
No assigned reading. The papers this week are yours.
25Tue Dec 1
Final project presentations, session I
26Thu Dec 3
Final project presentations, session II
December 4 is the last day of classes. The university exam period runs December 7–11 and 14–15.
Supplementary reading
Optional. Useful for project work and for filling gaps in background.
-
Tensor Fusion Network for Multimodal Sentiment Analysis
Zadeh, Chen, Poria, Cambria & Morency · EMNLP 2017, pp. 1103–1114
-
Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
Bagher Zadeh, Liang, Poria, Cambria & Morency · ACL 2018, pp. 2236–2246
-
Enhancing the Reliability of Out-of-Distribution Image Detection in Neural Networks (ODIN)
Liang, Li & Srikant · ICLR 2018
-
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Kong, Goel, Badlani, Ping, Valle & Catanzaro · ICML 2024
-
MiddleGAN: Generating Inter-domain Images with GANs
Macdonald, Chu, Stankovic, Zhou, Shao & Gao · Transactions on Machine Learning Research, 2024
-
Confidence-Aware Ranker Ensembles for Robust In-Context Knowledge Editing
Nair, Nafee, Jiang, Gao, Chen & Zhang · Findings of ACL 2026
-
ElectroMeter: The Practical Electrolyte Measurement System
Clapham, Zhou, MacDonald, Koltermann, Gao & Shao · IEEE/ACM CHASE 2025
-
Challenges to Recruiting Dementia Caregiving Dyads in Community-Based Settings
Ko, Gao, Wang, Wijayasingha, Wright, Gordon, Wang, Stankovic & Rose · JMIR Formative Research, 2025