William & Mary Fall 2026 Graduate seminar 3 credits

Affective Computing

Systems that recognize, model, generate, and respond to human emotion — read closely, argued over twice a week, and tested against what happens when they leave the lab.

Meetings
Tue & Thu
11:00 a.m. – 12:20 p.m.
Room
John E. Boswell Hall 220
Format
Paper presentation
+ discussion
Instructor
Ashley (Ye) Gao
ygao18@wm.edu

The semester, as a signal

Each frame is one meeting · height = papers assigned · gaps are weeks we don't meet

Open the schedule for all 26 meetings.

  • I · Foundations
  • II · Speech representations
  • III · Audio-language models
  • IV · Emotional speech generation
  • V · Interpretability & robustness
  • VI · Your projects

First meeting: Thursday, August 27.

What this course is

Affective computing is the study of systems that recognize, model, generate, and respond to human emotion. This seminar covers the modern, machine-learning-centered version of that field: how emotion is represented in learned features, how it is recognized from speech and other signals under realistic conditions, how it is controlled in generative models, and how the resulting systems fail once they leave the lab.

We read one to two technical papers per meeting, drawn mostly from ICML, NeurIPS, ICLR, ACL/EMNLP, ICASSP, and Interspeech. The first third of the semester builds shared ground — emotion theory and its critiques, corpora and labels, self-supervised speech representations. The rest moves through four active research areas: emotion-aware audio-language models, expressive and controllable speech synthesis, robustness under distribution shift, and interpretability of affective representations. The last week belongs to your projects.

This is a discussion course, not a lecture course. Come having read the paper.

How a meeting runs

35–40 min

The presenter

One student covers the problem, the method, the experiments, and the weaknesses — then facilitates the discussion that follows. You'll present once or twice over the semester; sign-ups happen on day one.

~30 min

The discussion

Everyone else arrives having read the paper. The instructor presents in Weeks 1–2 to model the format and adds context throughout, but the seminar belongs to the students.

Due 9:00 a.m.

The reading response

250–400 words on every paper you aren't presenting: the central claim, one substantive critique, one question for the room. Credit/no-credit. Your two lowest scores are dropped, no questions asked.

By December you should be able to

  1. Read a technical paper in affective computing critically — identify its claim, its evidence, its baselines, and the gap between what it demonstrates and what it asserts.
  2. Explain the major families of methods for emotion recognition and emotion-controllable generation, and articulate the tradeoffs among them.
  3. Present a research paper to a technical audience and lead a substantive discussion of it.
  4. Situate a specific research contribution within the broader literature.
  5. Design, execute, and report on an original research project, including a defensible experimental protocol.
  6. Reason about the validity and ethical stakes of inferring emotional state from behavioral signals.

Prerequisites

  • Graduate standing in Computer Science, or permission of the instructor.
  • Working knowledge of machine learning at the level of CSCI 516 or equivalent — neural network training, optimization, evaluation.
  • Comfort with Python and at least one deep learning framework. PyTorch preferred.
  • Enough mathematical maturity to follow derivations involving probability, linear algebra, and basic information theory.

No speech background needed

Prior work in speech or signal processing helps, but it isn't required. The audio concepts you need are introduced as they come up — the first time a mel spectrogram or a codec token matters, we stop and explain it.

Who's teaching

Ashley (Ye) Gao

Assistant Professor of Computer Science

Office hours: to be posted before the first day of class.