MA6225: Information-Theoretic Methods in Statistical Learning

Course code: MA6225

Official course title: Topics in Machine Learning

Location: S17-05-11

Time: 10.00am to ~12:30pm

Intended Audience: NUS graduate students (though advanced undergraduates are very welcome)

Instructor: Vincent Y. F. Tan (vtan@nus.edu.sg)

Assessment: Class Participation (25%), Quiz 1 (25%), Quiz 2 (25%), Project (25%)

References

There is no single required textbook. The course will draw from the following references:

  • Y. Polyanskiy and Y. Wu, Information Theory: From Coding to Learning, Cambridge University Press, 2025.

  • A. El Gamal and M. Raginsky, Information-theoretic Limits of Learning and Estimation, arXiv:2605.06710, 2026.

  • Selected research papers from the information theory, statistics, and machine learning literature.

Course Description

This graduate-level, proof-oriented course builds a principled, information-theoretic foundation for modern statistical learning. Starting with entropy, relative entropy, and mutual information, the course develops advanced tools—f-divergences, metric entropy, and strong data-processing inequalities—and highlights structural properties such as convexity, tensorization, and data processing that enable dimension-agnostic reasoning.

Applications connect these ideas to statistical decision theory, large-sample asymptotics, mutual-information-based generalization bounds, and fundamental lower bounds via hypothesis-testing reductions, including Le Cam's, Fano's, and Assouad's methods. Entropic techniques for estimation and the role of strong data processing in dependence, high-dimensional inference, and graph problems such as broadcasting and coloring are also treated.

The course is intended for graduate students and advanced undergraduates in mathematics, statistics, engineering, computer science, and analytics. A solid background in probability, including measure-theoretic intuition, is required. Prior exposure to convex optimization, information theory, and statistical learning is strongly recommended.

Tentative Schedule

Week Date Topic Reference
1 13/08/2026 Entropy, KL divergence, and data processing Chs. 1–2
2 20/08/2026 Mutual information; Variational principles Chs. 3–4
3 27/08/2026 Capacity, information radius, convexity, Sinkhorn Ch. 5
4 03/09/2026 Tensorization and f-divergences Chs. 6–7
5 10/09/2026 1-hour Quiz; Entropy methods in combinatorics Ch. 8
6 17/09/2026 Information-theoretic generalization (mutual information, Gibbs algorithm, PAC-Bayes) El Gamal & Raginsky (2026)
24/09/2026 Recess Week
7 01/10/2026 Student presentations
8 08/10/2026 Metric entropy Ch. 27
9 15/10/2026 Statistical decision theory and minimax estimation Ch. 28
10 22/10/2026 Information-theoretic lower bounds Chs. 29, 31
11 29/10/2026 1-hour Quiz; Mutual information methods in statistical learning Ch. 30
12 05/11/2026 Entropic methods for estimation and strong data-processing inequalities Chs. 32–33
13 12/11/2026 Student presentations

Schedule is tentative and may be adjusted depending on the pace of the class and students' interests.

Learning Outcomes

By the end of the course, students should be able to:

  • Explain and manipulate core information measures, including entropy, KL divergence, and mutual information, and derive key identities.

  • Apply f-divergences, metric entropy, and tensorization principles to quantify statistical complexity.

  • Use strong data-processing inequalities to reason about dependence and information propagation in high-dimensional and graph-structured problems.

  • Derive and interpret minimax lower bounds using hypothesis-testing reductions such as Le Cam's method, Fano's inequality, and Assouad's lemma.

  • Employ information-theoretic techniques to obtain generalization bounds and analyze estimation procedures.

  • Formulate and prove rigorous results connecting information measures to asymptotics and statistical decision problems.

  • Read and critique contemporary research at the interface of information theory and statistical learning.

Prerequisites

  • Solid probability background, with measure-theoretic intuition

  • Prior exposure to convex optimization, information theory, and statistical learning is strongly recommended