Concept-Based Interpretability with Performance Guarantees: From Vision Models to Large Language Models

PhD Thesis Proposal Defence


Title: "Concept-Based Interpretability with Performance Guarantees: From
Vision Models to Large Language Models"

by

Mr. Andong TAN


Abstract:

Deep neural networks are shown to be powerful in many tasks, yet they are 
often considered as black box models, raising concerns regarding their 
interpretability and trustworthiness when deployed in safety-critical 
scenarios. A rising research direction is concept-based interpretation due to 
their simplicity to be understood by humans. However, existing concept- based 
interpretation frameworks (e.g., part-prototype networks, sparse autoencoders, 
concept bottleneck models) often sacrifice the model's performance when 
constraining the reasoning to be based on only interpretable concepts. To 
address this concern, this thesis presents three approaches that can guarantee 
the model performance from principle while providing a human-centered 
interpretation interface in vision and large language models (LLMs). The three 
approaches achieve this goal via seeking for a precise interpretable model 
parameter decomposition, precise interpretable internal activation 
decomposition and obtaining an equally powerful yet interpretable model, 
respectively. All approaches contribute to converting a black box model into a 
model with interpretable internal reasoning process without performance 
degradation. To further address real-world trustworthiness concerns of 
clinicians regarding the application of LLMs in safety-critical healthcare 
settings, this thesis establishes the world's largest benchmark to evaluate 
guideline adherence capabilities of 8 LLMs across 9 countries/regions and 24 
medical specialties. The benchmark assesses the extent to which LLM's 
reasoning in the output space aligns with authoritative clinical standards. In 
the end, we conclude the thesis by summarizing the main findings and 
discussing the future works in this field.


Date:                   Monday, 20 July 2026

Time:                   3:00pm - 5:00pm

Venue:                  Room 5501
                        Lift 25/26

Committee Members:      Dr. Hao Chen (Supervisor)
                        Dr. Shuai Wang (Chairperson)
                        Dr. Dan Xu