Trustworthy and Explainable AI
Can a model's explanations be inspected and relied upon when the stakes are high?
We study self-interpretable models and why their explanations are unstable. Our recent work points to redundant input features rather than architecture alone as a main source of inconsistency, and shows that consensus distillation can improve explanation consistency and accuracy within a single model. We also work on causal debiasing for language models and on uncertainty-aware learning for IP geolocation and graph anomaly detection, where assessing predictive uncertainty is important.

Representative work
- TPAMI
Consensus-Driven Distillation for Trustworthy Explanations in Self-Interpretable GNNs
Wenxin Tai, Fan Zhou, Steve Azzolin, Goce Trajcevski, Ting Zhong, Kunpeng Zhang
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
- ICML
Redundancy Undermines the Trustworthiness of Self-Interpretable GNNs
Wenxin Tai, Ting Zhong, Goce Trajcevski, Fan Zhou
International Conference on Machine Learning, 2025
- ACL
Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning
Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, Ting Zhong
Annual Meeting of the Association for Computational Linguistics, 2023



