Kai Zhu 「祝凯」

I am a Ph.D. student in Pharmaceutical Science at Zhejiang University, advised by Prof. Hou Tingjun. I am also a visiting student at the Italian Institute of Technology, working with Luigi Bonati and Prof. Michele Parrinello.

My research focuses on developing machine learning-based enhanced sampling methods and their applications to atomic systems.

I enjoy designing new algorithms and making them accessible through open-source software, such as mlcolvar.

Research Interests

Machine-learning collective variables Enhanced sampling Atomistic simulations

Publications

Equal contribution * Corresponding author

Featured work

arXiv SelfTICA framework

Contrastive learning of dynamical representations for enhanced molecular sampling

Kai Zhu†, Jintu Zhang†, Pietro Novelli, Tingjun Hou*, Luigi Bonati*
arXiv arXiv:2606.15495 (2026)

scholar 0
Identifying collective variables that capture slow dynamical modes is essential for sampling rare events in complex systems. Existing machine-learning approaches often require predefined metastable states, carefully chosen descriptors, or training trajectories with high-quality kinetic information. Here, we introduce SelfTICA, a self-supervised contrastive-learning framework that reformulates collective-variable discovery as dynamical representation learning. SelfTICA defines positive and negative pairs from time-lagged molecular configurations, learns reusable features through a contrastive objective linked to spectral variational principles, and extracts orthogonal slow modes by applying time-lagged independent component analysis in the learned representation space. By decoupling representation learning from slow-mode extraction, SelfTICA avoids direct optimization of eigendecomposition-based objectives and enables spectra and collective variables to be evaluated across lag times without retraining. Across different atomistic systems, SelfTICA learns dynamical representations from limited, biased, or exploratory data and converts them into collective variables that accelerate rare-event exploration and improve free-energy convergence.
Nat. Commun. Illustration of AR-NTD enhanced sampling workflow

Targeting the intrinsically disordered AR-NTD through a machine learning-based enhanced sampling workflow

Kai Zhu†, Huating Wang†, Jintu Zhang†, Renling Hu, Linlong Jiang, Hui Zhang, Yu Kang, Tingjun Hou*, Dan Li*
Nat. Commun. 17, 7206 (2026)

scholar 0
Intrinsically disordered proteins are challenging drug targets because they lack stable three-dimensional structures and instead exist as dynamic conformational ensembles. In this work, we developed a machine learning-based enhanced sampling workflow to target the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD), a therapeutic target for drug-resistant prostate cancer. By combining enhanced sampling simulations with machine-learning collective variables, we characterized the conformational landscape of the Tau-5 region, identified druggable conformations, and revealed key molecular features involved in ligand recognition. These insights enabled structure-based virtual screening and led to the identification of K53, an AR-NTD antagonist with anti-proliferative activity in enzalutamide-resistant prostate cancer cells.
Chem. Rev. Cover image for MLES review

Enhanced Sampling in the Age of Machine Learning: Algorithms and Applications

Kai Zhu†, Enrico Trizio†, Jintu Zhang, Renling Hu, Linlong Jiang, Tingjun Hou*, Luigi Bonati*
Chem. Rev. 2026, 126, 1, 671-713

scholar 71
Molecular dynamics simulations hold great promise for providing insight into the microscopic behavior of complex molecular systems. However, their effectiveness is often constrained by long timescales associated with rare events. Enhanced sampling methods have been developed to address these challenges, and recent years have seen a growing integration with machine learning techniques. This Review provides a comprehensive overview of how they are reshaping the field, with a particular focus on the data-driven construction of collective variables. Furthermore, these techniques have also improved biasing schemes and unlocked novel strategies via reinforcement learning and generative approaches. In addition to methodological advances, we highlight applications spanning different areas, such as biomolecular processes, ligand binding, catalytic reactions, and phase transitions. We conclude by outlining future directions aimed at enabling more automated strategies for rareevent sampling.
Nat. Commun. Illustration for LiTEN work

A Scalable and Quantum-Accurate Foundation Model for Biomolecular Force Field via Linearly Tensorized Quadrangle Attention

Qun Su†, Kai Zhu†, Qiaolin Gou†, Jintu Zhang, Renling Hu, Yurong Li, Yongze Wang, Hui Zhang, Ziyi You, Linlong Jiang, Yu Kang, Jike Wang, Chang-Yu Hsieh, Tingjun Hou
Nat. Commun. 17, 3639 (2026).

scholar 2
Accurate atomistic biomolecular simulations are vital for disease mechanism understanding, drug discovery, and biomaterial design, but existing simulation methods exhibit significant limitations. Classical force fields are efficient but lack accuracy for transition states and fine conformational details critical in many chemical and biological processes. Quantum Mechanics (QM) methods are highly accurate but computationally infeasible for large-scale or long-time simulations. AI-based force fields (AIFFs) aim to achieve QM-level accuracy with efficiency but struggle to balance many-body modeling complexity, accuracy, and speed, often constrained by limited training data and insufficient validation for generalizability. To overcome these challenges, we introduce LiTEN, a novel equivariant neural network with Tensorized Quadrangle Attention (TQA). TQA efficiently models three- and four-body interactions with linear complexity by reparameterizing high-order tensor features via vector operations, avoiding costly spherical harmonics. Building on LiTEN, LiTEN-FF is a robust AIFF foundation model, pre-trained on the extensive nablaDFT dataset for broad chemical generalization and fine-tuned on SPICE for accurate solvated system simulations. LiTEN achieves state-of-the-art (SOTA) performance across most evaluation subsets of rMD17, MD22, and Chignolin, outperforming leading models such as MACE, NequIP, and EquiFormer. LiTEN-FF enables the most comprehensive suite of downstream biomolecular modeling tasks to date, including QM-level conformer searches, geometry optimization, and free energy surface construction, while offering 10x faster inference than MACE-OFF for large biomolecules (~1000 atoms). In summary, we present a physically grounded, highly efficient framework that advances complex biomolecular modeling, providing a versatile foundation for drug discovery and related applications.