Kangjie Zheng Photo

Kangjie Zheng Pronounced /kahng-dʒyeh dʒeng/

Postdoctoral Fellow · Wellcome Sanger Institute
Junior Research Fellow · Wolfson College, Cambridge
AI for Biology · Genomics · Biomolecules

I develop scalable, data-driven AI models that learn from large-scale biological data to decode the language of life and uncover novel fundamental principles of biological systems.

About Me

I am a Postdoctoral Researcher at the Wellcome Sanger Institute in Cambridge, UK, working with Dr. Mo Lotfollahi on developing scalable, generalizable foundation models for large-scale biological data. Since January 2026, I have also been a Junior Research Fellow at Wolfson College, University of Cambridge. My work spans sequence modeling (proteins, genomes) and structural modeling (3D molecular structures), with the goal of advancing integrative understanding across biological modalities. During my Ph.D., I conducted research on foundation models for molecular modeling, which formed the core of my doctoral thesis titled Research on Molecular Modeling Based on Pre-trained Models, and this work was recognized as the Winner of the ACM Beijing Doctoral Dissertation Award.

Academic Background

  • Junior Research Fellow , Wolfson College, University of Cambridge (Jan 2026 – Present)
  • Postdoctoral Fellow, Wellcome Sanger Institute (Sep 2025 – Present)
  • Ph.D. in Computer Science, Peking University (Aug 2020 – Jun 2025)
    Supervisor: Prof. Ming Zhang
  • B.Eng. in Computer Science, Harbin Institute of Technology (Aug 2016 – Jun 2020)
    College: The Honors School of HIT (Top 10 graduates)

Industry Experience

  • Research Intern, AIR in Tsinghua University (Aug 2022 – Nov 2024)
    Research Topic: AI for Protein and Drug Molecules Modeling
  • Research Intern, Tencent AI Lab (Aug 2021 – Aug 2022)
    Research Topic: Non-autoregressive Text Generation
  • Research Intern, Baidu Research (Aug 2019 – May 2020)
    Mentor: Dr. Mingming Sun
    Research Topic: Information Extraction

Research Highlights

Multi-scale foundation models for single-cell gene regulation and unified molecular modeling.

NEW RESEARCH

Single-cell Gene Regulation (MUSS)

A sequence-based, multi-scale model integrating DNA sequence with single-cell chromatin accessibility and gene expression to model gene regulation.

NeurIPS 2026 · Accepted (Poster)

Code will be released soon.

SELECTED RESEARCH

All-Atom Protein Language Model (ESM-AA)

Understands multi-scale molecular data (e.g. drug molecules and proteins), achieving state-of-the-art results on protein–molecule tasks.

ICML 2024

Selected Publications

* equal contribution · See a full list of papers on Google Scholar .

AI for Science

  1. MUSS: A Multi-scale and Sequence-based Model for Single-cell Gene Regulation. NeurIPS 2026 · Accepted (Poster).
    Kangjie Zheng*, Amirhossein Vahidi*, Liying Jin*, Joseph Clarke, Vijaya Baskar MS, Connor Rogerson, Mengji Zhang, Muzlifah Haniffa, Anshul Kundaje, Mohammad Lotfollahi.
    Code will be released soon.
  2. CrossMol: Cross-Modal Mask-Predict Pre-training For 3D Molecular Data. Bioinformatics · 2026.
    Kangjie Zheng, Junwei Yang, Siyu Long, Wei Ju, Wei-Ying Ma, Hao Zhou, Ming Zhang.
  3. ESM All-Atom: Multi-scale Protein Language Model for Unified Molecular Modeling. ICML 2024.
    Kangjie Zheng*, Siyu Long*, Tianyu Lu, Junwei Yang, Xinyu Dai, Ming Zhang, Zaiqing Nie, Wei-Ying Ma, Hao Zhou.
  4. SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision. ICLR 2025.
    Kangjie Zheng, Siyue Liang, Junwei Yang, Bin Feng, Zequn Liu, Wei Ju, Zhiping Xiao, Ming Zhang.
  5. Mol-AE: Auto-Encoder Based Molecular Representation Learning With 3D Cloze Test Objective. ICML 2024.
    Junwei Yang*, Kangjie Zheng*, Siyu Long, Zaiqing Nie, Ming Zhang, Xinyu Dai, Wei-Ying Ma, Hao Zhou.

Language Modeling

  1. ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models. ICML 2025.
    Kangjie Zheng, Junwei Yang, Siyue Liang, Bin Feng, Zequn Liu, Wei Ju, Zhiping Xiao, Ming Zhang.
  2. Towards A Unified Training for Levenshtein Transformer. ICASSP 2023.
    Kangjie Zheng, Longyue Wang, Zhihao Wang, Binqi Chen, Ming Zhang, Zhaopeng Tu.
  3. A Decoding Algorithm Based on Directed Acyclic Transformers for Length-Control Summarization. EMNLP Findings 2024.
    Chenyang Huang, Hao Zhou, Cameron Jen, Kangjie Zheng, Osmar Zaiane, Lili Mou.
  4. Gloss Matters: Unlocking the Potential of Non-Autoregressive Sign Language Translation. ACM MM 2025.
    Zhihao Wang, Shiyu Liu, Zhiwei He, Kangjie Zheng, Liangying Shao, Junfeng Yao, Jinsong Su.

Community Contributions

Academic Service

Program Committee / Reviewer

  • Conference on Neural Information Processing Systems (NeurIPS’25) — Reviewer
  • International Conference on Learning Representations (ICLR’24, ’25) — Reviewer
  • International Conference on Machine Learning (ICML’24, ’25) — Reviewer
  • AAAI Conference on Artificial Intelligence (AAAI’25) — Reviewer
  • ACL Rolling Review (ACL ARR) — Reviewer

Talks & Presentations

Contact Me

Feel free to reach out for collaboration, discussion, or to learn more about me.

Abstract

MUSS: A Multi-scale and Sequence-based Model for Single-cell Gene Regulation

NeurIPS 2026 · Accepted (Poster)

Kangjie Zheng*, Amirhossein Vahidi*, Liying Jin*, Joseph Clarke, Vijaya Baskar MS, Connor Rogerson, Mengji Zhang, Muzlifah Haniffa, Anshul Kundaje, Mohammad Lotfollahi.

* Equal contribution.

Code will be released soon.

Modeling gene regulation in single cells requires jointly representing DNA sequence, chromatin accessibility, and gene expression across regulatory scales. Sequence-based models capture DNA-level regulatory grammar but typically lack cell-level resolution, whereas single-cell omics models capture cell-state semantics but treat genes and peaks as fixed vocabulary tokens, and thus lack a unified sequence-based view of regulation. Here, we present MUSS, a sequence-based multi-scale model for single-cell gene regulation that preserves both DNA-level and cell-level resolution. MUSS derives peak- and gene-level representations from DNA using a frozen pre-trained sequence encoder, integrates them with a multi-scale encoder combining peak-peak self-attention and distance-based peak-gene interaction, and trains on paired scATAC-seq and scRNA-seq to align sequence-derived representations with cell-specific regulatory states. Across gene expression prediction, enhancer-gene linking, TF-target gene recovery, and zero-shot cell clustering, MUSS outperforms both sequence-based and single-cell omics baselines, and further generalizes to genes and peaks unseen during training. Together, these results demonstrate that MUSS learns biologically meaningful and broadly generalizable regulatory representations for decoding cell-type-specific regulatory programs from DNA sequence.