Abstract
An accurate prediction of protein-nucleic acid binding affinity is vital for deciphering genomic processes, yet existing approaches often struggle in reconciling high accuracy with interpretability and computational efficiency. In this study, we introduce commutative algebra prediction (CAP) framework, which couples persistent Stanley–Reisner theory with advanced sequence embedding for predicting protein-nucleic acid binding affinities. CAP encodes proteins through transformer-learned embeddings that retain long-range evolutionary context, and represents DNA and RNA with k-mer algebra embeddings derived from persistent facet ideals, which capture fine-scale nucleotide geometry. We demonstrate that CAP surpasses the SVSBI protein-nucleic acid benchmark and, in a further test, maintains reasonable performance on newly curated protein-RNA and protein-nucleic acid datasets. Leveraging only primary sequences, CAP generalizes to any protein-nucleic acid pair with minimal preprocessing, enabling genome-scale analyses without 3D structural data and promising faster virtual screening for drug discovery and protein engineering.
| Original language | English |
|---|---|
| Article number | 045068 |
| Journal | Machine Learning: Science and Technology |
| Volume | 6 |
| Issue number | 4 |
| DOIs | |
| State | Published - 30 Dec 2025 |
Bibliographical note
Publisher Copyright:© 2025 The Author(s). Published by IOP Publishing Ltd.
Keywords
- facet persistence barcodes
- machine learning
- persistent commutative algebra
- persistent ideals
- protein-nucleic acid binding
ASJC Scopus subject areas
- Software
- Human-Computer Interaction
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'CAP: Commutative algebra prediction of protein-nucleic acid binding affinities'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver