ribaoo Research Daily

Primary papers first · explicit date boundaries · study-design-aware claims

27 July 2026 Bioinformatics | 27 July primary-source batch
No. 2026-07-27 · date-scoped primary-source batch

PRIMARY SOURCE

Discovery of a phenazine-thiol conjugase from sparse data using genome-informed machine learning.

Observation and source date: 2026-07-27. This issue uses PubMed Date-Publication records; the same-day arXiv submitted-date window returned no relevant records.

Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (phenazine-thiol conjugase), an enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through nonenzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert "small data" typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.

Evidence boundary: 该案例证明稀疏数据下的酶发现路径;不等于所有小数据问题都会有相同可迁移性。

Observation

2026-07-27

Asia/Shanghai

Routine batch

10

PubMed Date-Publication

Selected

3

primary-source and boundary screen

Same-day arXiv

0

no older batch relabelled