ribaoo Research Daily

Primary papers first · explicit date boundaries · study-design-aware claims

28 July 2026 Bioinformatics | latest visible batch
No. 2026-07-28 · latest-visible observation

LATEST VISIBLE BATCH

Discovery of a phenazine-thiol conjugase from sparse data using genome-informed machine learning.

Observation date: 2026-07-28. The primary-source endpoint did not yield a verifiable new batch in this run; this issue uses the latest visible record from 2026-07-27, not a publication-date claim for the observation date.

Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (phenazine-thiol conjugase), an enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through nonenzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert "small data" typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.

Evidence boundary: 这是稀疏数据下的一个发现案例,不保证对所有小数据任务同样可迁移。

Observation

2026-07-28

Asia/Shanghai

Latest visible source

2026-07-27

not a 28 July publication claim

Selected

3

primary-source and boundary screen

Same-day batch

unverified

a failed endpoint is not a zero-result claim