PRIMARY SOURCE
Discovery of a phenazine-thiol conjugase from sparse data using genome-informed machine learning.
Observation and source date: 2026-07-27. This issue uses PubMed Date-Publication records; the same-day arXiv submitted-date window returned no relevant records.
Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (phenazine-thiol conjugase), an enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through nonenzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert "small data" typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.
Evidence boundary: 该案例证明稀疏数据下的酶发现路径;不等于所有小数据问题都会有相同可迁移性。
Community pulse · verified public sources
Social Network Radar
2026-07-27—2026-07-27. Items come from public APIs and pass original-link and Beijing-time window checks. They indicate public attention and discussion, not paper quality, clinical effectiveness, or scientific consensus.
BioTrain @biotrain-tv.bsky.social
Understanding your analysis changes how you design experiments. Training doesn’t just help with today’s data — it improves tomorrow’s science. #Science #Research #Data #Bioinformatics #Training #Analysis
Keith Bradnam 📈 @kbradnam@hachyderm.io
Bioinformatics and genomics resources on reddit — ACGT
Yesterday I updated a very old ACGT blog post to provide a more up-to-date view of subreddits that discuss # genomics and # bioinformatics : https://www. acgt.me/blog/2026/7/26/bioinfo rmatics-and-genomics-resources-on-reddit
Fios Genomics @fiosgenomics.bsky.social
Choosing the right #transcriptomic profiling method just got easier! 💡 Our comparison chart helps you quickly identify the best approach for your #research. 👀 View the chart at https://www.fiosgenomics.com/transcriptomic-profiling-methods #bioinformatics