LATEST VISIBLE BATCH
Discovery of a phenazine-thiol conjugase from sparse data using genome-informed machine learning.
Observation date: 2026-07-28. The primary-source endpoint did not yield a verifiable new batch in this run; this issue uses the latest visible record from 2026-07-27, not a publication-date claim for the observation date.
Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (phenazine-thiol conjugase), an enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through nonenzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert "small data" typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.
Evidence boundary: 这是稀疏数据下的一个发现案例,不保证对所有小数据任务同样可迁移。
Community pulse · verified public sources
Social Network Radar
2026-07-28—2026-07-28. Items come from public APIs and pass original-link and Beijing-time window checks. They indicate public attention and discussion, not paper quality, clinical effectiveness, or scientific consensus.
Casey Hand @cyanheads.bsky.social
uniprot-mcp-server: give an agent protein research over UniProtKB. search by GO term or organism, then map any ID across 100+ databases. keyless. github.com/cyanheads/uniprot-mcp-server #MCP #bioinformatics #proteomics
Dr. John White @molecular_geneticist
GeneMap Discovery
Gene sequencing shouldn’t require a degree in cryptography. GeneMap Discovery simplifies tangled data into intuitive maps that everyone can navigate. Try it free for the first week: https:// genemap-discovery.vercel.app # genomics # bioinformatics 🎬 Watch: https://www. youtube.c…
The AWS News Feed @aws-news.com
Kite Advances Scalable Bioinformatics for Cell Therapy Research in Collaboration with AWS HealthOmics
Kite Pharma migrated to AWS HealthOmics to build a scalable, event-driven bioinformatics platform that integrates with their Electronic Laboratory Notebook, enabling 100+ researchers to run genomics workflows without infrastructure expertise.
AWS Blogs on 🦋 @awsblogs.bsky.social
Kite Advances Scalable Bioinformatics for Cell Therapy Research in Collaboration with AWS HealthOmics
📰 New article by Nadeem Bulsara, Ryan Greene, Tanner MacPhee Kite Advances Scalable Bioinformatics for Cell Therapy Research in Collaboration with AWS HealthOmics #AWS #Industries