Large-scale screening and mapping efforts have produced vast libraries of perturbation-readout data. Converting these measurements into mechanistic insights requires models that link perturbation and genetic input to phenotypes, i.e., labels from experimental readouts, which are usually task specific. We propose a different, virtual experiment modeling approach: train generative models to recreate readouts conditioned on the experimental context, and then let established downstream models extract phenotypes from the synthetic data. As an illustrative case, we develop a bidirectional sequence-image generative framework, CELL-FM, that maps protein sequence and cellular context to fluorescence microscopy images and back, enabling in silico localization prediction, image-conditioned functional motif analysis and generation, and large-scale virtual mutagenesis revealing the amino acid features controlling condensate formation of intrinsically disordered peptides. This approach decouples representation learning from task-specific annotation, reuses rich experimental modalities across many downstream tasks, and preserves the spatial and organizational detail that hand-crafted labels often discard.
Zheng, D., Hong, K., Huang, B.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0
