Integrating histological, spatial protein, and transcriptomic information into a biologically grounded representation remains challenging because these modalities are rarely available as fully paired measurements, while existing computational approaches are commonly developed around individual modality pairs. To address this, we present CUBE (Colorectal Universal Representation & Bridge Encoder), a multimodal representation-learning framework that uses hematoxylin and eosin (H&E) histology as a bridge to integrate spatial protein phenotypes with transcriptome-associated information from incompletely paired data. CUBE independently learns representations from H&E-multiplex immunohistochemistry (mIHC) and H&E-pseudo-ST relationships and integrates them through attention-based fusion with biological grounding from mIHC-derived concepts. The H&E-mIHC representation supported competitive spatial protein reconstruction, while the pseudo-ST-supervised representation transferred to experimentally measured Visium HD spatial transcriptomics data and improved further after decoder-only calibration with the encoder frozen. Importantly, the fused representation retained biological information beyond its direct training targets, capturing immune-epithelial spatial organization and an independently measured ECM-receptor interaction transcriptomic program. Together, these findings demonstrate that separately paired spatial modalities can be organized through histology into a biologically structured and testable multimodal representation, providing a proof-of-concept strategy for multimodal tissue learning without requiring fully paired molecular measurements.
Ge, Z., Cai, H.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 3
- Comments 0
