Premium accounts now available! Sign up and create a premium account. Read more Close

Advertisement

Image

Adversarial random forests for omics synthesis

Preprint Created on 16 Sep 2026 bioRxiv

Data availability is critical for understanding complex disease pathways and developing robust predictive models. Although high-throughput omics technologies have improved insight into disease mechanisms, data acquisition from inaccessible tissues such as the central nervous system remains a major limitation, causing small sample sizes and complicating early prediction of neurodegenerative disorders such as Alzheimer's and Parkinson's diseases. Generative modeling has emerged as a powerful approach for synthesizing data to support downstream clustering and prediction with small sample size, but existing methods rarely handle high-dimensional tabular omics data effectively. Adversarial random forests (ARFs) provide a well-performing framework for tabular data generation but are not designed for high-dimensional settings. To address this limitation, we introduce high-dimensional ARF (h-ARF), an extension of ARF optimized for integrated clinical and high-dimensional omics data. Using benchmarks across nine datasets and eight performance metrics, we show that h-ARF better preserves both feature distributions, and downstream clustering and prediction utilities compared with ARFs. The method is implemented in the open-source R package harf, available on CRAN.

Kuete Fouodo, C. J., Kapar, J., Huels, A., Liang, D., Wright, M. N.

Advertisement

Stats

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 3
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement