sözaltı news Science
Science
EN AZ

Exploring transcriptomic and genomic latent variable correction approaches in differential expression analysis

nature.com 11.10.2026 02:00 4 views

Differential expression analysis is a central tool for studying the biological processes altered in human diseases via transcriptomic signatures. However, transcriptomic datasets are systematically confounded by latent variables from two distinct sources: unmeasured technical and biological heterogeneity within the expression data, and expression differences driven by population stratification. Correction using expression-based surrogate variables (SVs) and genotype-based principal components (PCs) addresses these sources independently, yet no study has directly evaluated their combined use against either method alone within a differential expression framework.

In this study we hypothesised that simultaneously including both correction layers would produce more biologically valid and reproducible results than either approach alone, and tested this in two independent post-mortem RNA-seq datasets of amyotrophic lateral sclerosis (ALS) cases and controls with paired genotype data. Four nested differential expression models (corrected for PC-only, SV-only, both SV and PC, and neither PCs nor SVs) were evaluated across the KCLBB (96 cases and 52 controls) and ALS Consortium (272 cases and 35 controls) datasets. Models were evaluated on cross-dataset effect size concordance, cross-dataset replicability quantified by the Jaccard Similarity Index, and biological recall against a curated reference set of 66 known ALS genes.

The combined SV + PC framework outperformed simpler models across most metrics. Replicability improved nearly ten-fold compared to the non-corrected model, (Jaccard index: 2.28% to 19.5%), and the combined framework exhibited a statistically significant 2.2% gain over the SV-only model. Biological recall doubled compared to SV correction alone.

Effect size magnitude consistency was preserved, though a reduction in Spearman’s rank correlation indicates PC correction refines magnitude rather than gene rankings. These findings remained generally robust to genotype PC number and sequencing platform adjustment, with more variable performance observed under ancestry restriction and upon extension to an independent tissue. This study found that SVs and genotype PCs address non-redundant sources of confounding, and we recommend their combined use as standard practice in differential expression analysis where paired genotype data are available, with the greatest benefit expected in ancestrally diverse cohorts.

Notably PCs capturing population structure can also be derived directly from RNA-seq data, potentially extending this framework’s applicability to studies lacking paired genotype data. Although this analysis was restricted to ALS datasets, we expect these findings to generalise to other traits. We would like to thank the London Neurodegenerative Diseases Brain Bank and the NYGC ALS Consortium for generating and maintaining the transcriptomic datasets.

All NYGC ALS Consortium activities are supported by the ALS Association (ALSA, 19-SI-459) and the Tow Foundation, as well as The Target ALS Human Postmortem Tissue Core, New York Genome Center for Genomics of Neurodegenerative Disease, Amyotrophic Lateral Sclerosis Association and TOW Foundation. We also acknowledge the relevant funding bodies that support ongoing ALS research, including the NIHR Maudsley Biomedical Research Centre and collaborating international organisations. We thank people with MND and their families for their participation in the initiatives that led to the generation of the data used in this project.

Extract — continue reading at the source.

Read full story