Identification of the molecular subgroups in Alzheimer's disease by transcriptomic data
Student14:44CCAI
paperi.ai
0:00 / 0:00
He Li, Meiqi Wei, Tianyuan Ye, Yiduan Liu, Dongmei Qi, Xiaorui Cheng
What if Alzheimer’s disease is not one molecular disease at all? By reading gene-expression patterns in hundreds of brain samples, this study separated Alzheimer’s into three subgroups with sharply different biological and clinical profiles.
Background: Alzheimer’s disease (AD) is a heterogeneous pathological disease with genetic background accompanied by aging. This inconsistency is present among molecular subtypes, which has led to diagnostic ambiguity and failure in drug development. We precisely distinguished patients of AD at the transcriptome level. Methods: We collected 1,240 AD brain tissue samples collected from the GEO dataset. Consensus clustering was used to identify molecular subtypes, and the clinical characteristics were focused on. To reveal transcriptome differences among subgroups, we certificated specific upregulated genes and annotated the biological function. According to RANK METRIC SCORE in GSEA, TOP10 was defined as the hub gene. In addition, the systematic correlation between the hub gene and “A/T/N” was analyzed. Finally, we used external data sets to verify the diagnostic value of hub genes. Results: We identified three molecular subtypes of AD from 743 AD samples, among which subtypes I and III had high-risk factors, and subtype II had protective factors. All three subgroups had higher neuritis plaque density, and subgroups I and III had higher clinical dementia scores and neurofibrillary tangles than subgroup II. Our results confirmed a positive association between neurofibrillary tangles and dementia, but not neuritis plaques. Subgroup I genes clustered in viral infection, hypoxia injury, and angiogenesis. Subgroup II showed heterogeneity in synaptic pathology, and we found several essential beneficial synaptic proteins. Due to presenilin one amplification, Subgroup III was a risk subgroup suspected of familial AD, involving abnormal neurogenic signals, glial cell differentiation, and proliferation. Among the three subgroups, the highest combined diagnostic value of the hub genes were 0.95, 0.92, and 0.83, respectively, indicating that the hub genes had sound typing and diagnostic ability. Conclusion: The transcriptome classification of AD cases played out the pathological heterogeneity of different subgroups. It throws daylight on the personalized diagnosis and treatment of AD.
Transcript
What if Alzheimer’s disease is not one molecular disease at all? By reading gene-expression patterns in hundreds of brain samples, this study separated Alzheimer’s into three subgroups with sharply different biological and clinical profiles. Alzheimer’s disease is a heterogeneous pathological disease with genetic background accompanied by aging.
That inconsistency is present among molecular subtypes and has led to diagnostic ambiguity and failure in drug development. The study therefore distinguished patients with Alzheimer’s disease at the transcriptome level, using gene-expression information to look beyond a single disease label.
Identifying the risk population and diagnosing the pathological course of Alzheimer’s disease in the preclinical stage is a major challenge. The Nincds-Adrad core symptom scale has an accuracy of sixty-five to ninety-six percent, but its specificity for distinguishing Alzheimer’s disease from other dementias is twenty-three to eighty-eight percent.
The A slash T slash N system describes amyloid deposition, tau protein disease, and neurodegeneration. But additional biomarkers may be needed to categorize participants before death into Alzheimer’s disease versus other dementias. Between ten and thirty percent of individuals clinically diagnosed with Alzheimer’s disease showed no Alzheimer’s-like neuropathological changes at autopsy.
Existing diagnostic techniques therefore cannot achieve accurate individualized medicine at the genetic level. High-throughput genome sequencing enables rapid analysis of genomic polymorphism in thousands of subjects. In transcriptomics, gene network analysis can identify co-expressed genes and significantly differentially expressed genes.
The study aimed to find more precise biomarkers for early diagnosis based on transcriptomics. It divided Alzheimer’s disease cases into three subgroups according to gene-expression profiles, then identified core genes for each subset. Three Alzheimer’s disease datasets—GSE1297, GSE29378, and GSE84422—were acquired from the Gene Expression Omnibus database.
GSE84422 included three subsets named GSE8442201, GSE8442202, and GSE8442203. Within GSE84422, the analysis used the NO-AD group and the identified Alzheimer’s disease group. Intersection genes from five datasets, including the three GSE84422 subsets, were obtained and merged.
The expression values were log two transformed before cross-platform normalization. ComBat was chosen to normalize the expression values after they were log two transformed, as part of the study’s cross-platform normalization process.
The analysis included one thousand two hundred forty brain samples: seven hundred forty-three Alzheimer’s disease samples and four hundred ninety-seven NO-AD samples from three independent studies. The studies used Affymetrix Human Genome U133A, Affymetrix Human Genome U133B, Illumina HumanHT-12 V3.0, and Affymetrix Human Genome U133 Plus 2.0 arrays.
Clinical features included CDR, Braak, NFT, CERAD, NDP, pH, age, and gender. Samples were distinguished as Alzheimer’s disease or NO-AD according to the dataset definitions of disease states. Clinical differences in CDR, Braak, NFT, CERAD, NPD, and pH were tested with pairwise Wilcoxon’s rank-sum tests.
For differential expression, the thresholds were a Benjamini–Hochberg adjusted P value below zero point zero five and an absolute difference of means greater than zero point two. Consensus clustering was used to classify the Alzheimer’s disease samples into different subgroups, with the maximum cluster number set to ten.
The maximum cluster number was set to ten, and a cluster consensus score greater than zero point eight was used for filtering the adjustment. Figure two uses principal component analysis to show how the gene-expression samples relate before and after batch correction.
In panel A, the five color-coded datasets form clearly separated clusters, consistent with differences caused by platforms and batches. After the authors apply ComBat, panel B shows the datasets intermingled across the principal-component space, supporting removal of the batch effect before clustering the seven hundred forty-three Alzheimer’s disease samples into molecular subgroups.
Consensus clustering classified the gene-expression profiles of seven hundred forty-three Alzheimer’s disease samples after batch-effect removal. The analysis tested between two and nine subgroups. The three-subgroup classification was robust, with each subgroup score higher than zero point eight.
Although the two-subgroup scores were also higher than zero point eight, the three-subgroup classification was selected for subsequent analysis. The three subgroups contained two hundred twenty-three, three hundred ninety-one, and one hundred twenty-nine samples in subgroups I, II, and III.
Expression patterns were highly similar within each subgroup and significantly different between subgroups. Figure four tests how consistently the Alzheimer’s samples separate into molecular subgroups. In panel A, the authors compare consensus scores for cluster counts from two through ten; the three-subgroup solution is identified as robust, with each subgroup’s score above zero point eight.
Panel B visualizes that choice as a consensus matrix for three clusters, where the darker blocks indicate samples repeatedly assigned together across clustering runs. The three subgroups had significantly increased CDR, CERAD, Braak, NFT, NPD, and age in the Alzheimer’s disease group, with P values below zero point zero zero one or below zero point zero one.
The proportion of women was significantly higher in subgroups I and II than in the NO-AD group, while subgroup III had no difference from the NO-AD group. The pH of all three subgroups was lower than the NO-AD group, with a P value below zero point zero zero one.
Figure five compares clinical characteristics across the normal-dementia group and three subgroups, using boxplots for disease and demographic measures and significance brackets from Wilcoxon rank-sum tests. The authors report that CDR, CERAD, Braak, NFT, NPD, and age differ significantly in subgroup comparisons, while pH is lower in all three subgroups than in the ND group.
Women are more common in subgroups one and two than in ND, whereas subgroup three shows no significant difference, helping distinguish the subgroups clinically. Subgroups I and III had a higher risk of cognitive decline than subgroup II.
The three subgroups had almost the same severe amyloid plaque load, but clinical dementia and NFT were higher in subgroups I and III than in subgroup II. The major clinical contrast was between those two higher-risk groups and subgroup II, which had a lower risk of cognitive decline.
The trends of the three-course labels were not completely parallel. In this research, NFT and clinical dementia score were correlated, with a correlation of zero point thirty-nine. Amyloid plaque deposition did not correlate with clinical dementia scores.
The text therefore distinguishes the relationship of NFT with dementia from the relationship of amyloid plaque deposition with dementia. Tau protein, a marker of neuronal injury, combined with mental symptoms, can be used as a standard for the severity classification of Alzheimer’s disease.
Figure seven presents three GSEA enrichment plots for genes upregulated in Alzheimer’s disease subgroups compared with the ND group. Panel A covers subgroup I with one hundred forty-nine genes, panel B subgroup II with four hundred three, and panel C subgroup III with four hundred ninety-one; all are reported at an FDR below zero point zero zero one.
The enrichment curves and gene-hit positions show how each subgroup’s upregulated genes align across the ranked dataset, providing the basis for the subgroup-specific biological-process and pathway interpretations that follow. Gene Ontology enrichment analysis illustrates gene function at the biological-process level.
In subgroup I, the main processes included nuclear-transcribed messenger RNA catabolic process, nonsense-mediated decay, and signal-recognition-particle-dependent cotranslational protein targeting to membrane. The analysis identified three molecular subtypes of Alzheimer’s disease from seven hundred forty-three Alzheimer’s disease samples.
Subtypes I and III had high-risk factors, while subtype II had protective factors. All three subgroups had higher neuritic plaque density, but subgroups I and III had higher clinical dementia scores and neurofibrillary tangles than subgroup II.
Neurofibrillary tangles were positively associated with dementia, but neuritic plaques were not. Subgroup I genes clustered in viral infection, hypoxia injury, and angiogenesis. Subgroup II showed heterogeneity in synaptic pathology and several essential beneficial synaptic proteins.
Subgroup III was a risk subgroup suspected of familial Alzheimer’s disease because of presenilin one amplification, involving abnormal neurogenic signals, glial-cell differentiation, and proliferation. The top ten genes of the groups were selected for follow-up studies according to core enrichment of Gene Set Enrichment Analysis.
Figure nine showed hub-gene expression in the NO-AD group and Alzheimer’s disease subgroups. Gene correlation heat maps showed gene-to-gene interactions in each subgroup. Clinical-relevance heat maps were mapped for single genes and combination genes in each subgroup.
Figure ten links the molecular subgroups to both gene interactions and clinical features. Panels A through C map subgroup-specific gene-to-gene correlations, while panel D shows how single-gene expression varies across clinical categories.
Panel E reports relationships between each subgroup and traits including CDR, Braak, NFT, CERAD, and NPD, and panels F through H extend the analysis to hub genes and CDR, NFT, and NPD. This matters because the authors use these associations to argue that the molecular subtypes capture distinct functional impairments in Alzheimer’s disease.
GSE5281 was used to verify diagnostic value. Receiver operating characteristic curve analysis investigated hub genes and gene unions for differentiating Alzheimer’s disease and NO-AD patients, with area under the curve quantified using the pROC and glmnet packages.
An area under the curve above zero point nine was regarded as outstanding specificity and sensitivity. An area under the curve from zero point seven to zero point nine was regarded as specificity and sensitivity. A new dataset, GSE5281, was used to verify the diagnostic values of the top ten genes in each subgroup.
Single-gene and combination-gene diagnostic values were tested, with three thousand sixty-nine options offered. The combination of eight genes was the optimal molecular-marker diagnostic scheme. Its area under the curve was zero point nine five zero in subgroup I, zero point nine one six in subgroup II, and zero point eight three four in subgroup III.
Figure eleven validates the diagnostic signal using eight-gene combinations across three subgroups. The ROC curves show areas under the curve of zero point nine four nine seven, zero point nine one five eight, and zero point eight three three six for subgroups one, two, and three, respectively.
Panel D further shows expression distributions for hub genes in normal and Alzheimer’s disease groups, with significance marked by Wilcoxon tests, supporting the authors’ proposed multi-gene diagnostic model. The majority of hub genes—twenty-nine out of thirty—were different from the normal group.
The findings indicated that a combination of eight marker genes could predict Alzheimer’s disease. The combination of eight marker genes provides a valuable model for the development of diagnostic chips. Repeatability of subtypes is an important index for detecting and evaluating effectiveness.
Experiments without out-of-sample validation tended to report near-perfect areas under the curve, while out-of-sample cross-validation produced milder and more convincing results. The analysis took the top ten genes from Gene Set Enrichment Analysis in each subgroup as core genes and introduced an additional independent group sample for model evaluation.
Subtypes in animal-model testing are tough to carry out. Like any single-omics approach, data-driven biomarkers do not directly consider the multi-gene and multi-factor mechanisms that influence Alzheimer’s disease. Current phenotypic and omics studies focus on association analysis of genome and disease rather than causal relationship exploration.
Exploration after intervention is required. The study’s central message is that transcriptomic subgroups expose biological differences hidden by a simple Alzheimer’s-versus-normal comparison, supporting more individualized diagnosis while still requiring causal and experimental validation.
A derivative work by Paperi · AI-generated script, voice and captions
· pages and figures unaltered
Made with Paperi.
Drop in a research PDF — get a narrated video walkthrough like this one,
with highlights that follow the narration. Free to start.
Could carrying more body fat sometimes protect bones—but only up to a point, and differently for women and men? This study finds that the answer is not a simple “more” or “less.”Can body fat be both protective and harmful to bone? This study finds a U-shaped relationship in postmenopausal women, but a linear pattern in men over fifty.
Daniel J. Taylor, Jeroen Feher, Krzysztof Czechowicz, Ian Halliday, D. Rodney Hose, Rebecca Gosling, Louise Aubinière-Robb, Marcel van’t Veer, Danielle Keulards, Pim A.L. Tonino, Michel Rochette, Julian Gunn, Paul Morris
A heart artery is not just one pipe carrying one stream of blood. Its smaller branches constantly draw blood away, and a computer model may now estimate where that flow goes—without measuring every drop directly.What if an angiogram could do more than show where a coronary artery narrows? This study tests whether a numerical model can estimate how blood flow is distributed along the artery—and into its side branches.
When cancer reaches bone, the consequences can be devastating: pain, fractures, loss of movement, and sometimes death. To understand how that journey happens—and how to stop it—researchers build living models of it.Cancer cells do not simply appear in bone: they invade, enter circulation, and adapt to a new tissue. This review shows how animal models try to recreate that journey—and where the imitation breaks down.