Knowledge discovery · Clinical machine learning
What can routine clinical data reveal about Alzheimer's that clinicians might miss?
A complete knowledge-discovery pipeline over 2,149 patients: explore, cluster, mine rules, classify, then validate that the same signals hold across independent methods. The aim was never raw accuracy; it was a model a clinician can read, audit, and trust.
91% sens · 95% spec
IF–THEN rules
35 features
methods
Pipeline
01 / Context
Diagnosis comes late. The data might come earlier.
Alzheimer's affects over 50 million people worldwide and is the most common cause of dementia. By the time most patients are diagnosed, the disease has progressed: the memory slips and missed appointments are noticed long before the clinical call arrives, and the window for early intervention has closed.
The question here is operational, not theoretical: can data collected at any standard visit surface patterns that flag risk sooner? And can it do so transparently enough that a clinician would actually use it?
02 / Exploration
Functional decline outranks the memory test.
Data balance
A realistic class split
35.4% of patients carry a diagnosis, a 1.83:1 imbalance that mirrors clinical reality and drove the choice of stratified sampling and class-weighted models downstream.
fig.01 · diagnosis distribution
What predicts diagnosis
Daily-living ability leads, not cognition alone
The strongest signals are FunctionalAssessment (r = −0.36) and ADL (r = −0.33), how well a person manages daily life, which outrank even MMSE (r = −0.24). Subjective MemoryComplaints ranks third (r = +0.31). No single feature exceeds 0.36: one test is never enough.
fig.02 · top five features correlated with diagnosis
03 / Clustering
A "failed" model that told us something true.
k-means across k = 2 to 7, four validation metrics, and hierarchical clustering all agreed: silhouette of 0.05 to 0.06 (far under the 0.30 bar), no elbow, and the two algorithms agreeing almost not at all (ARI = 0.025). The clusters were artifacts, not structure. So I reframed it: severity here lies on a continuum, not in discrete subtypes, consistent with the NIA-AA framework. Knowing when not to force a result is its own kind of rigor.
fig.03 · silhouette across k, no real separation at any k
04 / Association rules
Which combinations of symptoms predict it?
With no subtypes to find, the question shifted to co-occurrence. Apriori surfaced 39 interpretable rules, each at 60% confidence or higher with lift ≥ 1.5. The strongest:
IF MemoryComplaints AND MMSE = severe impairment → Alzheimer's
What the rules are built from
Three signals carry them
Across all 39 rules, MemoryComplaints dominates, then severe MMSE and BehavioralProblems. The two in flame also top the decision tree, the first sign of convergence.
fig.04 · antecedent frequency across the 39 rules
05 / Classification
A model a clinician can actually read.
A 12-configuration grid search landed on a decision tree (depth 5, min 10 samples per leaf): 93.8% accuracy, 0.912 F1, 91% sensitivity, 95% specificity, fully transparent, with a path you can follow for any patient.
The tree, recreated from the trained model
The first split is FunctionalAssessment, not a memory test. The structure mirrors how a clinician reasons: check daily function first, then let cognition and behavior refine the call.
fig.05 · top three levels of the depth-5 tree
645 held-out patients
no AD
AD
91% sensitivity · 95% specificity
Five features drive 98.4% of the prediction
The two marked also surface in the association rules, the convergence that makes them trustworthy.
Top three account for 62.9%; everything past the top five contributes under 1%.
06 / Validation
Three independent methods, one answer.
The strongest validation is not a single score; it is agreement. MemoryComplaints and BehavioralProblems rank high in correlation, in the association rules, and in the tree splits, three methods sharing no assumptions, pointing at the same clinically central signals.
Interpretable beats black box
A transparent tree outperforms logistic regression by 12 points; opacity buys nothing here.
No discrete subtypes exist
Silhouette ≈ 0.06, ARI = 0.025: severity varies continuously, consistent with current biology.
One actionable screening rule
Memory complaints plus severe MMSE impairment gives 84% confidence, a usable guideline.
Function outranks cognition
FunctionalAssessment (23.3%) beats MMSE (21.2%) in every method; daily-living ability is the earlier signal.
Subjective reports matter
MemoryComplaints appears in 51% of all rules; patient and caregiver concern is hard signal, not noise.
Convergent validation
Two features matter in correlation, rules, and tree splits at once, independent evidence of their centrality.
+ / How every number was produced
| Method | Question | Result | Verdict |
|---|---|---|---|
| Exploratory analysis | Which features relate to diagnosis? | FunctionalAssessment r = −0.36; no single predictor | foundation |
| K-means (k=2–7) | Do discrete subtypes exist? | Silhouette 0.05–0.06; no elbow | scientific finding |
| Hierarchical clustering | Does another algorithm agree? | ARI = 0.025 vs k-means | confirms continuum |
| Association rules (Apriori) | Which combinations predict it? | 39 rules; confidence 60–84%; lift ≤ 2.36 | actionable |
| Decision tree (depth 5) | Can it be accurate and readable? | 93.8% accuracy; 0.912 F1; 91% sensitivity | excellent |
| Cross-method integration | Do findings replicate? | MemoryComplaints & BehavioralProblems converge | validated |