Cynthia Mutua

Knowledge discovery · Clinical machine learning

What can routine clinical data reveal about Alzheimer's that clinicians might miss?

A complete knowledge-discovery pipeline over 2,149 patients: explore, cluster, mine rules, classify, then validate that the same signals hold across independent methods. The aim was never raw accuracy; it was a model a clinician can read, audit, and trust.

Role
Designed & built end to end
Scope
2,149 patients · 35 features
Pipeline
5 methods · cross-validated
93.8%
accuracy
91% sens · 95% spec
39
interpretable
IF–THEN rules
2,149
patients
35 features
5
independent
methods

Pipeline

01
Ingest
load, validate, 70/30 stratified split
pandas
02
Explore
correlations, distributions, balance
seaborn
03
Cluster
k-means + hierarchical, k=2–7
scikit-learn
04
Mine
association rules, Apriori
mlxtend
05
Classify
decision tree, grid search
scikit-learn
06
Validate
cross-method convergence
numpy

01  /  Context

Diagnosis comes late. The data might come earlier.

Alzheimer's affects over 50 million people worldwide and is the most common cause of dementia. By the time most patients are diagnosed, the disease has progressed: the memory slips and missed appointments are noticed long before the clinical call arrives, and the window for early intervention has closed.

The question here is operational, not theoretical: can data collected at any standard visit surface patterns that flag risk sooner? And can it do so transparently enough that a clinician would actually use it?

02  /  Exploration

Functional decline outranks the memory test.

Data balance

A realistic class split

35.4% of patients carry a diagnosis, a 1.83:1 imbalance that mirrors clinical reality and drove the choice of stratified sampling and class-weighted models downstream.

Donut: 35.4 percent of 2,149 patients have an Alzheimer's diagnosis

fig.01 · diagnosis distribution

What predicts diagnosis

Daily-living ability leads, not cognition alone

The strongest signals are FunctionalAssessment (r = −0.36) and ADL (r = −0.33), how well a person manages daily life, which outrank even MMSE (r = −0.24). Subjective MemoryComplaints ranks third (r = +0.31). No single feature exceeds 0.36: one test is never enough.

Bar chart of top correlations with diagnosis, FunctionalAssessment highest at -0.36

fig.02 · top five features correlated with diagnosis

03  /  Clustering

A "failed" model that told us something true.

k-means across k = 2 to 7, four validation metrics, and hierarchical clustering all agreed: silhouette of 0.05 to 0.06 (far under the 0.30 bar), no elbow, and the two algorithms agreeing almost not at all (ARI = 0.025). The clusters were artifacts, not structure. So I reframed it: severity here lies on a continuum, not in discrete subtypes, consistent with the NIA-AA framework. Knowing when not to force a result is its own kind of rigor.

Silhouette score across k from 2 to 7, flat near 0.05, far below the 0.30 threshold

fig.03 · silhouette across k, no real separation at any k

04  /  Association rules

Which combinations of symptoms predict it?

With no subtypes to find, the question shifted to co-occurrence. Apriori surfaced 39 interpretable rules, each at 60% confidence or higher with lift ≥ 1.5. The strongest:

IF MemoryComplaints AND MMSE = severe impairment → Alzheimer's

84%confidence
2.36×lift vs base rate
51%of rules use MemoryComplaints

What the rules are built from

Three signals carry them

Across all 39 rules, MemoryComplaints dominates, then severe MMSE and BehavioralProblems. The two in flame also top the decision tree, the first sign of convergence.

Bar chart of features in rule antecedents, MemoryComplaints 30, MMSE-severe 11, BehavioralProblems 9

fig.04 · antecedent frequency across the 39 rules

05  /  Classification

A model a clinician can actually read.

A 12-configuration grid search landed on a decision tree (depth 5, min 10 samples per leaf): 93.8% accuracy, 0.912 F1, 91% sensitivity, 95% specificity, fully transparent, with a path you can follow for any patient.

64.7%
baseline (predict "no AD")
81.6%
logistic regression
93.8%
interpretable tree

The tree, recreated from the trained model

The first split is FunctionalAssessment, not a memory test. The structure mirrors how a clinician reasons: check daily function first, then let cognition and behavior refine the call.

fig.05 · top three levels of the depth-5 tree

645 held-out patients

pred no AD
pred AD
actual
no AD
397cleared
20false alarm
actual
AD
20missed
208caught

91% sensitivity · 95% specificity

Five features drive 98.4% of the prediction

The two marked also surface in the association rules, the convergence that makes them trustworthy.

FunctionalAssessment
23.3%
MMSE
21.2%
ADL
18.4%
BehavioralProblems · also in rules
18.3%
MemoryComplaints · also in rules
17.2%

Top three account for 62.9%; everything past the top five contributes under 1%.

06  /  Validation

Three independent methods, one answer.

The strongest validation is not a single score; it is agreement. MemoryComplaints and BehavioralProblems rank high in correlation, in the association rules, and in the tree splits, three methods sharing no assumptions, pointing at the same clinically central signals.

01
93.8%

Interpretable beats black box

A transparent tree outperforms logistic regression by 12 points; opacity buys nothing here.

02
Continuum

No discrete subtypes exist

Silhouette ≈ 0.06, ARI = 0.025: severity varies continuously, consistent with current biology.

03
84%

One actionable screening rule

Memory complaints plus severe MMSE impairment gives 84% confidence, a usable guideline.

04
Function

Function outranks cognition

FunctionalAssessment (23.3%) beats MMSE (21.2%) in every method; daily-living ability is the earlier signal.

05
51%

Subjective reports matter

MemoryComplaints appears in 51% of all rules; patient and caregiver concern is hard signal, not noise.

06
×3

Convergent validation

Two features matter in correlation, rules, and tree splits at once, independent evidence of their centrality.

+  /  How every number was produced

MethodQuestionResultVerdict
Exploratory analysisWhich features relate to diagnosis?FunctionalAssessment r = −0.36; no single predictorfoundation
K-means (k=2–7)Do discrete subtypes exist?Silhouette 0.05–0.06; no elbowscientific finding
Hierarchical clusteringDoes another algorithm agree?ARI = 0.025 vs k-meansconfirms continuum
Association rules (Apriori)Which combinations predict it?39 rules; confidence 60–84%; lift ≤ 2.36actionable
Decision tree (depth 5)Can it be accurate and readable?93.8% accuracy; 0.912 F1; 91% sensitivityexcellent
Cross-method integrationDo findings replicate?MemoryComplaints & BehavioralProblems convergevalidated
stack
Python 3.10pandasNumPyscikit-learnmlxtendmatplotlibseaborn
demonstrates
pipeline designEDAunsupervised + supervised MLhyperparameter searchvalidation rigorinterpretabilityscientific framing

Conclusion

Transparent models can reach excellent clinical performance.

93.8% accuracy, 91% sensitivity, 95% specificity, all from a decision tree a clinician can read, audit, and trust. No black box required. The findings point to a shift in screening: weigh functional assessment alongside cognitive testing, take subjective complaints seriously, and watch for the behavioral and cognitive signals that converge across every method in this study.

Cynthia Mutua·LinkedIn·GitHub