Methodology — Clusters & Co-cluster Stability
Transparency is a feature. This page explains what the cluster panel shows and why — in plain terms. The precise formulas and numeric conventions are documented in our internal methodology; the summary here is deliberate, not an omission.
The correlation matrix above this panel is already ordered so that co-moving names sit next to each other. This panel expands that ordering into the tree it comes from — and, more importantly, tells you how much of that tree you should believe.
The stability matrix — the number to read
A clustering tree drawn from one sample of 10–40 stocks always looks confident, and much of it is sampling luck. So the load-bearing display here is not the tree: it is the co-cluster stability matrix. We resample your window a few hundred times (block bootstrap — blocks of consecutive days, so volatility clustering is preserved, and every name is resampled on the same days, so cross-correlation is preserved), rebuild the tree on each resample, cut every tree into the same fixed number of groups, and count: what fraction of resamples put each pair of names in the same group?
- Pairs near 100% group together in essentially every resample — that grouping is robust in this window.
- Pairs near 0% essentially never do.
- Middling numbers mean the tree's confident-looking merge is fragile — resampling dissolves it. Read those as estimation noise, not structure.
The resampling is deterministic: the same book and window always reproduce the same matrix.
The tree is one view
The dendrogram is drawn with one fixed, named method (average linkage on a standard correlation distance — the same method that orders the matrix above; we deliberately avoid "single linkage", which chains unrelated names together through intermediaries). It is presented as one view of one sample's hierarchy: useful for seeing the shape, honest only alongside the stability matrix.
How many groups? Shown, not asserted
The group count k is picked by a simple deterministic rule — cut the tree where the merge distances jump the most. That is a heuristic, not a discovery, so the panel also shows the k-instability distribution: applying the same rule to every resample, how often would the data have picked a different k? A spread-out distribution means the group count itself is sample-luck — we show that beside the matrix rather than hiding it inside one number.
The stress-day rebuild
The tree can also be rebuilt on the benchmark's worst-decile days — the same stress-day list every other stress number in this product uses, and only when that list is large enough to mean anything (at least 50 days spanning at least two distinct market episodes; otherwise you see the reason instead). Two honesty notes travel with it: correlations measured on selected days shift mechanically even when nothing changed (a well-documented statistical pitfall — Boyer, Gibson & Loretan, 1997), and no stability matrix is computed on the stress subset — there are too few days to resample honestly, so we abstain rather than fabricate one.
Group labels
Where every classified member of a group shares a GICS sector, the group is labeled with that shared fact (and how many members could be classified). Mixed groups show their composition instead. Labels state what the names have in common — they are not causes, and not a forecast.
Everything on this panel is a model output computed from the price window you loaded; excluded names and withheld displays always carry their reason.