Naive multiple-testing corrections manufacture false negatives on correlated candidates
How should a selection test handle strategies that share nearly identical trades?
Bonferroni remains valid when tests are dependent, although correlated configurations can make it conservative. For a frozen candidate population, replay the entire search under a dependence-preserving null and compare the selected maximum with maxima from those null runs.
Bonferroni remains valid under dependence; this maximum-null analysis applies to a frozen population of configurations.

Evidence map
| Aspect | Finding |
|---|---|
| What it is | Bonferroni controls the family-wise error rate by the union bound whether tests are independent or dependent. |
| The pitfall it addresses | A search over closely related configurations raises a selection question: how unusual is the best observed statistic relative to the best produced by that same search under a suitable null? |
| How to apply it | For a declared population of configurations, freeze the candidate family and selection rule. |
What it is
Its validity does not require independent candidate strategies. Strong correlation can make the bound conservative because many configurations respond to the same return path.
The pitfall it addresses
Dividing alpha by the raw configuration count can be more conservative than needed for that particular question, but it is never invalid merely because configurations are correlated. A false negative cannot be diagnosed from correlation alone; power and a positive control matter.
How to apply it
Resample in a way that preserves the relevant temporal and cross-candidate dependence, rerun the complete selection for each resample, and compare the observed winner with the null distribution of winners. Romano-Wolf stepdown methods offer family-wise inference under their bootstrap conditions. Check those conditions and report how many configurations were searched. The null of the maximum belongs to choosing a winner from a population of configurations; it is not an automatic hurdle for each newly added lever or a single predeclared sealed test.
Worked family
Consider 100 moving-average rules with neighboring windows. Their return paths may be similar. A Bonferroni bound across 100 tests remains valid, while a calibrated null of the maximum may be less conservative for choosing one winner from that frozen family. On each null path, evaluate all 100 rules and save the best statistic. Compare the observed best with the resulting distribution, keeping the same costs and selection procedure. A new asset, horizon or signal family changes the population and requires a newly declared analysis. Report the search history without converting the count into a universal threshold for unrelated tests.
Source
Romano and Wolf, Exact and Approximate Stepdown Methods for Multiple Hypothesis Testing, Journal of the American Statistical Association 100(469), 2005; doi:10.1198/016214504000000539. Primary source