Library / Machine learning in finance

Variable Importance Is Biased, and Permutation Extrapolates

Can feature-importance rankings change solely because inputs are correlated?

With correlated inputs, independently permuting one feature creates combinations the model never saw, so the resulting importance can depend on extrapolation. Compare conditional permutations or refit the model without that feature, and repeat the ranking across resamples before using it to remove an input.

Variable-importance bias depends on predictor properties and dependence structure.

Evidence map

AspectFinding
What it isTwo corrections to the importance measures that tree models report.
Key result / formulaStrobl, Boulesteix, Zeileis and Hothorn show that the impurity-based importance of a random forest is biased in favour of variables with many possible split points, continuous variables and high-cardinality categorical ones, and that the bias arises from the combination of the splitting criterion with bootstrap sampling with replacement.
Why it matters for backtestingFinancial features are strongly correlated with each other by construction, since most are transformations of the same price series, so both problems bite hard here.

What it is

The built-in measure is biased toward certain kinds of variable regardless of their predictive value, and the permutation alternative, applied naively, evaluates the model on combinations of feature values that do not occur. It extends [Feature importance: MDA vs MDI + SHAP].

Key result / formula

A variable with no relationship to the outcome can therefore outrank a genuinely predictive one purely because it offers more places to split. They propose using subsampling without replacement together with an unbiased split criterion, and a conditional permutation scheme when predictors are correlated. Hooker and Mentch address the permutation measure, which is often recommended as the unbiased alternative: permuting one feature independently of the others breaks their joint distribution, so the model is evaluated at feature combinations that do not exist in the data, and its behaviour there is extrapolation. With correlated features this produces importance rankings that reflect the model's arbitrary behaviour off the data manifold rather than its reliance on the feature.

Why it matters for backtesting

An importance ranking should therefore not be read as a statement about which feature matters, and in particular a decision to drop features based on such a ranking can remove a genuinely useful one. The defensible procedures are to compare models with and without a feature, refitting each time, which is expensive but answers the actual question, or to use a conditional scheme that respects the correlation structure. A cheaper diagnostic available to any user is to check whether the ranking is stable across resamples: an unstable ranking is measuring noise.

Source

Strobl, Boulesteix, Zeileis & Hothorn, "Bias in random forest variable importance measures: Illustrations, sources and a solution", BMC Bioinformatics 8, 2007, 25; Hooker & Mentch, "Please Stop Permuting Features: An Explanation and Alternatives", arXiv:1905.03151, 2019. Primary source

Related in the library