Rule Cross-Validation#
Functions#
validate_rules_cv#
- iguanas.rule_cv.validate_rules_cv(X: polars.DataFrame, y: polars.Series, rules: list[str], n_folds: int = 5, cv_metrics: list[str] | None = None, weight_column: str | None = None, shuffle: bool = True, random_state: int | None = None) polars.DataFrame[source]#
Evaluate rule stability across K folds.
Validates already-generated rules on K folds without re-generating them. For each fold the rules are evaluated on the validation split, and the mean, standard deviation, and minimum of each requested metric across folds are returned. Rules with a high
{metric}_cv_stdor a low{metric}_cv_minare likely over-fitted to the training data.Warning
The folds are not truly held out, and the reported statistics are optimistically biased. This function does not generate rules; it takes rules that were already produced from the entire dataset and then re-scores them on subsets of that same dataset. Every validation fold was therefore part of the data used to choose the rules’ features and thresholds.
{metric}_cv_mean,{metric}_cv_stdand{metric}_cv_minmust not be reported as out-of-sample or generalisation estimates. See Notes for an unbiased protocol.- Parameters:
X (pl.DataFrame) – Feature DataFrame. Must contain all columns referenced in
rules.y (pl.Series) – Target series (boolean or binary).
rules (list[str]) – Rule expressions to validate. Typically obtained from
apply_filter_and_deduplicate_rules()or similar.n_folds (int, default=5) – Number of CV folds. The data is split into
n_foldscontiguous blocks (after optional shuffling).cv_metrics (list[str] | None, default=None) – Metric names to compute CV statistics for. When
None, defaults to["precision", "recall", "f1"].weight_column (str | None, default=None) – Name of a column in
Xto use as sample weights when computing metrics. IfNone, all samples are weighted equally.shuffle (bool, default=True) – Whether to shuffle the row indices before splitting into folds.
random_state (int | None, default=None) – Random seed for reproducibility when
shuffle=True.
- Returns:
One row per rule with columns:
rule{metric}_cv_mean— mean of the metric across folds{metric}_cv_std— standard deviation across folds{metric}_cv_min— worst-fold value (lowest)
Sorted by
rulename.- Return type:
pl.DataFrame
Examples
>>> import polars as pl >>> X = pl.DataFrame({"age": [25, 30, 35, 40, 45, 50, 55, 60, 65, 70]}) >>> y = pl.Series([0, 0, 0, 0, 0, 1, 1, 1, 1, 1]) >>> rules = ['(X["age"] >= 50)'] >>> validate_rules_cv(X, y, rules, n_folds=2, random_state=0)
Notes
Optimism bias. The intended usage is: generate rules on
X/y, then call this function on the sameX/y. Rule selection has therefore already seen every fold, which leaks information into each “validation” split. The consequences are:{metric}_cv_meanis inflated relative to true held-out performance.{metric}_cv_stdis deflated and{metric}_cv_minis inflated, so the numbers understate how badly a rule can degrade on genuinely unseen data. They are a lower bound on overfitting, not a measure of it.A rule that looks stable here can still fail out of sample; a rule that looks unstable here is almost certainly unstable.
Use these statistics as a relative screen to rank and discard fragile rules within a candidate set — never as an estimate of deployment performance, and never as a headline result.
For an unbiased estimate, use a nested protocol in which rule generation happens strictly inside the training split of each outer fold:
Split the data into outer folds (or a single held-out test set).
For each outer fold, run the full pipeline — weight transformations, grid search, filtering, deduplication — on the training portion only.
Evaluate the resulting rules on the outer fold, which no step of the pipeline has seen.
Aggregate across outer folds.
Only step 3 yields a defensible generalisation estimate.
See also
- apply_filter_and_deduplicate_rules
Complete evaluation pipeline that produces the
rulesinput for this function.
identify_unstable_rules#
- iguanas.rule_cv.identify_unstable_rules(cv_result: polars.DataFrame, metric: str = 'f1', max_std: float = 0.05, min_mean: float | None = None) polars.DataFrame[source]#
Return rules whose cross-fold metric is unstable or consistently poor.
Filters the output of
validate_rules_cv()to surface rules that are likely over-fitted (high variance across folds) or simply weak (low mean metric).Warning
Inherits the optimism bias of
validate_rules_cv(): the folds were already seen during rule generation, so{metric}_cv_stdis deflated and{metric}_cv_meaninflated. This function is therefore a one-sided screen — rules it flags are genuinely unstable, but rules it does not flag are not thereby shown to generalise. Absence from the returned set is not evidence of stability.- Parameters:
cv_result (pl.DataFrame) – Output of
validate_rules_cv(). Must contain columns{metric}_cv_stdand (ifmin_meanis set){metric}_cv_mean.metric (str, default="f1") – Metric prefix to inspect (must match one used in
validate_rules_cv()).max_std (float, default=0.05) – Rules whose
{metric}_cv_stdexceeds this threshold are flagged as unstable.min_mean (float | None, default=None) – If provided, also flag rules whose
{metric}_cv_meanis below this threshold (consistently weak rules).
- Returns:
Subset of
cv_resultcontaining only flagged rules, sorted by{metric}_cv_stddescending (most unstable first).- Return type:
pl.DataFrame
- Raises:
ValueError – If required columns are absent from
cv_result.
Examples
>>> cv = validate_rules_cv(X, y, rules, n_folds=5) >>> identify_unstable_rules(cv, metric="f1", max_std=0.05, min_mean=0.3)