Rule Cross-Validation#

Functions#

validate_rules_cv#

iguanas.rule_cv.validate_rules_cv(X: polars.DataFrame, y: polars.Series, rules: list[str], n_folds: int = 5, cv_metrics: list[str] | None = None, weight_column: str | None = None, shuffle: bool = True, random_state: int | None = None) → polars.DataFrame[source]#

Evaluate rule stability across K folds.

Validates already-generated rules on K folds without re-generating them. For each fold the rules are evaluated on the validation split, and the mean, standard deviation, and minimum of each requested metric across folds are returned. Rules with a high {metric}_cv_std or a low {metric}_cv_min are likely over-fitted to the training data.

Warning

The folds are not truly held out, and the reported statistics are optimistically biased. This function does not generate rules; it takes rules that were already produced from the entire dataset and then re-scores them on subsets of that same dataset. Every validation fold was therefore part of the data used to choose the rules’ features and thresholds. {metric}_cv_mean, {metric}_cv_std and {metric}_cv_min must not be reported as out-of-sample or generalisation estimates. See Notes for an unbiased protocol.

Parameters:
  • X (pl.DataFrame) – Feature DataFrame. Must contain all columns referenced in rules.

  • y (pl.Series) – Target series (boolean or binary).

  • rules (list[str]) – Rule expressions to validate. Typically obtained from apply_filter_and_deduplicate_rules() or similar.

  • n_folds (int, default=5) – Number of CV folds. The data is split into n_folds contiguous blocks (after optional shuffling).

  • cv_metrics (list[str] | None, default=None) – Metric names to compute CV statistics for. When None, defaults to ["precision", "recall", "f1"].

  • weight_column (str | None, default=None) – Name of a column in X to use as sample weights when computing metrics. If None, all samples are weighted equally.

  • shuffle (bool, default=True) – Whether to shuffle the row indices before splitting into folds.

  • random_state (int | None, default=None) – Random seed for reproducibility when shuffle=True.

Returns:

One row per rule with columns:

  • rule

  • {metric}_cv_mean — mean of the metric across folds

  • {metric}_cv_std — standard deviation across folds

  • {metric}_cv_min — worst-fold value (lowest)

Sorted by rule name.

Return type:

pl.DataFrame

Examples

>>> import polars as pl
>>> X = pl.DataFrame({"age": [25, 30, 35, 40, 45, 50, 55, 60, 65, 70]})
>>> y = pl.Series([0, 0, 0, 0, 0, 1, 1, 1, 1, 1])
>>> rules = ['(X["age"] >= 50)']
>>> validate_rules_cv(X, y, rules, n_folds=2, random_state=0)

Notes

Optimism bias. The intended usage is: generate rules on X/y, then call this function on the same X/y. Rule selection has therefore already seen every fold, which leaks information into each “validation” split. The consequences are:

  • {metric}_cv_mean is inflated relative to true held-out performance.

  • {metric}_cv_std is deflated and {metric}_cv_min is inflated, so the numbers understate how badly a rule can degrade on genuinely unseen data. They are a lower bound on overfitting, not a measure of it.

  • A rule that looks stable here can still fail out of sample; a rule that looks unstable here is almost certainly unstable.

Use these statistics as a relative screen to rank and discard fragile rules within a candidate set — never as an estimate of deployment performance, and never as a headline result.

For an unbiased estimate, use a nested protocol in which rule generation happens strictly inside the training split of each outer fold:

  1. Split the data into outer folds (or a single held-out test set).

  2. For each outer fold, run the full pipeline — weight transformations, grid search, filtering, deduplication — on the training portion only.

  3. Evaluate the resulting rules on the outer fold, which no step of the pipeline has seen.

  4. Aggregate across outer folds.

Only step 3 yields a defensible generalisation estimate.

See also

apply_filter_and_deduplicate_rules

Complete evaluation pipeline that produces the rules input for this function.

identify_unstable_rules#

iguanas.rule_cv.identify_unstable_rules(cv_result: polars.DataFrame, metric: str = 'f1', max_std: float = 0.05, min_mean: float | None = None) → polars.DataFrame[source]#

Return rules whose cross-fold metric is unstable or consistently poor.

Filters the output of validate_rules_cv() to surface rules that are likely over-fitted (high variance across folds) or simply weak (low mean metric).

Warning

Inherits the optimism bias of validate_rules_cv(): the folds were already seen during rule generation, so {metric}_cv_std is deflated and {metric}_cv_mean inflated. This function is therefore a one-sided screen — rules it flags are genuinely unstable, but rules it does not flag are not thereby shown to generalise. Absence from the returned set is not evidence of stability.

Parameters:
  • cv_result (pl.DataFrame) – Output of validate_rules_cv(). Must contain columns {metric}_cv_std and (if min_mean is set) {metric}_cv_mean.

  • metric (str, default="f1") – Metric prefix to inspect (must match one used in validate_rules_cv()).

  • max_std (float, default=0.05) – Rules whose {metric}_cv_std exceeds this threshold are flagged as unstable.

  • min_mean (float | None, default=None) – If provided, also flag rules whose {metric}_cv_mean is below this threshold (consistently weak rules).

Returns:

Subset of cv_result containing only flagged rules, sorted by {metric}_cv_std descending (most unstable first).

Return type:

pl.DataFrame

Raises:

ValueError – If required columns are absent from cv_result.

Examples

>>> cv = validate_rules_cv(X, y, rules, n_folds=5)
>>> identify_unstable_rules(cv, metric="f1", max_std=0.05, min_mean=0.3)