Rule Monitoring#

Functions#

compare_rule_metrics#

iguanas.rule_monitoring.compare_rule_metrics(ref_metrics: polars.DataFrame, curr_metrics: polars.DataFrame, thresholds: dict[str, float] | None = None) → polars.DataFrame[source]#

Flag per-rule performance degradation between two periods.

Takes two compute_metrics() outputs and returns the per-rule delta for every shared metric column, together with a boolean flag indicating whether the metric dropped by more than an optional threshold.

Parameters:
  • ref_metrics (pl.DataFrame) – Baseline metrics from compute_metrics(). Must contain a rule column.

  • curr_metrics (pl.DataFrame) – Current-period metrics from compute_metrics(). Must contain a rule column. Only rules present in both DataFrames are compared (inner join on rule).

  • thresholds (dict[str, float] | None, default=None) – Maximum allowed drop per metric, e.g. {"precision": 0.05} flags rules whose precision fell by more than 5 pp. When None, any negative delta is flagged.

Returns:

One row per rule with columns:

  • rule

  • {metric}_ref — metric value in the reference period

  • {metric}_curr — metric value in the current period

  • {metric}_delta — curr - ref (negative means degradation)

  • {metric}_degraded — True when the drop exceeds the threshold

Return type:

pl.DataFrame

Notes

What this does: an inner join on rule, an arithmetic difference per shared metric column, and a comparison of that difference against a fixed threshold.

What this does not do:

  • It is not a statistical drift test. No Kolmogorov-Smirnov test, Population Stability Index, Jensen-Shannon/KL divergence, or change-point detection is computed.

  • It performs no significance testing and returns no p-value or confidence interval, so a flagged drop may be sampling noise — particularly for low-volume rules. Inspect the TP/FP counts in the source metric tables before acting on a flag.

  • It compares supervised metrics only, so both periods must be labelled. It cannot detect feature or covariate shift on unlabelled data.

  • It compares exactly two snapshots. There is no trend estimation over a series of periods.

  • Rules absent from either input are silently dropped by the inner join.

Choosing thresholds is a judgement call: with the default (None) every negative delta is flagged, including a one-sample fluctuation.

Examples

>>> import polars as pl
>>> from iguanas.metrics import compute_metrics
>>> from iguanas.rule_monitoring import compare_rule_metrics
>>> R_ref = pl.DataFrame({"rule_A": [True, False, True]})
>>> y_ref = pl.Series([True, True, True])
>>> R_curr = pl.DataFrame({"rule_A": [True, False, False]})
>>> y_curr = pl.Series([True, True, True])
>>> ref = compute_metrics(R_ref, y_ref)
>>> curr = compute_metrics(R_curr, y_curr)
>>> compare_rule_metrics(ref, curr, thresholds={"precision": 0.1})