Rule Monitoring#
Functions#
compare_rule_metrics#
- iguanas.rule_monitoring.compare_rule_metrics(ref_metrics: polars.DataFrame, curr_metrics: polars.DataFrame, thresholds: dict[str, float] | None = None) polars.DataFrame[source]#
Flag per-rule performance degradation between two periods.
Takes two
compute_metrics()outputs and returns the per-rule delta for every shared metric column, together with a boolean flag indicating whether the metric dropped by more than an optional threshold.- Parameters:
ref_metrics (pl.DataFrame) – Baseline metrics from
compute_metrics(). Must contain arulecolumn.curr_metrics (pl.DataFrame) – Current-period metrics from
compute_metrics(). Must contain arulecolumn. Only rules present in both DataFrames are compared (inner join onrule).thresholds (dict[str, float] | None, default=None) – Maximum allowed drop per metric, e.g.
{"precision": 0.05}flags rules whose precision fell by more than 5 pp. WhenNone, any negative delta is flagged.
- Returns:
One row per rule with columns:
rule{metric}_ref— metric value in the reference period{metric}_curr— metric value in the current period{metric}_delta—curr - ref(negative means degradation){metric}_degraded—Truewhen the drop exceeds the threshold
- Return type:
pl.DataFrame
Notes
What this does: an inner join on
rule, an arithmetic difference per shared metric column, and a comparison of that difference against a fixed threshold.What this does not do:
It is not a statistical drift test. No Kolmogorov-Smirnov test, Population Stability Index, Jensen-Shannon/KL divergence, or change-point detection is computed.
It performs no significance testing and returns no p-value or confidence interval, so a flagged drop may be sampling noise — particularly for low-volume rules. Inspect the
TP/FPcounts in the source metric tables before acting on a flag.It compares supervised metrics only, so both periods must be labelled. It cannot detect feature or covariate shift on unlabelled data.
It compares exactly two snapshots. There is no trend estimation over a series of periods.
Rules absent from either input are silently dropped by the inner join.
Choosing
thresholdsis a judgement call: with the default (None) every negative delta is flagged, including a one-sample fluctuation.Examples
>>> import polars as pl >>> from iguanas.metrics import compute_metrics >>> from iguanas.rule_monitoring import compare_rule_metrics >>> R_ref = pl.DataFrame({"rule_A": [True, False, True]}) >>> y_ref = pl.Series([True, True, True]) >>> R_curr = pl.DataFrame({"rule_A": [True, False, False]}) >>> y_curr = pl.Series([True, True, True]) >>> ref = compute_metrics(R_ref, y_ref) >>> curr = compute_metrics(R_curr, y_curr) >>> compare_rule_metrics(ref, curr, thresholds={"precision": 0.1})