Rule Formatting#

Functions#

simplify_rule#

iguanas.rule_formatting.simplify_rule(rule: str) → str[source]#

Simplify a rule by removing redundant conditions on the same column.

When multiple conditions exist on the same column, keeps only the most restrictive:

  • For lower bounds (>, >=): keeps the highest threshold, preferring > over >= when equal

  • For upper bounds (<, <=): keeps the lowest threshold, preferring < over <= when equal

Parameters:

rule (str) – Rule string with conditions like (X[“col”] > val) & (X[“col”] >= val).

Returns:

Simplified rule string with redundant conditions removed. Column order is preserved based on first appearance.

Return type:

str

Examples

>>> simplify_rule('(X["amount"] >= 100.0) & (X["amount"] > 100.0)')
'(X["amount"] > 100.0)'
>>> simplify_rule('(X["amount"] < 100.0) & (X["amount"] <= 100.0)')
'(X["amount"] < 100.0)'
>>> simplify_rule('(X["a"] >= 50) & (X["b"] < 10) & (X["a"] > 100)')
'(X["a"] > 100) & (X["b"] < 10)'

rule_to_sql#

iguanas.rule_formatting.rule_to_sql(rule: str, table_alias: str | None = None) → str[source]#

Convert a rule expression string to a SQL WHERE clause.

Translates Iguanas rule notation (X["col"] op value) into standard SQL predicate syntax suitable for use in a WHERE or CASE WHEN clause.

Parameters:
  • rule (str) – Rule expression using X["col"] notation with & / | operators, e.g. '(X["age"] > 30) & (X["income"] < 50000)'.

  • table_alias (str | None, default=None) – Optional table or CTE alias to prefix column references with. For example, table_alias="t" turns age > 30 into t.age > 30.

Returns:

SQL WHERE clause string.

Return type:

str

Examples

>>> rule_to_sql('(X["age"] > 30) & (X["income"] < 50000)')
'(age > 30.0) AND (income < 50000.0)'
>>> rule_to_sql('(X["age"] > 30) | (X["flag"] == 1)', table_alias="t")
'(t.age > 30.0) OR (t.flag = 1.0)'

format_floats_as_integers#

iguanas.rule_formatting.format_floats_as_integers(rule: str, int_columns: list[str]) → str[source]#

Convert float thresholds to integers for the given columns.

Uses ceiling for >=/< and floor for >/<=, so the integer boundary preserves the original condition’s semantics. ==/!= are left unchanged since there is no boundary to round.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • int_columns (list[str]) – Columns whose float thresholds should be converted to integers.

Returns:

Rule string with integer thresholds for the given columns.

Return type:

str

Examples

>>> format_floats_as_integers('(X["a"] >= 0.1) & (X["b"] >= 9.1)', ["a"])
'(X["a"] >= 1) & (X["b"] >= 9.1)'

add_missing_value_conditions#

iguanas.rule_formatting.add_missing_value_conditions(rule: str, mapping: dict[str, float]) → str[source]#

Append an is_null() clause to conditions satisfied by an imputed value.

When a column’s nulls were filled with a value that also satisfies an existing condition, that condition implicitly matches originally-null rows too. This makes that explicit by OR-ing in X[col].is_null().

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • mapping (dict[str, float]) – Maps column name to the value nulls were imputed with.

Returns:

Rule string with is_null() clauses added where relevant.

Return type:

str

Examples

>>> add_missing_value_conditions('(X["a"] < 1)', {"a": 0})
'((X["a"] < 1) | X["a"].is_null())'

decode_string_imputation#

iguanas.rule_formatting.decode_string_imputation(rule: str, mapping: dict[str, str]) → str[source]#

Convert an equality on a string-imputed placeholder into is_null().

For columns where nulls were filled with a placeholder string (e.g. gators StringImputer’s default "MISSING"), rewrites an equality condition on that placeholder back to an is-null check.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • mapping (dict[str, str]) – Maps column name to the placeholder value strings were imputed with.

Returns:

Rule string with placeholder equality conditions decoded.

Return type:

str

Examples

>>> decode_string_imputation('(X["status"] == "MISSING")', {"status": "MISSING"})
'X["status"].is_null()'

decode_numeric_encodings#

iguanas.rule_formatting.decode_numeric_encodings(rule: str, mapping: dict[str, dict[str, float]]) → str[source]#

Reverse a numeric category encoding back to the original category labels.

For encodings where each category maps to a numeric statistic (e.g. WOE score, category count, mean target value), finds which categories satisfy the condition’s operator/threshold and replaces it with an equality (single match) or .is_in() (multiple matches) condition on the original values.

Parameters:
  • rule (str) – Rule expression using X["col"] notation, where col holds the encoded numeric values.

  • mapping (dict[str, dict[str, float]]) – Maps column name to a dict of {category: encoded_value}.

Returns:

Rule string with conditions decoded back to category labels.

Return type:

str

Examples

>>> mapping = {"A": {"a": 1, "b": 2, "c": 3}}
>>> decode_numeric_encodings('(X["A"] >= 2)', mapping)
'(X["A"].is_in(["b", "c"]))'

format_as_boolean_conditions#

iguanas.rule_formatting.format_as_boolean_conditions(rule: str, bool_columns: list[str]) → str[source]#

Convert True/False-like condition values to Python booleans.

Recognises two condition shapes for the given columns:

  • “True”/”true”/”1” and “False”/”false”/”0” (quoted or bare), combined with ==/!=.

  • Raw numeric threshold splits on >=/>/</<= (as produced by a model trained on a boolean column cast to float), where the threshold unambiguously selects only the 0.0 or only the 1.0 value.

Both shapes are rewritten as == True/== False.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • bool_columns (list[str]) – Columns to treat as boolean.

Returns:

Rule string with boolean conditions normalised.

Return type:

str

Examples

>>> format_as_boolean_conditions('(X["flag"] != "False")', ["flag"])
'(X["flag"] == True)'
>>> format_as_boolean_conditions('(X["flag"] >= 1.0)', ["flag"])
'(X["flag"] == True)'
>>> format_as_boolean_conditions('(X["flag"] < 1.0)', ["flag"])
'(X["flag"] == False)'

decode_onehot_encodings#

iguanas.rule_formatting.decode_onehot_encodings(rule: str, mapping: dict[str, tuple[str, str]], null_category: str | None = None) → str[source]#

Reverse a one-hot encoding back to a categorical condition.

One-hot encoders typically produce a binary column per category, split by the model at 0.5. This converts those splits back to equality/ inequality conditions on the original categorical column. If one category represents “value was null” (null_category), it’s rendered as is_null()/~is_null() instead of a literal category comparison.

Parameters:
  • rule (str) – Rule expression using X["col"] notation, where col is the one-hot encoded binary column.

  • mapping (dict[str, tuple[str, str]]) – Maps encoded column name to (original_col, category).

  • null_category (str | None, optional) – Category value that represents an originally-null value, by default None.

Returns:

Rule string with one-hot conditions decoded back to category labels.

Return type:

str

Examples

>>> mapping = {"status__active": ("status", "active")}
>>> decode_onehot_encodings('(X["status__active"] >= 0.5)', mapping)
'(X["status"] == "active")'

decode_null_indicators#

iguanas.rule_formatting.decode_null_indicators(rule: str, mapping: dict[str, str]) → str[source]#

Convert null-indicator binary columns to is_null() conditions.

Parameters:
  • rule (str) – Rule expression using X["col"] notation, where col is a binary column indicating whether the original column was null.

  • mapping (dict[str, str]) – Maps encoded column name to the original column name.

Returns:

Rule string with is-null conditions decoded.

Return type:

str

Examples

>>> decode_null_indicators('(X["amount__is_null"] >= 0.5)', {"amount__is_null": "amount"})
'X["amount"].is_null()'

decode_discretized_bins#

iguanas.rule_formatting.decode_discretized_bins(rule: str, mapping: dict[str, list[float]]) → str[source]#

Reverse a discretizer’s bin index back to a threshold on the original column.

Discretizers (equal-width, equal-frequency, quantile, k-means, tree-based, geometric, custom bin edges, …) all replace a numeric column with an integer bin index. Given the fitted bin edges, this converts a condition on the bin index back to a condition on the original numeric column.

Parameters:
  • rule (str) – Rule expression using X["col"] notation, where col holds the bin index.

  • mapping (dict[str, list[float]]) – Maps column name to its sorted bin edges, where edges[i] is the lower boundary of bin i.

Returns:

Rule string with bin-index conditions decoded to original thresholds.

Return type:

str

Examples

>>> decode_discretized_bins('(X["amount"] >= 2)', {"amount": [0, 10, 50, 200]})
'(X["amount"] >= 50)'

decode_scaled_thresholds#

iguanas.rule_formatting.decode_scaled_thresholds(rule: str, mapping: dict[str, collections.abc.Callable[[float], float]]) → str[source]#

Reverse a monotonic numeric scaling back to a threshold on the original column.

Works for any strictly increasing scaler (standardisation, min-max, log1p, Box-Cox, Yeo-Johnson, arcsinh, power, robust, …) since the >=/</etc. ordering is preserved - pass the scaler’s inverse transform as the mapping value.

Parameters:
  • rule (str) – Rule expression using X["col"] notation, where col holds the scaled values.

  • mapping (dict[str, Callable[[float], float]]) – Maps column name to a function that inverts the scaling (e.g. scaler.inverse_transform).

Returns:

Rule string with scaled thresholds decoded back to original values.

Return type:

str

Examples

>>> decode_scaled_thresholds('(X["amount"] >= 5.0)', {"amount": lambda x: x * 2})
'(X["amount"] >= 10.0)'

quote_string_values#

iguanas.rule_formatting.quote_string_values(rule: str, columns: list[str]) → str[source]#

Wrap bare (unquoted) condition values in double quotes.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • columns (list[str]) – Columns whose values should be quoted.

Returns:

Rule string with bare values quoted.

Return type:

str

Examples

>>> quote_string_values('(X["col"] == retail)', ["col"])
'(X["col"] == "retail")'

round_thresholds#

iguanas.rule_formatting.round_thresholds(rule: str, columns: list[str], ndigits: int = 2) → str[source]#

Round numeric thresholds to a fixed number of decimal places.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • columns (list[str]) – Columns to round.

  • ndigits (int, optional) – Decimal places, by default 2.

Returns:

Rule string with rounded thresholds.

Return type:

str

Examples

>>> round_thresholds('(X["amount"] >= 1234.56789)', ["amount"])
'(X["amount"] >= 1234.57)'

drop_null_clauses#

iguanas.rule_formatting.drop_null_clauses(rule: str, columns: list[str]) → str[source]#

Strip | X[col].is_null() clauses added for always-imputed columns.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • columns (list[str]) – Columns for which to strip the is-null clause, e.g. because the column is never actually null once imputed.

Returns:

Rule string with the is-null clauses removed.

Return type:

str

Examples

>>> drop_null_clauses('((X["amount"] >= 5.0) | X["amount"].is_null())', ["amount"])
'(X["amount"] >= 5.0)'

drop_not_null_conditions#

iguanas.rule_formatting.drop_not_null_conditions(rule: str, columns: list[str]) → str[source]#

Drop standalone (~X[col].is_null()) conditions for given columns.

Parameters:
  • rule (str) – Rule expression using X["col"] notation.

  • columns (list[str]) – Columns for which a not-null condition is trivially true and can be dropped, e.g. because the column is never actually null.

Returns:

Rule string with the not-null conditions removed.

Return type:

str

Examples

>>> drop_not_null_conditions('(X["a"] > 1) & (~X["b"].is_null())', ["b"])
'(X["a"] > 1)'

prettify_rules#

iguanas.rule_formatting.prettify_rules(rules: list[str], steps: list[collections.abc.Callable[[str], str]], column_name_mapping: dict[str, str] | None = None) → list[str][source]#

Apply an ordered list of rule-string transformations to each rule.

Each step is a plain function taking and returning a rule string (e.g. simplify_rule, or functools.partial(decode_numeric_encodings, mapping=woe_mapping)), applied in order.

Parameters:
  • rules (list[str]) – Raw rule strings to prettify.

  • steps (list[Callable[[str], str]]) – Ordered transformations to apply to each rule.

  • column_name_mapping (dict[str, str] | None, optional) – Maps column name to a display name, applied last, by default None.

Returns:

Prettified rule strings.

Return type:

list[str]

Examples

>>> from functools import partial
>>> steps = [
...     partial(decode_numeric_encodings, mapping={"A": {"x": 1, "y": 2}}),
...     partial(round_thresholds, columns=["amount"]),
... ]
>>> prettify_rules(['(X["A"] >= 2) & (X["amount"] > 1.239)'], steps)
['(X["A"] == "y") & (X["amount"] > 1.24)']