Rule Formatting#
Functions#
simplify_rule#
- iguanas.rule_formatting.simplify_rule(rule: str) str[source]#
Simplify a rule by removing redundant conditions on the same column.
When multiple conditions exist on the same column, keeps only the most restrictive:
For lower bounds (>, >=): keeps the highest threshold, preferring > over >= when equal
For upper bounds (<, <=): keeps the lowest threshold, preferring < over <= when equal
- Parameters:
rule (str) – Rule string with conditions like (X[“col”] > val) & (X[“col”] >= val).
- Returns:
Simplified rule string with redundant conditions removed. Column order is preserved based on first appearance.
- Return type:
str
Examples
>>> simplify_rule('(X["amount"] >= 100.0) & (X["amount"] > 100.0)') '(X["amount"] > 100.0)'
>>> simplify_rule('(X["amount"] < 100.0) & (X["amount"] <= 100.0)') '(X["amount"] < 100.0)'
>>> simplify_rule('(X["a"] >= 50) & (X["b"] < 10) & (X["a"] > 100)') '(X["a"] > 100) & (X["b"] < 10)'
rule_to_sql#
- iguanas.rule_formatting.rule_to_sql(rule: str, table_alias: str | None = None) str[source]#
Convert a rule expression string to a SQL WHERE clause.
Translates Iguanas rule notation (
X["col"] op value) into standard SQL predicate syntax suitable for use in aWHEREorCASE WHENclause.- Parameters:
rule (str) – Rule expression using
X["col"]notation with&/|operators, e.g.'(X["age"] > 30) & (X["income"] < 50000)'.table_alias (str | None, default=None) – Optional table or CTE alias to prefix column references with. For example,
table_alias="t"turnsage > 30intot.age > 30.
- Returns:
SQL WHERE clause string.
- Return type:
str
Examples
>>> rule_to_sql('(X["age"] > 30) & (X["income"] < 50000)') '(age > 30.0) AND (income < 50000.0)'
>>> rule_to_sql('(X["age"] > 30) | (X["flag"] == 1)', table_alias="t") '(t.age > 30.0) OR (t.flag = 1.0)'
format_floats_as_integers#
- iguanas.rule_formatting.format_floats_as_integers(rule: str, int_columns: list[str]) str[source]#
Convert float thresholds to integers for the given columns.
Uses ceiling for
>=/<and floor for>/<=, so the integer boundary preserves the original condition’s semantics.==/!=are left unchanged since there is no boundary to round.- Parameters:
rule (str) – Rule expression using
X["col"]notation.int_columns (list[str]) – Columns whose float thresholds should be converted to integers.
- Returns:
Rule string with integer thresholds for the given columns.
- Return type:
str
Examples
>>> format_floats_as_integers('(X["a"] >= 0.1) & (X["b"] >= 9.1)', ["a"]) '(X["a"] >= 1) & (X["b"] >= 9.1)'
add_missing_value_conditions#
- iguanas.rule_formatting.add_missing_value_conditions(rule: str, mapping: dict[str, float]) str[source]#
Append an
is_null()clause to conditions satisfied by an imputed value.When a column’s nulls were filled with a value that also satisfies an existing condition, that condition implicitly matches originally-null rows too. This makes that explicit by OR-ing in
X[col].is_null().- Parameters:
rule (str) – Rule expression using
X["col"]notation.mapping (dict[str, float]) – Maps column name to the value nulls were imputed with.
- Returns:
Rule string with
is_null()clauses added where relevant.- Return type:
str
Examples
>>> add_missing_value_conditions('(X["a"] < 1)', {"a": 0}) '((X["a"] < 1) | X["a"].is_null())'
decode_string_imputation#
- iguanas.rule_formatting.decode_string_imputation(rule: str, mapping: dict[str, str]) str[source]#
Convert an equality on a string-imputed placeholder into
is_null().For columns where nulls were filled with a placeholder string (e.g. gators
StringImputer’s default"MISSING"), rewrites an equality condition on that placeholder back to an is-null check.- Parameters:
rule (str) – Rule expression using
X["col"]notation.mapping (dict[str, str]) – Maps column name to the placeholder value strings were imputed with.
- Returns:
Rule string with placeholder equality conditions decoded.
- Return type:
str
Examples
>>> decode_string_imputation('(X["status"] == "MISSING")', {"status": "MISSING"}) 'X["status"].is_null()'
decode_numeric_encodings#
- iguanas.rule_formatting.decode_numeric_encodings(rule: str, mapping: dict[str, dict[str, float]]) str[source]#
Reverse a numeric category encoding back to the original category labels.
For encodings where each category maps to a numeric statistic (e.g. WOE score, category count, mean target value), finds which categories satisfy the condition’s operator/threshold and replaces it with an equality (single match) or
.is_in()(multiple matches) condition on the original values.- Parameters:
rule (str) – Rule expression using
X["col"]notation, where col holds the encoded numeric values.mapping (dict[str, dict[str, float]]) – Maps column name to a dict of {category: encoded_value}.
- Returns:
Rule string with conditions decoded back to category labels.
- Return type:
str
Examples
>>> mapping = {"A": {"a": 1, "b": 2, "c": 3}} >>> decode_numeric_encodings('(X["A"] >= 2)', mapping) '(X["A"].is_in(["b", "c"]))'
format_as_boolean_conditions#
- iguanas.rule_formatting.format_as_boolean_conditions(rule: str, bool_columns: list[str]) str[source]#
Convert True/False-like condition values to Python booleans.
Recognises two condition shapes for the given columns:
“True”/”true”/”1” and “False”/”false”/”0” (quoted or bare), combined with
==/!=.Raw numeric threshold splits on
>=/>/</<=(as produced by a model trained on a boolean column cast to float), where the threshold unambiguously selects only the 0.0 or only the 1.0 value.
Both shapes are rewritten as
== True/== False.- Parameters:
rule (str) – Rule expression using
X["col"]notation.bool_columns (list[str]) – Columns to treat as boolean.
- Returns:
Rule string with boolean conditions normalised.
- Return type:
str
Examples
>>> format_as_boolean_conditions('(X["flag"] != "False")', ["flag"]) '(X["flag"] == True)'
>>> format_as_boolean_conditions('(X["flag"] >= 1.0)', ["flag"]) '(X["flag"] == True)'
>>> format_as_boolean_conditions('(X["flag"] < 1.0)', ["flag"]) '(X["flag"] == False)'
decode_onehot_encodings#
- iguanas.rule_formatting.decode_onehot_encodings(rule: str, mapping: dict[str, tuple[str, str]], null_category: str | None = None) str[source]#
Reverse a one-hot encoding back to a categorical condition.
One-hot encoders typically produce a binary column per category, split by the model at 0.5. This converts those splits back to equality/ inequality conditions on the original categorical column. If one category represents “value was null” (
null_category), it’s rendered asis_null()/~is_null()instead of a literal category comparison.- Parameters:
rule (str) – Rule expression using
X["col"]notation, where col is the one-hot encoded binary column.mapping (dict[str, tuple[str, str]]) – Maps encoded column name to (original_col, category).
null_category (str | None, optional) – Category value that represents an originally-null value, by default None.
- Returns:
Rule string with one-hot conditions decoded back to category labels.
- Return type:
str
Examples
>>> mapping = {"status__active": ("status", "active")} >>> decode_onehot_encodings('(X["status__active"] >= 0.5)', mapping) '(X["status"] == "active")'
decode_null_indicators#
- iguanas.rule_formatting.decode_null_indicators(rule: str, mapping: dict[str, str]) str[source]#
Convert null-indicator binary columns to
is_null()conditions.- Parameters:
rule (str) – Rule expression using
X["col"]notation, where col is a binary column indicating whether the original column was null.mapping (dict[str, str]) – Maps encoded column name to the original column name.
- Returns:
Rule string with is-null conditions decoded.
- Return type:
str
Examples
>>> decode_null_indicators('(X["amount__is_null"] >= 0.5)', {"amount__is_null": "amount"}) 'X["amount"].is_null()'
decode_discretized_bins#
- iguanas.rule_formatting.decode_discretized_bins(rule: str, mapping: dict[str, list[float]]) str[source]#
Reverse a discretizer’s bin index back to a threshold on the original column.
Discretizers (equal-width, equal-frequency, quantile, k-means, tree-based, geometric, custom bin edges, …) all replace a numeric column with an integer bin index. Given the fitted bin edges, this converts a condition on the bin index back to a condition on the original numeric column.
- Parameters:
rule (str) – Rule expression using
X["col"]notation, where col holds the bin index.mapping (dict[str, list[float]]) – Maps column name to its sorted bin edges, where
edges[i]is the lower boundary of bini.
- Returns:
Rule string with bin-index conditions decoded to original thresholds.
- Return type:
str
Examples
>>> decode_discretized_bins('(X["amount"] >= 2)', {"amount": [0, 10, 50, 200]}) '(X["amount"] >= 50)'
decode_scaled_thresholds#
- iguanas.rule_formatting.decode_scaled_thresholds(rule: str, mapping: dict[str, collections.abc.Callable[[float], float]]) str[source]#
Reverse a monotonic numeric scaling back to a threshold on the original column.
Works for any strictly increasing scaler (standardisation, min-max, log1p, Box-Cox, Yeo-Johnson, arcsinh, power, robust, …) since the
>=/</etc. ordering is preserved - pass the scaler’s inverse transform as the mapping value.- Parameters:
rule (str) – Rule expression using
X["col"]notation, where col holds the scaled values.mapping (dict[str, Callable[[float], float]]) – Maps column name to a function that inverts the scaling (e.g.
scaler.inverse_transform).
- Returns:
Rule string with scaled thresholds decoded back to original values.
- Return type:
str
Examples
>>> decode_scaled_thresholds('(X["amount"] >= 5.0)', {"amount": lambda x: x * 2}) '(X["amount"] >= 10.0)'
quote_string_values#
- iguanas.rule_formatting.quote_string_values(rule: str, columns: list[str]) str[source]#
Wrap bare (unquoted) condition values in double quotes.
- Parameters:
rule (str) – Rule expression using
X["col"]notation.columns (list[str]) – Columns whose values should be quoted.
- Returns:
Rule string with bare values quoted.
- Return type:
str
Examples
>>> quote_string_values('(X["col"] == retail)', ["col"]) '(X["col"] == "retail")'
round_thresholds#
- iguanas.rule_formatting.round_thresholds(rule: str, columns: list[str], ndigits: int = 2) str[source]#
Round numeric thresholds to a fixed number of decimal places.
- Parameters:
rule (str) – Rule expression using
X["col"]notation.columns (list[str]) – Columns to round.
ndigits (int, optional) – Decimal places, by default 2.
- Returns:
Rule string with rounded thresholds.
- Return type:
str
Examples
>>> round_thresholds('(X["amount"] >= 1234.56789)', ["amount"]) '(X["amount"] >= 1234.57)'
drop_null_clauses#
- iguanas.rule_formatting.drop_null_clauses(rule: str, columns: list[str]) str[source]#
Strip
| X[col].is_null()clauses added for always-imputed columns.- Parameters:
rule (str) – Rule expression using
X["col"]notation.columns (list[str]) – Columns for which to strip the is-null clause, e.g. because the column is never actually null once imputed.
- Returns:
Rule string with the is-null clauses removed.
- Return type:
str
Examples
>>> drop_null_clauses('((X["amount"] >= 5.0) | X["amount"].is_null())', ["amount"]) '(X["amount"] >= 5.0)'
drop_not_null_conditions#
- iguanas.rule_formatting.drop_not_null_conditions(rule: str, columns: list[str]) str[source]#
Drop standalone
(~X[col].is_null())conditions for given columns.- Parameters:
rule (str) – Rule expression using
X["col"]notation.columns (list[str]) – Columns for which a not-null condition is trivially true and can be dropped, e.g. because the column is never actually null.
- Returns:
Rule string with the not-null conditions removed.
- Return type:
str
Examples
>>> drop_not_null_conditions('(X["a"] > 1) & (~X["b"].is_null())', ["b"]) '(X["a"] > 1)'
prettify_rules#
- iguanas.rule_formatting.prettify_rules(rules: list[str], steps: list[collections.abc.Callable[[str], str]], column_name_mapping: dict[str, str] | None = None) list[str][source]#
Apply an ordered list of rule-string transformations to each rule.
Each step is a plain function taking and returning a rule string (e.g.
simplify_rule, orfunctools.partial(decode_numeric_encodings, mapping=woe_mapping)), applied in order.- Parameters:
rules (list[str]) – Raw rule strings to prettify.
steps (list[Callable[[str], str]]) – Ordered transformations to apply to each rule.
column_name_mapping (dict[str, str] | None, optional) – Maps column name to a display name, applied last, by default None.
- Returns:
Prettified rule strings.
- Return type:
list[str]
Examples
>>> from functools import partial >>> steps = [ ... partial(decode_numeric_encodings, mapping={"A": {"x": 1, "y": 2}}), ... partial(round_thresholds, columns=["amount"]), ... ] >>> prettify_rules(['(X["A"] >= 2) & (X["amount"] > 1.239)'], steps) ['(X["A"] == "y") & (X["amount"] > 1.24)']