Rule Generation#
Supported backends: XGBClassifier, LGBMClassifier and scikit-learn’s
RandomForestClassifier. All three are converted to the same canonical
per-node tree table before rule extraction, so every function below works
unchanged regardless of which fitted estimator is passed in.
Functions#
extract_max_gain_rule#
- iguanas.rule_generation.extract_max_gain_rule(tree_X: pandas.DataFrame) str[source]#
Extract the rule path to the leaf with maximum gain using bottom-to-top approach.
Finds the leaf node with highest gain value and traces back to the root node, building the rule by reconstructing conditions from child to parent.
- Parameters:
tree_X (pd.DataFrame) – Output from estimator._Booster.trees_to_dataframe() filtered for a single tree. Required columns: Tree, Node, ID, Feature, Split, Yes, No, Missing, Gain, Cover.
- Returns:
Rule string in format (X[“feat1”] >= Split1) & (X[“feat2”] < Split2). Returns empty string if tree is empty or has no valid leaves.
- Return type:
str
extract_rule_with_monotone_constraints#
- iguanas.rule_generation.extract_rule_with_monotone_constraints(tree_X: pandas.DataFrame, monotone_constraints: dict[str, int]) str[source]#
Extract rule path following monotone constraints using top-to-bottom approach.
Starts from root and follows tree structure based on monotone constraints. NOTE: Only applicable if ALL features have a monotone constraint of -1 or +1. Features with constraint 0 will raise a ValueError.
- Parameters:
tree_X (pd.DataFrame) – Output from estimator._Booster.trees_to_dataframe() filtered for a single tree. Required columns: Tree, Node, ID, Feature, Split, Yes, No, Missing.
monotone_constraints (dict[str, int]) –
Dictionary mapping feature names to constraint values:
+1 (positive): follow “No” branch (feature >= threshold)
-1 (negative): follow “Yes” branch (feature < threshold)
0 (none): raises ValueError - not supported
- Returns:
Rule string in format (X[“feat1”] >= Split1) & (X[“feat2”] < Split2). Returns empty string if tree is empty or starts with a leaf.
- Return type:
str
- Raises:
ValueError – If a feature has no constraint defined or has constraint 0.
extract_rules#
- iguanas.rule_generation.extract_rules(estimator: xgboost.sklearn.XGBClassifier, all_features_constrained: bool, leaf_selection: str = 'max_gain', **kwargs: Any) pandas.DataFrame[source]#
Generate rules extracted from XGBoost, LightGBM or RandomForest trees.
- Parameters:
estimator (XGBClassifier | LGBMClassifier | RandomForestClassifier) – Fitted tree-based classifier. XGBoost, LightGBM and scikit-learn’s RandomForestClassifier are supported.
all_features_constrained (bool) – If True, uses monotone constraint-based extraction (top-to-bottom), which always yields one rule per tree;
leaf_selectionis ignored. If False,leaf_selectionpicks the extraction strategy.leaf_selection ({"max_gain", "all_positive"}, default="max_gain") –
Strategy used when
all_features_constrainedis False:"max_gain": one rule per tree, the single highest-gain leaf (seeextract_max_gain_rule())."all_positive": zero or more rules per tree, one for every leaf whose value favours the positive class (seeextract_positive_gain_rules()).
**kwargs (dict) – Additional metadata columns added to the output DataFrame (e.g., transformation name, scale_pos_weight value).
- Returns:
DataFrame with columns:
rule,tree, and anykwargscolumns.- Return type:
pd.DataFrame
- Raises:
ValueError – If
leaf_selectionis not"max_gain"or"all_positive", or ifall_features_constrainedis True but the estimator was not fitted with a non-zero monotone constraint for every feature.
rule_grid_search#
- iguanas.rule_generation.rule_grid_search(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, n_jobs: int = -1, verbose: int = 0, leaf_selection: str = 'max_gain') polars.DataFrame[source]#
Grid search over scale_pos_weight values and sample weight transformations.
Dispatches to
rule_grid_search_parallel_scales()orrule_grid_search_parallel_weights()depending on which axis of the grid is larger, so that the parallelised loop is the longer one.This function systematically trains XGBoost models with different combinations of: - sample weights - scale_pos_weight values
For each combination, it extracts rules from the fitted models and returns them as a Polars DataFrame. The results from all combinations are pooled and deduplicated; the grid is a mechanism for producing a diverse candidate set, not a search for a single “best” configuration, and no optimality over the space of rules is claimed. Parallelism is
joblib.Parallelwith the"threading"backend — single-node, thread-parallel only.- Parameters:
estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.
X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.
y_train (pl.Series | pd.Series) – Training target values.
scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try.
sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.
n_jobs (int, default=-1) – Number of joblib worker threads. -1 means one per available core.
verbose (int, default=0) –
Controls the verbosity level:
0: silent (no output)
1: progress information (start/end summary)
>=2: detailed progress with live updates from joblib Parallel backend
leaf_selection ({"max_gain", "all_positive"}, default="max_gain") – Forwarded to
extract_rules()for every unconstrained tree.
- Returns:
Same schema as
rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.- Return type:
pl.DataFrame
rule_grid_search_sequential#
- iguanas.rule_generation.rule_grid_search_sequential(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, verbose: int = 0, leaf_selection: str = 'max_gain') polars.DataFrame[source]#
Sequential (single-threaded) variant of rule_grid_search.
Identical behaviour to
rule_grid_search()but runs the grid in a plain loop without joblib. Useful for debugging, for deterministic profiling, or for small workloads where the thread-dispatch overhead outweighs the benefit of parallelism.- Parameters:
estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.
X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.
y_train (pl.Series | pd.Series) – Training target values.
scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try.
sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.
verbose (int, default=0) – Controls verbosity. 0 = silent, 1 = summary.
leaf_selection ({"max_gain", "all_positive"}, default="max_gain") – Forwarded to
extract_rules()for every unconstrained tree; see that function for what each option does. Ignored for trees where every feature carries a monotone constraint.
- Returns:
Same schema as
rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.- Return type:
pl.DataFrame
rule_grid_search_parallel_weights#
- iguanas.rule_generation.rule_grid_search_parallel_weights(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, n_jobs: int = -1, verbose: int = 0, leaf_selection: str = 'max_gain') polars.DataFrame[source]#
Grid search over sample weight transformations and scale_pos_weight values, parallelised over weight transformations.
This function systematically trains XGBoost models with different combinations of: - sample weights - scale_pos_weight values
For each combination, it extracts rules from the fitted models and returns them as a Polars DataFrame. The results from all combinations are pooled and deduplicated; the grid is a mechanism for producing a diverse candidate set, not a search for a single “best” configuration, and no optimality over the space of rules is claimed. The weight-transformation loop is parallelised with
joblib.Parallelusing the"threading"backend — single-node, thread-parallel only.- Parameters:
estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.
X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.
y_train (pl.Series | pd.Series) – Training target values.
scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try. Iterated sequentially within each worker thread.
sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.
n_jobs (int, default=-1) – Number of joblib worker threads. -1 means one per available core.
verbose (int, default=0) –
Controls the verbosity level:
0: silent (no output)
1: progress information (start/end summary)
>=2: detailed progress with live updates from joblib Parallel backend
- Returns:
Same schema as
rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.- Return type:
pl.DataFrame
Examples
>>> weights_train = generate_sample_weight_transformations(X_train["amount"]) >>> scale_pos_weights = np.logspace(0, np.log10(imbalance_ratio*2), 20) >>> results = rule_grid_search( ... estimator, X_train, y_train, ... scale_weights, weights_train, n_jobs=-1, verbose=1 ... )
rule_grid_search_parallel_scales#
- iguanas.rule_generation.rule_grid_search_parallel_scales(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, n_jobs: int = -1, verbose: int = 0, leaf_selection: str = 'max_gain') polars.DataFrame[source]#
Grid search parallelised over scale_pos_weight values.
This function systematically trains XGBoost models with different combinations of: - sample weights - scale_pos_weight values
For each combination, it extracts rules from the fitted models and returns them as a Polars DataFrame. The results from all combinations are pooled and deduplicated; the grid is a mechanism for producing a diverse candidate set, not a search for a single “best” configuration, and no optimality over the space of rules is claimed. The scale_pos_weight loop is parallelised with
joblib.Parallelusing the"threading"backend — single-node, thread-parallel only.- Parameters:
estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.
X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.
y_train (pl.Series | pd.Series) – Training target values.
scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try. Distributed across worker threads.
sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.
n_jobs (int, default=-1) – Number of joblib worker threads. -1 means one per available core.
verbose (int, default=0) –
Controls the verbosity level:
0: silent (no output)
1: progress information (start/end summary)
>=2: detailed progress with live updates from joblib Parallel backend
- Returns:
Same schema as
rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.- Return type:
pl.DataFrame