Rule Generation#

Supported backends: XGBClassifier, LGBMClassifier and scikit-learn’s RandomForestClassifier. All three are converted to the same canonical per-node tree table before rule extraction, so every function below works unchanged regardless of which fitted estimator is passed in.

Functions#

extract_max_gain_rule#

iguanas.rule_generation.extract_max_gain_rule(tree_X: pandas.DataFrame) → str[source]#

Extract the rule path to the leaf with maximum gain using bottom-to-top approach.

Finds the leaf node with highest gain value and traces back to the root node, building the rule by reconstructing conditions from child to parent.

Parameters:

tree_X (pd.DataFrame) – Output from estimator._Booster.trees_to_dataframe() filtered for a single tree. Required columns: Tree, Node, ID, Feature, Split, Yes, No, Missing, Gain, Cover.

Returns:

Rule string in format (X[“feat1”] >= Split1) & (X[“feat2”] < Split2). Returns empty string if tree is empty or has no valid leaves.

Return type:

str

extract_rule_with_monotone_constraints#

iguanas.rule_generation.extract_rule_with_monotone_constraints(tree_X: pandas.DataFrame, monotone_constraints: dict[str, int]) → str[source]#

Extract rule path following monotone constraints using top-to-bottom approach.

Starts from root and follows tree structure based on monotone constraints. NOTE: Only applicable if ALL features have a monotone constraint of -1 or +1. Features with constraint 0 will raise a ValueError.

Parameters:
  • tree_X (pd.DataFrame) – Output from estimator._Booster.trees_to_dataframe() filtered for a single tree. Required columns: Tree, Node, ID, Feature, Split, Yes, No, Missing.

  • monotone_constraints (dict[str, int]) –

    Dictionary mapping feature names to constraint values:

    • +1 (positive): follow “No” branch (feature >= threshold)

    • -1 (negative): follow “Yes” branch (feature < threshold)

    • 0 (none): raises ValueError - not supported

Returns:

Rule string in format (X[“feat1”] >= Split1) & (X[“feat2”] < Split2). Returns empty string if tree is empty or starts with a leaf.

Return type:

str

Raises:

ValueError – If a feature has no constraint defined or has constraint 0.

extract_rules#

iguanas.rule_generation.extract_rules(estimator: xgboost.sklearn.XGBClassifier, all_features_constrained: bool, leaf_selection: str = 'max_gain', **kwargs: Any) → pandas.DataFrame[source]#

Generate rules extracted from XGBoost, LightGBM or RandomForest trees.

Parameters:
  • estimator (XGBClassifier | LGBMClassifier | RandomForestClassifier) – Fitted tree-based classifier. XGBoost, LightGBM and scikit-learn’s RandomForestClassifier are supported.

  • all_features_constrained (bool) – If True, uses monotone constraint-based extraction (top-to-bottom), which always yields one rule per tree; leaf_selection is ignored. If False, leaf_selection picks the extraction strategy.

  • leaf_selection ({"max_gain", "all_positive"}, default="max_gain") –

    Strategy used when all_features_constrained is False:

    • "max_gain": one rule per tree, the single highest-gain leaf (see extract_max_gain_rule()).

    • "all_positive": zero or more rules per tree, one for every leaf whose value favours the positive class (see extract_positive_gain_rules()).

  • **kwargs (dict) – Additional metadata columns added to the output DataFrame (e.g., transformation name, scale_pos_weight value).

Returns:

DataFrame with columns: rule, tree, and any kwargs columns.

Return type:

pd.DataFrame

Raises:

ValueError – If leaf_selection is not "max_gain" or "all_positive", or if all_features_constrained is True but the estimator was not fitted with a non-zero monotone constraint for every feature.

rule_grid_search_sequential#

iguanas.rule_generation.rule_grid_search_sequential(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, verbose: int = 0, leaf_selection: str = 'max_gain') → polars.DataFrame[source]#

Sequential (single-threaded) variant of rule_grid_search.

Identical behaviour to rule_grid_search() but runs the grid in a plain loop without joblib. Useful for debugging, for deterministic profiling, or for small workloads where the thread-dispatch overhead outweighs the benefit of parallelism.

Parameters:
  • estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.

  • X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.

  • y_train (pl.Series | pd.Series) – Training target values.

  • scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try.

  • sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.

  • verbose (int, default=0) – Controls verbosity. 0 = silent, 1 = summary.

  • leaf_selection ({"max_gain", "all_positive"}, default="max_gain") – Forwarded to extract_rules() for every unconstrained tree; see that function for what each option does. Ignored for trees where every feature carries a monotone constraint.

Returns:

Same schema as rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.

Return type:

pl.DataFrame

rule_grid_search_parallel_weights#

iguanas.rule_generation.rule_grid_search_parallel_weights(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, n_jobs: int = -1, verbose: int = 0, leaf_selection: str = 'max_gain') → polars.DataFrame[source]#

Grid search over sample weight transformations and scale_pos_weight values, parallelised over weight transformations.

This function systematically trains XGBoost models with different combinations of: - sample weights - scale_pos_weight values

For each combination, it extracts rules from the fitted models and returns them as a Polars DataFrame. The results from all combinations are pooled and deduplicated; the grid is a mechanism for producing a diverse candidate set, not a search for a single “best” configuration, and no optimality over the space of rules is claimed. The weight-transformation loop is parallelised with joblib.Parallel using the "threading" backend — single-node, thread-parallel only.

Parameters:
  • estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.

  • X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.

  • y_train (pl.Series | pd.Series) – Training target values.

  • scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try. Iterated sequentially within each worker thread.

  • sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.

  • n_jobs (int, default=-1) – Number of joblib worker threads. -1 means one per available core.

  • verbose (int, default=0) –

    Controls the verbosity level:

    • 0: silent (no output)

    • 1: progress information (start/end summary)

    • >=2: detailed progress with live updates from joblib Parallel backend

Returns:

Same schema as rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.

Return type:

pl.DataFrame

Examples

>>> weights_train = generate_sample_weight_transformations(X_train["amount"])
>>> scale_pos_weights = np.logspace(0, np.log10(imbalance_ratio*2), 20)
>>> results = rule_grid_search(
...     estimator, X_train, y_train,
...     scale_weights, weights_train, n_jobs=-1, verbose=1
... )

rule_grid_search_parallel_scales#

iguanas.rule_generation.rule_grid_search_parallel_scales(estimator: xgboost.sklearn.XGBClassifier, X_train: polars.DataFrame | pandas.DataFrame, y_train: polars.Series | pandas.Series, scale_pos_weights: list[float] | numpy.ndarray, sample_weights_df: polars.DataFrame | pandas.DataFrame | None = None, n_jobs: int = -1, verbose: int = 0, leaf_selection: str = 'max_gain') → polars.DataFrame[source]#

Grid search parallelised over scale_pos_weight values.

This function systematically trains XGBoost models with different combinations of: - sample weights - scale_pos_weight values

For each combination, it extracts rules from the fitted models and returns them as a Polars DataFrame. The results from all combinations are pooled and deduplicated; the grid is a mechanism for producing a diverse candidate set, not a search for a single “best” configuration, and no optimality over the space of rules is claimed. The scale_pos_weight loop is parallelised with joblib.Parallel using the "threading" backend — single-node, thread-parallel only.

Parameters:
  • estimator (XGBClassifier) – Base XGBoost classifier to use as a template for rule extraction.

  • X_train (pl.DataFrame | pd.DataFrame) – Training feature matrix.

  • y_train (pl.Series | pd.Series) – Training target values.

  • scale_pos_weights (list | np.ndarray) – Array of scale_pos_weight values to try. Distributed across worker threads.

  • sample_weights_df (pl.DataFrame | pd.DataFrame | None, default=None) – DataFrame mapping transformation names to sample weight arrays. If None, uses baseline weights of 1.0 for all samples.

  • n_jobs (int, default=-1) – Number of joblib worker threads. -1 means one per available core.

  • verbose (int, default=0) –

    Controls the verbosity level:

    • 0: silent (no output)

    • 1: progress information (start/end summary)

    • >=2: detailed progress with live updates from joblib Parallel backend

Returns:

Same schema as rule_grid_search(): columns rule, tree, scale_pos_weight, transformation.

Return type:

pl.DataFrame