Quick Start#
This guide will get you started with Iguanas in minutes.
Basic Example#
Here’s a simple example showing the core Iguanas workflow:
import polars as pl
import numpy as np
from xgboost import XGBClassifier
# Import Iguanas modules
from iguanas.rule_generation import rule_grid_search_parallel_scales
from iguanas.rule_evaluation import apply_rules
from iguanas.rule_selection import filter_correlated_rules
from iguanas.metrics import compute_metrics
# 1. Load your data (example with synthetic data)
X_train = pl.DataFrame({
'age': [25, 45, 35, 50, 30, 55, 40, 28],
'income': [30000, 80000, 50000, 90000, 40000, 95000, 70000, 35000],
'credit_score': [650, 720, 680, 750, 660, 780, 710, 640]
})
y_train = pl.Series([0, 1, 0, 1, 0, 1, 1, 0])
# 3. Configure XGBoost estimator for rule extraction
estimator = XGBClassifier(
max_depth=1, # Decision stumps for simple rules
n_estimators=10,
random_state=42
)
# 4. Generate rules using grid search
scale_pos_weights = np.logspace(0, 1, 5) # Try different class balance weights
rules_df = rule_grid_search_parallel_scales(
estimator=estimator,
X_train=X_train,
y_train=y_train,
scale_pos_weights=scale_pos_weights,
n_jobs=-1,
verbose=1
)
# 5. Apply rules to your data
rules = rules_df['rule'].unique().to_list()
R_train = apply_rules(X_train, rules)
# 6. Compute performance metrics
metrics = compute_metrics(R_train, y_train)
print(metrics.select(['rule', 'precision', 'recall', 'f1']).head(10))
# 7. Filter correlated rules to keep only diverse, high-performing rules
importance = dict(metrics[['rule', 'f1']].rows())
uncorrelated_rules = filter_correlated_rules(R_train, importance, max_corr=0.8)
print(f"Original rules: {len(rules)}")
print(f"Filtered rules: {len(uncorrelated_rules)}")
Understanding the API#
Iguanas is organized into modular components that work together in a typical workflow:
- 1. Rule Generation (Rule Generation)
Generate rules from your data using XGBoost, LightGBM or RandomForest decision trees:
rule_grid_search(): Thread-parallel grid search over weight transformations and scale_pos_weight values (single node; joblib"threading"backend)extract_rules(): Extract rules from a fitted XGBoost/LightGBM/RandomForest modelextract_max_gain_rule(): Extract the highest-gain rule path from a single tree
- 2. Rule Evaluation (Rule Evaluation)
Apply rules to data and evaluate their performance:
apply_rules(): Evaluate rule expressions on DataFrames (compiles rules witheval()— use trusted rule sources only)apply_and_filter_by_performance(): Filter rules by precision/recall thresholdsselect_diverse_top_rules(): Select top performing non-correlated rules
- 3. Metrics (Metrics)
Compute comprehensive performance metrics:
compute_metrics(): Calculate precision, recall, F-scores, TPVE metricsSupports both count-based and weighted metrics
- 4. Rule Selection (Rule Selection)
Filter and select rules based on similarity and correlation:
filter_correlated_rules(): Remove highly correlated rulesfilter_rules_by_feature_overlap(): Filter rules with similar feature usageextract_feature_names_from_rule(): Extract features used in rules
- 5. Rule Combination (Rule Combination)
Combine rules to create more powerful composite rules:
combine_rules_full_search(): Generate all combinationscombine_rules_greedy(): Greedy search for best combinationscombine_rules_beam_search(): Beam search algorithmcombine_rules_a_star(): A* search algorithm
- 6. Rule Analysis (Rule Analysis)
Analyze rules at hierarchical levels:
generate_rule_performance_report(): Metrics at rule, component, and condition levels
- 7. Rule Formatting (Rule Formatting)
Transform and simplify rules, and reverse encoded-feature rules (e.g. from a gators preprocessing pipeline) back to the original columns:
simplify_rule(): Remove redundant conditions (e.g. duplicate bounds on the same feature)rule_to_sql(): Convert a rule expression to a SQLWHEREclause with an optional table aliasdecode_numeric_encodings(),decode_onehot_encodings(),decode_string_imputation(),decode_discretized_bins(),decode_scaled_thresholds(),decode_null_indicators(): Reverse encoder/imputer/discretizer/scaler outputs back to conditions on the original columnsadd_missing_value_conditions(),format_floats_as_integers(),format_as_boolean_conditions(),quote_string_values(),round_thresholds(): Normalise condition values for readabilitydrop_null_clauses(),drop_not_null_conditions(): Remove redundant is-null clausesprettify_rules(): Chain any of the above into a single rule-cleanup pipeline
- 8. Utilities
Supporting utilities for rule generation:
Weight Transformations: Generate and select sample weight schedules (
generate_weights,select_uncorrelated_weights)Monotone Constraints: Infer monotone constraints for XGBoost/LightGBM/RandomForest
- 9. Rule Cross-Validation (Rule Cross-Validation)
Check rule stability across folds:
validate_rules_cv(): Evaluate rules on K folds and return per-metric mean, std, and minidentify_unstable_rules(): Flag rules with high cross-fold variance
Warning
Rules are generated on the full dataset before being passed to
validate_rules_cv(), so the folds are not truly held out and the reported cv std/min are optimistically biased. Treat them as a relative screen for fragile rules, not as an out-of-sample estimate; use a nested protocol (full pipeline inside each outer training split) for unbiased numbers.- 10. Rule Explanation (Rule Explanation)
Interpret and explain individual rule predictions:
verbalize_rule(): Convert a rule expression to plain Englishcompute_coverage_overlap(): Compute pairwise Jaccard overlap between rule predictionscompute_counterfactual(): Find the minimal feature changes to un-flag a sample
- 11. Rule Fairness (Rule Fairness)
Post-hoc bias measurement across demographic subgroups. Fairness is measured, not optimised — nothing here changes how rules are generated or selected:
compute_subgroup_metrics(): Per-subgroup precision, recall, and all other metricscompute_disparate_impact_ratio(): Surface disparate impact relative to a reference group (screening heuristic, not a statistical test)
- 12. Rule Monitoring (Rule Monitoring)
Flag per-rule performance degradation between two labelled periods. This is a threshold on metric deltas, not statistical drift detection (no KS test, PSI, or JS divergence):
compare_rule_metrics(): Compare per-rule metrics between a reference and current period; flag degraded rules
- 13. Rule Registry (Rule Registry)
Store and compare named ruleset snapshots across experiments (keyed by name; no revision history):
RuleRegistry: Save, load, delete, and compare named rule snapshots (JSON persistence)filter_rule_pairs_by_overlap(): Filter rule pairs by Jaccard overlap range
- 14. Classifiers (Rule Classifier)
scikit-learn compatible end-to-end pipelines:
RuleClassifier: Fit → generate → filter → select single best ruleRulesetClassifier: Fit → generate → filter → correlate → greedily combine rules
- 15. ONNX Export (ONNX Converter)
Deploy rules to any ONNX-compatible runtime. Rules are parsed with
astand emitted as a static graph, so scoring executes no Python — the recommended production path:rules_to_onnx(): Convert rule strings to a portable ONNX binary classifier
Next Steps#
Explore Rule Generation for generating rules from your data
Check out Rule Evaluation for applying and filtering rules
Check out Metrics for evaluating rule performance
Learn about Rule Combination for combining rules
Dive into Rule Selection for deduplicating rule sets
See Rule Classifier for the scikit-learn compatible pipeline
See ONNX Converter for exporting rules to ONNX
Browse the API Reference for complete API documentation