Iguanas is a rule generation and evaluation library built on top of Polars, designed to streamline the entire rule-based system development workflow — from raw data to production-ready rules.
Note
For data preprocessing and feature engineering prior to rule generation, we recommend using Gators — a complementary library built on top of Polars by the same team at PayPal, providing 70+ transformers for cleaning, encoding, imputation, scaling, and more.
Built by the PSP Data Team at PayPal, Iguanas makes rule generation, evaluation, and selection simpler, with a single-node execution model that uses multi-threading where it helps.
Key Features#
⚙️ Vectorised evaluation: Rules are compiled once to Polars expressions (cached) and applied as columnar, multi-threaded operations
🧵 Thread-parallel grid search: Grid search is parallelised on a single machine via joblib’s threading backend
🎯 End-to-End: Generate, evaluate, combine, and select rules in one library
📦 Production Ready: Lightweight rule strings that deploy anywhere, plus ONNX export for runtimes that must not execute Python
🔧 Flexible: Sequential and thread-parallel grid search strategies
🔗 Composable: Chain generation → evaluation → selection with a few function calls
🎓 Easy to Learn: Simple functional API with clear, consistent signatures
Note
Scope of parallelism. Iguanas runs on a single node. Grid search uses
joblib.Parallel with the "threading" backend, and rule evaluation relies on
Polars’ internal multi-threading. There is no multiprocessing, no distributed or
cluster execution (no Dask, Ray or Spark), and no GPU code path. No published
benchmark accompanies this release, so no throughput or speed-up figure is claimed.
Quick Start#
import polars as pl
import numpy as np
from xgboost import XGBClassifier
from iguanas.weight_transformations import generate_weights
from iguanas.rule_generation import rule_grid_search_parallel_weights
from iguanas.rule_evaluation import apply_filter_and_deduplicate_rules
# 1. Load your data
X_train = pl.DataFrame({
"age": [25, 45, 35, 50, 30, 55, 40, 28],
"income": [30000, 80000, 50000, 90000, 40000, 95000, 70000, 35000],
})
y_train = pl.Series([0, 1, 0, 1, 0, 1, 1, 0])
# 2. Generate sample weight transformations
weights = generate_weights(X_train["income"])
# 3. Run a parallel grid search to extract rules
estimator = XGBClassifier(max_depth=2, n_estimators=5, random_state=42)
scale_pos_weights = np.logspace(0, 1, 5)
rules_df = rule_grid_search_parallel_weights(
estimator, X_train, y_train,
scale_pos_weights=scale_pos_weights,
sample_weights_df=weights,
n_jobs=-1,
)
# 4. Evaluate, filter, and deduplicate rules
R, metrics, selected_rules = apply_filter_and_deduplicate_rules(
X_train, y_train, rules_df,
metric_thresholds=[
{"name": "precision", "operator": ">=", "value": 0.6},
{"name": "recall", "operator": ">=", "value": 0.5},
],
max_corr=0.8,
)
print(selected_rules)
What Can Iguanas Do?#
⚙️ Rule Generation - Extract rules from XGBoost/LightGBM models with grid search
📊 Metrics - Precision, recall, F-beta, MCC, and weighted variants
🔍 Rule Evaluation - Evaluate, filter, and deduplicate rule sets; lazy scoring via
apply_rules_lazy(compiles rules witheval()— trusted rule sources only)🔁 Rule Cross-Validation - Check rule stability across K folds (cv mean, std, min per metric; optimistically biased — see the module notes)
🔀 Rule Combination - Combine rules with greedy, beam, and A* search
✂️ Rule Selection - Prune by feature overlap and correlation
🔬 Rule Analysis - Inspect and report on rule structure
💬 Rule Explanation - Verbalize rules, compute coverage overlap, and generate counterfactual explanations
🖊️ Rule Formatting - Simplify rules, export to SQL, and reverse encoded-feature rules (e.g. from gators) back to the original columns
⚖️ Rule Fairness - Post-hoc bias measurement across demographic subgroups (measured, not optimised)
🗂️ Rule Registry - Store named ruleset snapshots and filter pairs by Jaccard overlap
📈 Rule Monitoring - Flag per-rule performance degradation between reference and current periods (threshold on metric deltas, not a statistical drift test)
🤖 Classifiers - scikit-learn compatible
RuleClassifierandRulesetClassifier📦 ONNX Export - Convert rules to portable ONNX models for any runtime (parses with
ast; executes no Python at scoring time)📐 Monotone Constraints - Infer feature directionality
🔧 Weight Transformations - Generate and select sample weight schedules
Use Cases#
Iguanas is intended for:
Fraud Detection — Generate high-precision rules to flag suspicious transactions
Risk Scoring — Build interpretable rule sets for credit or operational risk
Compliance & Policy — Encode business policies as auditable rule expressions
Anomaly Detection — Surface rare but meaningful patterns in labelled data
Model Explainability — Extract human-readable rules from gradient boosted models
Credits#
Developed by the PSP Data Team at PayPal.
⚡ Built by data scientists, for data scientists