Iguanas is a rule generation and evaluation library built on top of Polars, designed to streamline the entire rule-based system development workflow — from raw data to production-ready rules.

Note

For data preprocessing and feature engineering prior to rule generation, we recommend using Gators — a complementary library built on top of Polars by the same team at PayPal, providing 70+ transformers for cleaning, encoding, imputation, scaling, and more.

Built by the PSP Data Team at PayPal, Iguanas makes rule generation, evaluation, and selection simpler, with a single-node execution model that uses multi-threading where it helps.

Key Features#

  • ⚙️ Vectorised evaluation: Rules are compiled once to Polars expressions (cached) and applied as columnar, multi-threaded operations

  • 🧵 Thread-parallel grid search: Grid search is parallelised on a single machine via joblib’s threading backend

  • 🎯 End-to-End: Generate, evaluate, combine, and select rules in one library

  • 📦 Production Ready: Lightweight rule strings that deploy anywhere, plus ONNX export for runtimes that must not execute Python

  • 🔧 Flexible: Sequential and thread-parallel grid search strategies

  • 🔗 Composable: Chain generation → evaluation → selection with a few function calls

  • 🎓 Easy to Learn: Simple functional API with clear, consistent signatures

Note

Scope of parallelism. Iguanas runs on a single node. Grid search uses joblib.Parallel with the "threading" backend, and rule evaluation relies on Polars’ internal multi-threading. There is no multiprocessing, no distributed or cluster execution (no Dask, Ray or Spark), and no GPU code path. No published benchmark accompanies this release, so no throughput or speed-up figure is claimed.

Quick Start#

import polars as pl
import numpy as np
from xgboost import XGBClassifier

from iguanas.weight_transformations import generate_weights
from iguanas.rule_generation import rule_grid_search_parallel_weights
from iguanas.rule_evaluation import apply_filter_and_deduplicate_rules

# 1. Load your data
X_train = pl.DataFrame({
    "age":    [25, 45, 35, 50, 30, 55, 40, 28],
    "income": [30000, 80000, 50000, 90000, 40000, 95000, 70000, 35000],
})
y_train = pl.Series([0, 1, 0, 1, 0, 1, 1, 0])

# 2. Generate sample weight transformations
weights = generate_weights(X_train["income"])

# 3. Run a parallel grid search to extract rules
estimator = XGBClassifier(max_depth=2, n_estimators=5, random_state=42)
scale_pos_weights = np.logspace(0, 1, 5)

rules_df = rule_grid_search_parallel_weights(
    estimator, X_train, y_train,
    scale_pos_weights=scale_pos_weights,
    sample_weights_df=weights,
    n_jobs=-1,
)

# 4. Evaluate, filter, and deduplicate rules
R, metrics, selected_rules = apply_filter_and_deduplicate_rules(
    X_train, y_train, rules_df,
    metric_thresholds=[
        {"name": "precision", "operator": ">=", "value": 0.6},
        {"name": "recall",    "operator": ">=", "value": 0.5},
    ],
    max_corr=0.8,
)

print(selected_rules)

What Can Iguanas Do?#

  • ⚙️ Rule Generation - Extract rules from XGBoost/LightGBM models with grid search

  • 📊 Metrics - Precision, recall, F-beta, MCC, and weighted variants

  • 🔍 Rule Evaluation - Evaluate, filter, and deduplicate rule sets; lazy scoring via apply_rules_lazy (compiles rules with eval() — trusted rule sources only)

  • 🔁 Rule Cross-Validation - Check rule stability across K folds (cv mean, std, min per metric; optimistically biased — see the module notes)

  • 🔀 Rule Combination - Combine rules with greedy, beam, and A* search

  • ✂️ Rule Selection - Prune by feature overlap and correlation

  • 🔬 Rule Analysis - Inspect and report on rule structure

  • 💬 Rule Explanation - Verbalize rules, compute coverage overlap, and generate counterfactual explanations

  • 🖊️ Rule Formatting - Simplify rules, export to SQL, and reverse encoded-feature rules (e.g. from gators) back to the original columns

  • ⚖️ Rule Fairness - Post-hoc bias measurement across demographic subgroups (measured, not optimised)

  • 🗂️ Rule Registry - Store named ruleset snapshots and filter pairs by Jaccard overlap

  • 📈 Rule Monitoring - Flag per-rule performance degradation between reference and current periods (threshold on metric deltas, not a statistical drift test)

  • 🤖 Classifiers - scikit-learn compatible RuleClassifier and RulesetClassifier

  • 📦 ONNX Export - Convert rules to portable ONNX models for any runtime (parses with ast; executes no Python at scoring time)

  • 📐 Monotone Constraints - Infer feature directionality

  • 🔧 Weight Transformations - Generate and select sample weight schedules

Use Cases#

Iguanas is intended for:

  • Fraud Detection — Generate high-precision rules to flag suspicious transactions

  • Risk Scoring — Build interpretable rule sets for credit or operational risk

  • Compliance & Policy — Encode business policies as auditable rule expressions

  • Anomaly Detection — Surface rare but meaningful patterns in labelled data

  • Model Explainability — Extract human-readable rules from gradient boosted models

Credits#

Developed by the PSP Data Team at PayPal.

⚡ Built by data scientists, for data scientists