Skip to content

Repository files navigation

PhilanthroPy logo

Rank your donors by who is most likely to make a major gift — leakage-safe scikit-learn models for nonprofit and hospital fundraising.

PyPI version Python versions Tests Coverage at least 92 percent sklearn compatible documentation License

🚀 View the Full Documentation Site


What is PhilanthroPy?

PhilanthroPy is a Python library that slots directly into sklearn.pipeline.Pipeline. It covers the full predictive workflow for nonprofit and academic medical center (AMC) fundraising — from raw CRM cleaning and wealth imputation to major-gift propensity scoring, lapse prediction, and planned-giving intent.

Who it's for

One toolkit, two audiences:

  • General nonprofit & university advancement teams (no PHI in scope) — CRM cleaning, RFM segmentation, wealth-screening imputation, and donor-propensity / lapse / planned-giving scoring. Start with CRMCleaner, RFMTransformer, WealthScreeningImputer, and DonorPropensityModel.
  • Academic medical center (AMC) foundations running grateful-patient programs (PHI in scope, higher scrutiny) — clinical-encounter featurization via EncounterTransformer, GratefulPatientFeaturizer, and DischargeToSolicitationWindowTransformer. Before production use, read Compliance Considerations: the PII handling here is a name-based heuristic, not formal HIPAA de-identification.

Maturity

Single-maintainer MIT project. pip install philanthropy gives you 0.6.0, the current release.

main is ahead of it: 0.7.0 (removes the 0.6.0 deprecations) and 1.0.0 (freezes the API) are merged and green but deliberately unreleased, so the 0.6.0 deprecation warnings get a real migration window rather than a token one. Read the CHANGELOG for what is queued.

Preprocessing and the core classifiers are Tier 1; grateful-patient featurization and philanthropy.ingest are Tier 2 (Beta); FinancialForecastModel and philanthropy.experimental.* are Tier 3 (Experimental) and carry no API guarantees. From 1.0.0, Tier 1 becomes semver-protected — breaking one requires a major release preceded by a full published minor of DeprecationWarning. Per-symbol tiers are in the API reference.

Maintenance: maintained by one person on a best-effort basis. For vendor / OSS risk reviews: the bus factor is 1.


Installation

pip install philanthropy
From source (for development)
git clone https://github.com/PhilanthroPy-Project/PhilanthroPy.git
cd PhilanthroPy
pip install -e ".[dev]"

Quick Start

from philanthropy.datasets import generate_synthetic_donor_data
from philanthropy.models import DonorPropensityModel

df = generate_synthetic_donor_data(n_samples=500, random_state=42)
X = df[["total_gift_amount", "years_active", "event_attendance_count"]].to_numpy()

model = DonorPropensityModel(n_estimators=200, random_state=0)
model.fit(X, df["is_major_donor"].to_numpy())
df["affinity_score"] = model.predict_affinity_score(X)   # 0–100, not a raw probability

print(df.groupby("is_major_donor")["affinity_score"].describe()[["count", "mean", "min", "max"]])
                count       mean   min    max
is_major_donor
0               165.0   9.636364   0.0   39.0
1               335.0  94.865672  65.0  100.0

Non-major donors top out at 39; no major donor scores below 65. That gap is the whole product: a gift-officer call list, sorted.

Try it now — zero install: Open In Colab

Runnable scripts: examples/quickstart.py and examples/unischema_to_scores.py run end to end and are smoke-tested in CI.

Affinity score distribution separating major from non-major donors
Output of plot_affinity_distribution(): the 0–100 affinity scores cleanly separate major from non-major donors.

No Python? Use the CLI

pip install philanthropy also puts a philanthropy command on your PATH. CSV in, scored CSV out, no Python file to write.

philanthropy train --data gifts.csv --target is_major_donor \
  --features total_gift_amount,years_active,event_attendance_count \
  --out model.joblib

philanthropy score --data prospects.csv --model model.joblib --out scored.csv

philanthropy validate reports precision/recall/F1/ROC-AUC on a labelled CSV — point it at a holdout year, not the year you trained on. Full walkthrough: Use the CLI.

Your data never leaves your machine

PhilanthroPy makes no network calls. There is no telemetry, no license check, no phone-home, and no third-party data append — the package imports no HTTP client at all, and tests/test_no_network.py enforces that in CI by making every socket raise. It models only what is already in your database. See Compliance considerations and the security review Q&A for the questions an institutional review will ask.


From your CRM to scores

philanthropy.ingest is the on-ramp: it turns what your donor system already emits into the donor-level feature table the estimators expect, with no glue code in between.

CiviCRM. A contribution export (or an APIv4 Contribution.get result) → read_civicrm_contributions()civicrm_contributions_to_features()predict_affinity_score(). The bridge drops payment-processor test transactions and counts only Completed contributions, which is the difference between a lifetime-giving number you can brief a gift officer on and one inflated by refunds. Worked version: Ingest CiviCRM contributions.

UniSchema. PhilanthroPy is also the modeling half of an ecosystem. UniSchema normalizes fragmented advancement webhooks (GiveCampus, Slate, NPSP, Cvent, …) into a single ConstituentEvent stream. Webhooks → UniSchema egress → read_constituent_events()constituent_events_to_features()predict_affinity_score(). Worked, runnable version with the full diagram: Ingest UniSchema events.


Feature overview

Full parameter documentation for every symbol below is rendered in the API reference.

🧹 Preprocessing

Transformer Description
CRMCleaner Standardise raw CRM exports — coerce gift_date to datetime64 and gift_amount to float64
WealthScreeningImputer Leakage-safe wealth imputation (median / mean / zero), fill stats frozen at fit()
WealthScreeningImputerKNN Leakage-safe KNN imputation for wealth-screening vendor columns
WealthPercentileTransformer Per-column wealth percentile rank (0–100); NaN-in → NaN-out
FiscalYearTransformer Fiscal year & quarter from gift dates; configurable start month
RFMTransformer Recency–Frequency–Monetary features for donor segmentation
ShareOfWalletScorer Normalised Share-of-Wallet score + capacity_tier encoding
MatchingGiftFeaturizer Employer matching-gift eligibility and expected-match features
EncounterTransformer Bridge EHR encounters with the CRM; drops identifier-like columns by name
EncounterRecencyTransformer Encounter-date columns → predictive recency features
GratefulPatientFeaturizer Clinical gravity score + service-line capacity weights
DischargeToSolicitationWindowTransformer in_solicitation_window (0/1) and window_position_score [0,1]
SolicitationWindowTransformer Supported alias of DischargeToSolicitationWindowTransformer
PlannedGivingSignalTransformer Bequest / legacy-gift intent vector

🤖 Models

Model Description
DonorPropensityModel Random Forest with predict_affinity_score() on a 0–100 scale
MajorGiftClassifier Calibrated HistGradientBoostingClassifier — NaN-native
LapsePredictor Random Forest for donor lapse, with predict_lapse_score()
PlannedGivingIntentScorer Calibrated bequest-intent scorer, predict_intent_score()
ShareOfWalletRegressor Total giving capacity and untapped-potential ratio
AskAmountRecommender Conservative / target / stretch ask ladder via ask_ladder()
MovesManagementClassifier Multi-class portfolio stage predictor
FinancialForecastModel Hybrid LSTM-ARIMA revenue forecaster, dependency-free
PropensityScorer Constant-probability baseline — a floor to beat, not a scorer

📊 Metrics, splitters, and the rest

Symbol Module Description
donor_lifetime_value metrics Discounted LTV annuity
donor_retention_rate, donor_acquisition_cost metrics Core campaign KPIs
cost_per_dollar_raised, fundraising_roi metrics Campaign efficiency
gift_concentration_gini, top_donor_share metrics Portfolio concentration
disparate_impact_ratio, selection_rate_by_group metrics Four-fifths-rule fairness audit
FiscalYearGroupedSplitter model_selection Walk-forward fiscal-year CV
donor_feature_importance inspection Permutation importance for any fitted estimator
constituent_events_to_features, read_constituent_events ingest UniSchema bridge
civicrm_contributions_to_features, read_civicrm_contributions ingest CiviCRM contribution-export bridge
generate_synthetic_donor_data, load_ciob_fundraising datasets Synthetic pool and a real CIOB series
make_donor_dataset, save_model, load_model utils Labelled fixtures and pipeline persistence
plot_affinity_distribution, plot_retention_waterfall visualisation Matplotlib is imported lazily, per function
UpliftTLearner experimental T-learner appeal uplift — no API guarantees

Guides

TutorialsBuilding your first model · Avoiding temporal data leakage · Building a grateful-patient pipeline

How-toUse the CLI · Ingest UniSchema events · Ingest CiviCRM contributions · Handle missing wealth data · Build grateful-patient features · Recommend ask amounts · Score matching-gift eligibility · Measure campaign efficiency · Audit score fairness · Estimate appeal uplift · Save and load models · Develop and test

ExplanationDesign principles · Capacity and loyalty · Fundraising metrics · Compliance considerations · Benchmarks


Roadmap

🔜 Next

  • philanthropy.visualisation.plot_capacity_heatmap()
  • EnsemblePropensityModel (stacked LapsePredictor + DonorPropensityModel)

Research

S. A. Lalakiya, "AI for Advancement: Predictive Donor Analytics and Fundraising Intelligence at Scale," 2025 IEEE 11th ICCED, IEEE, 2025, doi: 10.1109/ICCED68324.2025.11325064.

This is the library author's own related work on the same problem space, using a different dataset and its own models. It is not an independent evaluation or a benchmark of PhilanthroPy. To cite the software itself, see CITATION.cff.


Generative AI disclosure

AI assistance (Claude Code) was used during development of this package, across implementation, tests, and documentation, in an agentic workflow rather than line completion alone.

What was not generated. The design constraints are the author's and predate any generated code: the leakage-safety contract (every fitted statistic is computed on training data inside fit and frozen before transform/predict), the dependency rule (scikit-learn, pandas, numpy, matplotlib, seaborn — no deep learning frameworks), the estimator conventions, and the stability tiers. Generated code that violated them was rejected rather than merged.

Human review. Nothing lands without the full gate green. Locally, make ci runs flake8, mypy, the docstring examples, and the test suite against a 92% coverage floor (pyproject.toml). CI additionally enforces a 93% coverage floor on the risk-tier subtree, runs the suite across an OS and Python-version matrix, installs at the declared dependency floors on Python 3.9, builds the distributions and checks their metadata with twine, and verifies the package imports without a plotting stack installed. Every public estimator passes sklearn.utils.estimator_checks.check_estimator.

No approximate scale is attached to the AI use here. The author has not measured the split and will not estimate one; the mechanism and the review gate above are stated instead. This follows the pyOpenSci generative AI policy.


Contributing

Contributions are welcome, and a first PR does not need to be big — docs fixes, missing tests, and clearer error messages all count.

Start with a good first issue. Each one names the files to touch, the steps, and the single command that proves it is done. Comment on the issue to claim it; ask there if anything is unclear.

See CONTRIBUTING.md for the fork-and-PR workflow, the full local test gate, and pre-push hook setup. In short: fork, branch, run make ci before every push, and never use git push --no-verify. Setup plus a first green make ci takes about eight minutes.

Everyone who has landed a change is credited in CONTRIBUTORS.md — add yourself in the same PR.

Questions are welcome in Discussions.


License

MIT License — see LICENSE for details.

About

scikit-learn–native predictive analytics for nonprofit & academic-medical-center fundraising: donor propensity, lapse, planned giving, wealth screening, and revenue forecasting.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages