Every model in the toolkit values a customer with the same multi-horizon identity — an actuarial annuity with retention in the role of survival:
Every term, in plain language (subscript i = one customer; t = a future policy year):
| Term | Name | What it means |
|---|---|---|
Σt=1..T | sum over the horizon | Add up each future year's contribution, year 1 through the horizon T (here 5 years). |
Si(t) | survival / retention | Probability the customer is still on the books (has kept renewing) in year t. Declines as t grows. This is the single piece each of the six models estimates differently. |
Pi(t) | premium | Expected premium collected in year t. |
Li(t) | losses | Expected losses (claims + loss-adjustment expense) in year t. |
Ei(t) | expense | Expenses in year t — the higher acquisition load in year 1, the renewal expense thereafter (§2.2). |
Xi(t) | growth value | Cross-sell + upsell (+ referral) value in year t. = 0 in the filing variant; added in the planning variant. |
1/(1+d)t | discount factor | Converts a year-t dollar to today's value; d = annual discount rate (here 8%). |
Because premium, losses and expense are roughly level per renewal, the identity factors into two pieces you can read off separately:
| Piece | Formula | What it means |
|---|---|---|
| DERT Discounted Expected Residual Transactions |
Σt=1..T S(t)/(1+d)t |
The discounted count of future ("residual" = remaining, from now forward) renewals the customer is expected to make — a survival-weighted, discounted annuity factor. This is the number each of the six models produces differently. |
| annual net margin |
P − L − E ( + X ) |
The dollar contribution of one renewal year. Identical across all models. Filing variant excludes X; planning variant adds it. |
One word, two meanings — in DERT, "residual" means future / remaining transactions. Separately, "residual CLV" in the expense discussion (§2.2) means the value of an already-in-force customer, charging renewal expense only. Same word, different idea.
So the whole of Phase 3 is: one annuity, six independent estimates of the retention curve
S(t).
Formula: docs/01_clv_framework.md §2. Two variants (filing vs planning),
§2.1. Expense timing (residual charges renewal expense only), §2.2.
Before any model, here is the calculation an actuary can do on the back of an envelope, using only published market figures. A full-coverage auto household:
Public sources (all loaded 2026-07-02, see research/data/calibration_benchmarks.md):
premium — NerdWallet/Quadrant Jul 2026 & NAIC; loss ratio — NAIC 2024 P&C Industry Analysis
Report (personal lines pure LR 60–66%); expense — NAIC industry expense ratio 25.2%, split
per agency commission benchmarks (Agentero); retention 88% — ASNOA independent-agency personal-lines
benchmark (85–90%); LexisNexis 2024 carrier auto retention 78% (hard market); discount 8% —
Gupta, Lehmann & Stuart (2004) CLV range 8–16%, P&C cost of equity ~6% (Damodaran).
Where the 2.82 comes from: with a constant 88% annual retention, the survival curve is just
S(t) = 0.88t (renew 5 years running = 0.88×0.88×… ). Plug into DERT:
0.88/1.08 + 0.882/1.082 + … + 0.885/1.085 =
0.815 + 0.664 + 0.541 + 0.441 + 0.359 = 2.82. That flat-retention DERT is the actuary's
baseline annuity. Each model below replaces the flat 88% with a data-driven
retention/survival curve — and, reassuringly, they all land close to it.
Toolkit output on the synthetic book (5,000 customers, 32,728 policy-terms, 2015–2025,
set_seed(42)), horizon 5 yr, discount 8%, expense split 40%/20%, cross-sell $150 +
upsell $100. Representative in-force customer #3848: 4 active years, renewed 3×,
avg premium $2,896, avg loss $879, so annual margin = $1,438 (filing) / $1,688 (planning).
| Model family | DERT | CLV filing | CLV planning | book mean (filing) | book mean (planning) |
|---|---|---|---|---|---|
| BG/NBD | 3.14 | $4,513 | $5,298 | $2,805 | $3,499 |
| Pareto/NBD | 3.14 | $4,515 | $5,300 | $2,759 | $3,394 |
| Survival — Cox PH | 3.52 | $5,068 | $5,949 | $2,879 | $3,499 |
| Survival — Weibull AFT | 3.21 | $4,609 | $5,410 | $2,689 | $3,257 |
| Markov — product states | 3.09 | $4,439 | $5,211 | $2,575 | $3,129 |
| Markov — HMM (3 latent) | 3.07 | $4,413 | $5,181 | $2,601 | $3,163 |
| Ensemble — GBM | 3.11 | $4,479 | $5,257 | $2,799 | $3,387 |
| Hybrid — BG/NBD + ML HEADLINE | 3.19 | $4,589 | $5,386 | $2,887 | $3,494 |
The comfort point for reviewers: six mathematically unrelated engines — closed-form BTYD, semiparametric hazard, parametric AFT, multi-state chains, a latent HMM, and gradient boosting — value the same customer within ±7% ($4,413–$5,068 filing basis) and the same book within a tight band. CLV here is not an artifact of one model choice. Note DERT × margin reproduces each CLV exactly (e.g. BG/NBD 3.14 × $1,438 ≈ $4,515).
Reproduce: python docs/calls/phase3_worked_examples.py.
There is deliberately no single winning number. For CLV in ratemaking, model choice balances two different things — raw predictive accuracy, and fitness for a rate filing. The toolkit measures both.
Computed on a 20% customer-disjoint holdout, targeting each customer's realized discounted renewals (5-yr horizon, 8% discount, reference book). Bold = best on that metric.
| Metric | What it asks | Best | Where models land |
|---|---|---|---|
| Holdout MAE Mean Absolute Error |
Average size of the prediction miss (discounted renewal-years) | lower ↓ | Ensemble 0.632 · Hybrid 0.655 |
| Validation R² | Share of variance in realized value the model explains | higher ↑ | Ensemble 0.617 · Hybrid 0.560 (>½ of what BG/NBD misses is recoverable) |
| 1-yr renewal calibration | Does predicted near-term retention match the empirical 0.924? | closest | Ensemble 0.923 · Cox 0.956 (over) · Pareto 0.875 · BG/NBD 0.807 (under) |
| Rank correlation Spearman |
Do models agree on who is high- vs low-value? | higher ↑ | BG/NBD vs Pareto 0.946 · vs Hybrid 0.895 · vs Cox 0.820 |
The subtlety that matters for pricing: for ratemaking, ranking (Spearman) usually matters more than absolute level (R²/MAE) — you are segmenting customers, not quoting one number. A model can have modest R² and still rank customers correctly, which is what drives the rate relativities.
Pure accuracy does not pick the winner here. These three override it:
| Factor | Why it decides |
|---|---|
| Filing defensibility / interpretability | Can you show a regulator the mechanism? BTYD closed forms and Cox hazard ratios are transparent; a pure GBM is not. This is exactly why the POG chose the Hybrid — it keeps the probabilistic base for filing exhibits and lets ML only correct the residual: ensemble-grade accuracy without surrendering explainability. |
| Loss-correlation | Does the CLV ranking track loss experience? On this book the loss-ratio gradient runs 1.68 in the bottom quintile → 0.16 in the top — a CLV segmentation that correlates with loss cost is filing-relevant (NY §2304 / CA Prop 103); one that ranks on willingness-to-pay is not. |
| Reproducibility & runtime | Deterministic (set_seed(42)), pinned dependencies, CI, fast enough to refit — itself
a governance answer for the NAIC AI Model Bulletin and Colorado SB21-169. |
Bottom line: the headline performance scorecard is holdout MAE + validation R² + 1-yr renewal calibration + rank correlation. But model selection for this contract is won by the model that pairs strong accuracy with interpretability and loss-correlation — which is why the Hybrid is the headline: best-in-class ML accuracy, probabilistic transparency for filings. (Which of these columns lead the published comparison table is exactly poll question Q5.)
Accuracy metrics from EnsembleCLV.params_ (val R²/MAE) and paper §6.6;
loss-ratio gradient from cas_clv.ratemaking.clv_rate_relativities (paper §7.1).
Treats each annual renewal as a "purchase occasion." Renewals follow a Poisson process while the customer is active (rate Gamma-distributed across the book); after each renewal there is a Beta-distributed chance the relationship silently ends. Closed-form (Fader–Hardie–Lee 2005), so every step is auditable — no black box. Output: P(still active) × expected future renewals → DERT. Use: the probabilistic base of the headline hybrid; filing-defensible.
DERT 3.14CLV filing $4,513 p(alive) & renewal count fully explicitThe original "counting your customers" model (Schmittlein–Morrison–Colombo 1987). Same Poisson renewals, but the customer's unobserved lifetime is exponential (dropout can happen at any time, not only just after a renewal). We fit it natively beside BG/NBD on the identical RFM panel. Why show it: two independent BTYD formulations landing on the same DERT (3.14) tells reviewers the retention estimate is robust to the dropout assumption, not a modeling accident.
DERT 3.14CLV filing $4,515 agrees with BG/NBD to the dollarThe most familiar framing for a pricing actuary: model the hazard of lapse exactly as you would model mortality. Covariates (age, household size, # products, log premium) shift the survival curve; Cox PH makes no shape assumption (semiparametric), Weibull AFT is fully parametric. An in-force customer is valued by conditional residual survival — "given they've kept the policy d years, what's the discounted expected future renewal stream." Cox lands a touch higher (DERT 3.52) because tenure lowers the lapse hazard — the seasoning effect actuaries already price. Use: directly comparable to your retention/persistency analyses (Brockett et al. 2008 household portfolio).
Cox DERT 3.52 → $5,068 Weibull DERT 3.21 → $4,609 covariate hazard ratios exportableStates = number of products the household holds (1, 2, 3) plus an absorbing "lapsed" state. We estimate the annual transition matrix by MLE (Laplace-smoothed) straight from the book, then value each customer by discounted expected active years from their current state. Because multi-product households retain materially better, a higher state carries real retention signal — the cross-sell– retention link, made mechanical. Use: the natural home for upsell/cross-sell dynamics and the multi-state churn literature (Dong–Frees, ASTIN 2022).
DERT 3.09CLV filing $4,439 transition matrix inspectableSame multi-state machinery, but when the meaningful "state" isn't directly observed. A Gaussian HMM decodes latent engagement states from each year's behavior (# products, coverage tier), then the transition/valuation runs on the decoded states. It reproduces the product-state Markov result almost exactly (DERT 3.07 vs 3.09) — a useful cross-check that the observed product-count states already capture most of the latent structure. Use: where engagement is richer than a simple product count.
DERT 3.07CLV filing $4,413 latent states data-decodedThe model the POG selected at kickoff. The BG/NBD closed form supplies the per-year renewal base; gradient boosting (XGBoost + LightGBM, customer-disjoint splits) learns only the residual the probabilistic model misses. Filing exhibits keep the probabilistic base; the ML layer sharpens it. Crucially, the interpretability exhibit shows the ML layer leans on tenure (0.35), renewal frequency (0.23), premium (0.16) — i.e. it refines the retention signal, it does not smuggle in new rating variables. Validation R² ≈ 0.53 (hybrid) / 0.60 (pure ensemble). DERT 3.19 sits squarely in the pack — the correction is a nudge, not an override.
DERT 3.19CLV filing $4,589 top feature: tenure (T) 0.35val R² 0.53⚠ One letter, two meanings — T: in the CLV
formula, T = the projection horizon (5 years here). In the customer summary (the
T=3 for customer #3848), T = the customer's observed "age" — years elapsed
since their first policy. Different quantities, same historical letter (kept to match the source
papers).
| Symbol / term | Meaning |
|---|---|
| CLV | Customer Lifetime Value — present value of a customer's future net cash flows. |
| DERT | Discounted Expected Residual Transactions — the discounted count of a customer's future ("residual") renewals; the survival-weighted annuity factor. CLV = DERT × margin. |
S(t) | Survival / retention probability — chance the customer is still renewing in year t. |
d | Discount rate (8% here) — the cost of waiting a year for a dollar. |
P, L, E, X | Per-year Premium, Losses, Expense, and growth value X (cross-sell + upsell + referral). |
| annuity | A stream of level future payments; CLV is an annuity in which each payment is weighted by survival S(t). |
| Term | Meaning |
|---|---|
| RFM | Recency, Frequency, Monetary — the compact per-customer history the models are fit on. |
| frequency | Number of repeat renewals after the first policy year. |
| recency | How long ago the most recent renewal was, measured from entry. |
| T (RFM age) | Years observed since the customer's first policy (≠ the horizon T above). |
| p(alive) | BTYD probability the customer has not silently churned — has renewals still ahead. |
| Term | Meaning |
|---|---|
| BTYD | "Buy-Till-You-Die" — the family that models how many purchases (here, renewals) a customer makes before silently lapsing. |
| BG/NBD | Beta-Geometric / Negative Binomial Distribution (Fader–Hardie–Lee 2005). |
| NBD | Negative Binomial Distribution — the count model for number of renewals. |
| Pareto/NBD | Pareto / Negative Binomial Distribution (Schmittlein et al. 1987) — the classic BTYD benchmark. |
| Poisson process | A model for events (renewals) occurring randomly over time at some rate. |
| Gamma / Beta | Distributions used to spread the renewal rate / dropout chance across customers (heterogeneity). |
| Term | Meaning |
|---|---|
| hazard | The instantaneous rate of lapse at a given tenure — the lapse analog of a mortality rate. |
| Cox PH | Cox Proportional Hazards — a semiparametric survival model (leaves the baseline curve shape free). |
| Weibull AFT | Weibull Accelerated Failure Time — a parametric survival model (assumes a Weibull shape). |
| conditional residual survival | S(d+t)/S(d) — chance of surviving t more years given the customer has already survived d. |
| MLE | Maximum Likelihood Estimation — the standard method for fitting the model parameters to data. |
| multi-state / Markov chain | Model of year-to-year moves between discrete states (here, number of products held). |
| absorbing state | A state you can't leave — here, "lapsed." |
| transition matrix | The table of probabilities of moving from each state to each other in one year. |
| Laplace smoothing | Adding a small count so no transition probability is exactly zero (numerical robustness). |
| HMM | Hidden Markov Model — a multi-state model whose states are inferred from behavior, not directly observed. |
| Term | Meaning |
|---|---|
| GBM | Gradient Boosting Machine — an ensemble of small decision trees; strong tabular predictor. |
| XGBoost / LightGBM | The two industry-standard GBM implementations we average. |
| residual (ML sense) | The part of the outcome the base (BG/NBD) model didn't capture, which the GBM then learns. |
| credibility | The actuarial idea of blending a base estimate with data-driven signal — the intuition behind the hybrid. |
| val R² | Validation R-squared — share of variance explained on held-out customers (higher = better; 0.53 for the hybrid). |
| MAE | Mean Absolute Error — average size of the prediction miss. |
| feature importance | How much each input variable drives the model's predictions (the interpretability exhibit). |
| Term | Meaning |
|---|---|
| LR | Loss Ratio — losses ÷ premium. |
| LAE | Loss Adjustment Expense — the cost of handling/settling claims. |
| filing / unconstrained CLV | P − L − E only (X = 0) — the rate-filing-defensible base. |
| planning / constrained CLV | Adds cross-sell + upsell value — the business-planning view. |
| residual CLV | Value of an already-in-force customer — future terms charge renewal expense only (acquisition is sunk). |
| inception / new-business CLV | Residual CLV minus the acquisition expense — the acquisition-decision view. |
| acquisition vs renewal expense | First-year load (40% of premium) vs ongoing renewal expense (20%). |
set_seed(42) | Fixes all randomness so every number here is exactly reproducible. |
python docs/calls/phase3_worked_examples.py on the seeded
synthetic book (set_seed(42)) — deterministic, 60+ tests / ruff / mypy green.research/data/calibration_benchmarks.md
(NAIC 2024 report, III/ISO, LexisNexis, NerdWallet/Quadrant, Gupta 2004, Damodaran).