Superseded, 7 September 2026. This page was built for the 16 July 2026 call and keeps the vocabulary and figures of paper v1.x: unconstrained / constrained are now base / with-growth, filing / planning basis are now revenue-neutral / level-effect basis, and the cross-sell value is credited once rather than every term (so the growth gap and cohort means quoted here are pre-v3.0). The current numbers and glossary are in the 10 September worked examples and the paper v4.0.
← Project hub

Phase 3 — CLV worked examples, model family by model family

CAS Ratemaking Working Group · Call #3 · Jul 16, 2026 · every number is reproducible from the toolkit and anchored to public benchmarks

One formula, six ways to estimate it

Every model in the toolkit values a customer with the same multi-horizon identity — an actuarial annuity with retention in the role of survival:

CLVi = Σt=1..T  Si(t) · [ Pi(t) − Li(t) − Ei(t) + Xi(t) ] / (1 + d)t

Every term, in plain language (subscript i = one customer; t = a future policy year):

TermNameWhat it means
Σt=1..Tsum over the horizon Add up each future year's contribution, year 1 through the horizon T (here 5 years).
Si(t)survival / retention Probability the customer is still on the books (has kept renewing) in year t. Declines as t grows. This is the single piece each of the six models estimates differently.
Pi(t)premium Expected premium collected in year t.
Li(t)losses Expected losses (claims + loss-adjustment expense) in year t.
Ei(t)expense Expenses in year t — the higher acquisition load in year 1, the renewal expense thereafter (§2.2).
Xi(t)growth value Cross-sell + upsell (+ referral) value in year t. = 0 in the filing variant; added in the planning variant.
1/(1+d)tdiscount factor Converts a year-t dollar to today's value; d = annual discount rate (here 8%).

Because premium, losses and expense are roughly level per renewal, the identity factors into two pieces you can read off separately:

CLV = DERT × annual net margin
PieceFormulaWhat it means
DERT
Discounted Expected Residual Transactions
Σt=1..T S(t)/(1+d)t The discounted count of future ("residual" = remaining, from now forward) renewals the customer is expected to make — a survival-weighted, discounted annuity factor. This is the number each of the six models produces differently.
annual net
margin
P − L − E ( + X ) The dollar contribution of one renewal year. Identical across all models. Filing variant excludes X; planning variant adds it.

One word, two meanings — in DERT, "residual" means future / remaining transactions. Separately, "residual CLV" in the expense discussion (§2.2) means the value of an already-in-force customer, charging renewal expense only. Same word, different idea.

So the whole of Phase 3 is: one annuity, six independent estimates of the retention curve S(t).

Formula: docs/01_clv_framework.md §2. Two variants (filing vs planning), §2.1. Expense timing (residual charges renewal expense only), §2.2.

The anchor example — a personal-auto household, hand-checkable from public data

Before any model, here is the calculation an actuary can do on the back of an envelope, using only published market figures. A full-coverage auto household:

Annual premium (full-coverage auto, national avg)$2,300
− Losses @ 66% loss+LAE ratio (personal lines, NAIC 2024)− $1,518
− Renewal expense @ 20% (renewal commission + servicing)− $460
= Annual net margin (unconstrained / filing basis)$322
DERT = Σt=1..5 0.88t/1.08t  (retention 88%, discount 8%)2.82
CLV (filing basis) = 2.82 × $322≈ $908
+ cross-sell $150 + upsell $100 / yr → margin $572 → CLV (planning)≈ $1,613
− acquisition expense (40% × $2,300 = $920) → inception CLV≈ $693

Public sources (all loaded 2026-07-02, see research/data/calibration_benchmarks.md): premium — NerdWallet/Quadrant Jul 2026 & NAIC; loss ratio — NAIC 2024 P&C Industry Analysis Report (personal lines pure LR 60–66%); expense — NAIC industry expense ratio 25.2%, split per agency commission benchmarks (Agentero); retention 88% — ASNOA independent-agency personal-lines benchmark (85–90%); LexisNexis 2024 carrier auto retention 78% (hard market); discount 8% — Gupta, Lehmann & Stuart (2004) CLV range 8–16%, P&C cost of equity ~6% (Damodaran).

Where the 2.82 comes from: with a constant 88% annual retention, the survival curve is just S(t) = 0.88t (renew 5 years running = 0.88×0.88×… ). Plug into DERT: 0.88/1.08 + 0.882/1.082 + … + 0.885/1.085 = 0.815 + 0.664 + 0.541 + 0.441 + 0.359 = 2.82. That flat-retention DERT is the actuary's baseline annuity. Each model below replaces the flat 88% with a data-driven retention/survival curve — and, reassuringly, they all land close to it.

All six families on one real customer — and they agree

Toolkit output on the synthetic book (5,000 customers, 32,728 policy-terms, 2015–2025, set_seed(42)), horizon 5 yr, discount 8%, expense split 40%/20%, cross-sell $150 + upsell $100. Representative in-force customer #3848: 4 active years, renewed 3×, avg premium $2,896, avg loss $879, so annual margin = $1,438 (filing) / $1,688 (planning).

Model familyDERTCLV filing CLV planningbook mean (filing) book mean (planning)
BG/NBD3.14$4,513$5,298$2,805$3,499
Pareto/NBD3.14$4,515$5,300$2,759$3,394
Survival — Cox PH3.52$5,068$5,949$2,879$3,499
Survival — Weibull AFT3.21$4,609$5,410$2,689$3,257
Markov — product states3.09$4,439$5,211$2,575$3,129
Markov — HMM (3 latent)3.07$4,413$5,181$2,601$3,163
Ensemble — GBM3.11$4,479$5,257$2,799$3,387
Hybrid — BG/NBD + ML HEADLINE3.19$4,589$5,386$2,887$3,494

The comfort point for reviewers: six mathematically unrelated engines — closed-form BTYD, semiparametric hazard, parametric AFT, multi-state chains, a latent HMM, and gradient boosting — value the same customer within ±7% ($4,413–$5,068 filing basis) and the same book within a tight band. CLV here is not an artifact of one model choice. Note DERT × margin reproduces each CLV exactly (e.g. BG/NBD 3.14 × $1,438 ≈ $4,515).

Reproduce: python docs/calls/phase3_worked_examples.py.

How we judge which model is "best"

There is deliberately no single winning number. For CLV in ratemaking, model choice balances two different things — raw predictive accuracy, and fitness for a rate filing. The toolkit measures both.

A. Accuracy scorecard — backtest on held-out customers

Computed on a 20% customer-disjoint holdout, targeting each customer's realized discounted renewals (5-yr horizon, 8% discount, reference book). Bold = best on that metric.

MetricWhat it asksBestWhere models land
Holdout MAE
Mean Absolute Error
Average size of the prediction miss (discounted renewal-years)lower ↓ Ensemble 0.632 · Hybrid 0.655
Validation R² Share of variance in realized value the model explainshigher ↑ Ensemble 0.617 · Hybrid 0.560 (>½ of what BG/NBD misses is recoverable)
1-yr renewal calibration Does predicted near-term retention match the empirical 0.924?closest Ensemble 0.923 · Cox 0.956 (over) · Pareto 0.875 · BG/NBD 0.807 (under)
Rank correlation
Spearman
Do models agree on who is high- vs low-value?higher ↑ BG/NBD vs Pareto 0.946 · vs Hybrid 0.895 · vs Cox 0.820

The subtlety that matters for pricing: for ratemaking, ranking (Spearman) usually matters more than absolute level (R²/MAE) — you are segmenting customers, not quoting one number. A model can have modest R² and still rank customers correctly, which is what drives the rate relativities.

B. Fitness for a rate filing — usually the deciding factors

Pure accuracy does not pick the winner here. These three override it:

FactorWhy it decides
Filing defensibility / interpretability Can you show a regulator the mechanism? BTYD closed forms and Cox hazard ratios are transparent; a pure GBM is not. This is exactly why the POG chose the Hybrid — it keeps the probabilistic base for filing exhibits and lets ML only correct the residual: ensemble-grade accuracy without surrendering explainability.
Loss-correlation Does the CLV ranking track loss experience? On this book the loss-ratio gradient runs 1.68 in the bottom quintile → 0.16 in the top — a CLV segmentation that correlates with loss cost is filing-relevant (NY §2304 / CA Prop 103); one that ranks on willingness-to-pay is not.
Reproducibility & runtime Deterministic (set_seed(42)), pinned dependencies, CI, fast enough to refit — itself a governance answer for the NAIC AI Model Bulletin and Colorado SB21-169.

Bottom line: the headline performance scorecard is holdout MAE + validation R² + 1-yr renewal calibration + rank correlation. But model selection for this contract is won by the model that pairs strong accuracy with interpretability and loss-correlation — which is why the Hybrid is the headline: best-in-class ML accuracy, probabilistic transparency for filings. (Which of these columns lead the published comparison table is exactly poll question Q5.)

Accuracy metrics from EnsembleCLV.params_ (val R²/MAE) and paper §6.6; loss-ratio gradient from cas_clv.ratemaking.clv_rate_relativities (paper §7.1).

Model-by-model — what it is, and the actuarial analog

BG/NBD

Beta-Geometric / NBD

≈ frequency–retention counting model

Treats each annual renewal as a "purchase occasion." Renewals follow a Poisson process while the customer is active (rate Gamma-distributed across the book); after each renewal there is a Beta-distributed chance the relationship silently ends. Closed-form (Fader–Hardie–Lee 2005), so every step is auditable — no black box. Output: P(still active) × expected future renewals → DERT. Use: the probabilistic base of the headline hybrid; filing-defensible.

DERT 3.14CLV filing $4,513 p(alive) & renewal count fully explicit
Pareto/NBD

Pareto / NBD

≈ the classic BTYD benchmark

The original "counting your customers" model (Schmittlein–Morrison–Colombo 1987). Same Poisson renewals, but the customer's unobserved lifetime is exponential (dropout can happen at any time, not only just after a renewal). We fit it natively beside BG/NBD on the identical RFM panel. Why show it: two independent BTYD formulations landing on the same DERT (3.14) tells reviewers the retention estimate is robust to the dropout assumption, not a modeling accident.

DERT 3.14CLV filing $4,515 agrees with BG/NBD to the dollar
SURVIVAL

Cox PH & Weibull AFT (time-to-lapse)

≈ a lapse table / mortality study

The most familiar framing for a pricing actuary: model the hazard of lapse exactly as you would model mortality. Covariates (age, household size, # products, log premium) shift the survival curve; Cox PH makes no shape assumption (semiparametric), Weibull AFT is fully parametric. An in-force customer is valued by conditional residual survival — "given they've kept the policy d years, what's the discounted expected future renewal stream." Cox lands a touch higher (DERT 3.52) because tenure lowers the lapse hazard — the seasoning effect actuaries already price. Use: directly comparable to your retention/persistency analyses (Brockett et al. 2008 household portfolio).

Cox DERT 3.52 → $5,068 Weibull DERT 3.21 → $4,609 covariate hazard ratios exportable
MARKOV

Multi-state chain over product-portfolio states

≈ a multi-state life / disability model

States = number of products the household holds (1, 2, 3) plus an absorbing "lapsed" state. We estimate the annual transition matrix by MLE (Laplace-smoothed) straight from the book, then value each customer by discounted expected active years from their current state. Because multi-product households retain materially better, a higher state carries real retention signal — the cross-sell– retention link, made mechanical. Use: the natural home for upsell/cross-sell dynamics and the multi-state churn literature (Dong–Frees, ASTIN 2022).

DERT 3.09CLV filing $4,439 transition matrix inspectable
HMM

Hidden-Markov latent engagement states

≈ multi-state model when the state is unobserved

Same multi-state machinery, but when the meaningful "state" isn't directly observed. A Gaussian HMM decodes latent engagement states from each year's behavior (# products, coverage tier), then the transition/valuation runs on the decoded states. It reproduces the product-state Markov result almost exactly (DERT 3.07 vs 3.09) — a useful cross-check that the observed product-count states already capture most of the latent structure. Use: where engagement is richer than a simple product count.

DERT 3.07CLV filing $4,413 latent states data-decoded
HYBRID

BG/NBD base + ML residual correction POG HEADLINE

≈ a credibility-style blend

The model the POG selected at kickoff. The BG/NBD closed form supplies the per-year renewal base; gradient boosting (XGBoost + LightGBM, customer-disjoint splits) learns only the residual the probabilistic model misses. Filing exhibits keep the probabilistic base; the ML layer sharpens it. Crucially, the interpretability exhibit shows the ML layer leans on tenure (0.35), renewal frequency (0.23), premium (0.16) — i.e. it refines the retention signal, it does not smuggle in new rating variables. Validation R² ≈ 0.53 (hybrid) / 0.60 (pure ensemble). DERT 3.19 sits squarely in the pack — the correction is a nudge, not an override.

DERT 3.19CLV filing $4,589 top feature: tenure (T) 0.35val R² 0.53

Full glossary — every symbol & abbreviation used on this page

⚠ One letter, two meanings — T: in the CLV formula, T = the projection horizon (5 years here). In the customer summary (the T=3 for customer #3848), T = the customer's observed "age" — years elapsed since their first policy. Different quantities, same historical letter (kept to match the source papers).

Core CLV notation

Symbol / termMeaning
CLVCustomer Lifetime Value — present value of a customer's future net cash flows.
DERTDiscounted Expected Residual Transactions — the discounted count of a customer's future ("residual") renewals; the survival-weighted annuity factor. CLV = DERT × margin.
S(t)Survival / retention probability — chance the customer is still renewing in year t.
dDiscount rate (8% here) — the cost of waiting a year for a dollar.
P, L, E, XPer-year Premium, Losses, Expense, and growth value X (cross-sell + upsell + referral).
annuityA stream of level future payments; CLV is an annuity in which each payment is weighted by survival S(t).

Customer-history summary (feeds every model)

TermMeaning
RFMRecency, Frequency, Monetary — the compact per-customer history the models are fit on.
frequencyNumber of repeat renewals after the first policy year.
recencyHow long ago the most recent renewal was, measured from entry.
T (RFM age)Years observed since the customer's first policy (≠ the horizon T above).
p(alive)BTYD probability the customer has not silently churned — has renewals still ahead.

The BTYD (probabilistic counting) models

TermMeaning
BTYD"Buy-Till-You-Die" — the family that models how many purchases (here, renewals) a customer makes before silently lapsing.
BG/NBDBeta-Geometric / Negative Binomial Distribution (Fader–Hardie–Lee 2005).
NBDNegative Binomial Distribution — the count model for number of renewals.
Pareto/NBDPareto / Negative Binomial Distribution (Schmittlein et al. 1987) — the classic BTYD benchmark.
Poisson processA model for events (renewals) occurring randomly over time at some rate.
Gamma / BetaDistributions used to spread the renewal rate / dropout chance across customers (heterogeneity).

Survival & multi-state models

TermMeaning
hazardThe instantaneous rate of lapse at a given tenure — the lapse analog of a mortality rate.
Cox PHCox Proportional Hazards — a semiparametric survival model (leaves the baseline curve shape free).
Weibull AFTWeibull Accelerated Failure Time — a parametric survival model (assumes a Weibull shape).
conditional residual survivalS(d+t)/S(d) — chance of surviving t more years given the customer has already survived d.
MLEMaximum Likelihood Estimation — the standard method for fitting the model parameters to data.
multi-state / Markov chainModel of year-to-year moves between discrete states (here, number of products held).
absorbing stateA state you can't leave — here, "lapsed."
transition matrixThe table of probabilities of moving from each state to each other in one year.
Laplace smoothingAdding a small count so no transition probability is exactly zero (numerical robustness).
HMMHidden Markov Model — a multi-state model whose states are inferred from behavior, not directly observed.

Machine-learning & hybrid

TermMeaning
GBMGradient Boosting Machine — an ensemble of small decision trees; strong tabular predictor.
XGBoost / LightGBMThe two industry-standard GBM implementations we average.
residual (ML sense)The part of the outcome the base (BG/NBD) model didn't capture, which the GBM then learns.
credibilityThe actuarial idea of blending a base estimate with data-driven signal — the intuition behind the hybrid.
val R²Validation R-squared — share of variance explained on held-out customers (higher = better; 0.53 for the hybrid).
MAEMean Absolute Error — average size of the prediction miss.
feature importanceHow much each input variable drives the model's predictions (the interpretability exhibit).

Economics, variants & general

TermMeaning
LRLoss Ratio — losses ÷ premium.
LAELoss Adjustment Expense — the cost of handling/settling claims.
filing / unconstrained CLVP − L − E only (X = 0) — the rate-filing-defensible base.
planning / constrained CLVAdds cross-sell + upsell value — the business-planning view.
residual CLVValue of an already-in-force customer — future terms charge renewal expense only (acquisition is sunk).
inception / new-business CLVResidual CLV minus the acquisition expense — the acquisition-decision view.
acquisition vs renewal expenseFirst-year load (40% of premium) vs ongoing renewal expense (20%).
set_seed(42)Fixes all randomness so every number here is exactly reproducible.

Reproducibility & provenance

This research project has been funded by the Casualty Actuarial Society.