Everything in the paper — six model families, four ratemaking exhibits, five
regulatory packages — reduces to one identity. The families differ only in how they produce
nt.
| Term | Name | What it means, in words |
|---|---|---|
nt | Discounted renewal increment | The discounted probability that the customer is still on the books for term t specifically — differenced from the model's own cumulative expectation, so it sums back exactly. Max absolute reconstruction error across all six families: 0.0. |
DERT(t) | Discounted Expected Residual Transactions | Cumulative discounted renewals expected through term t. This is the quantity each model family estimates differently. |
Pt | Premium | Written premium for the term. |
Lt | Loss | Projected loss including ALAE (Economics.lae_ratio). May vary with
tenure; by default it does not — see §4. |
Et | Expense | Variable expense, including ULAE. Split acquisition (0.40) vs renewal (0.20): residual CLV charges renewal expense only, because acquisition is sunk for an in-force customer. Fixed expense is excluded from a marginal CLV by design. |
DERT ×
margin product cannot express a margin that varies by year. Writing it as a sum over
terms can — which is what let the tenure-varying loss work land inside all six families without
re-estimating anything.Before any synthetic book, the identity should survive a back-of-envelope test on published industry figures. Every input below is a public number.
| Input | Value | Source |
|---|---|---|
| Full-coverage personal auto premium | $2,300 | Industry average, full-coverage annual |
| Loss ratio | 66% | NAIC personal auto, incurred |
| Variable expense ratio | 20% | Renewal-expense convention |
| Annual retention | 88% | Industry personal-lines retention |
| Discount rate | 8% | Cost of capital (Damodaran) |
Step 1 — the margin on one term.
Loss = $2,300 × 0.66 = $1,518 · Expense = $2,300 × 0.20 = $460
margin = $2,300 − $1,518 − $460 = $322
Step 2 — how many renewal terms, discounted.
v = r / (1 + d) = 0.88 / 1.08 = 0.81481
DERT(5y) = v(1 − v5) / (1 − v) = 2.8197
Step 3 — multiply.
CLVbase = $322 × 2.8197 = $907.93
This is the paper's central credibility claim. Six structurally different model
families, one predict_clv() contract, fitted to the same seeded synthetic book
(5,000 customers · 32,728 policy terms · 5,617 claims · $103.3M written premium).
| Model family | Mean CLV | Median | DERT (5y) | 1-yr in-force renewal | % CLV positive |
|---|---|---|---|---|---|
| BG/NBD | $2,236.27 | $3,004.27 | 2.77 | 0.8052 | 75.7% |
| Pareto/NBD | $2,232.75 | $2,558.86 | 2.54 | 0.8546 | 75.7% |
| Cox PH | $2,340.37 | $1,098.80 | 2.48 | 0.9555 | 52.4% |
| Weibull AFT | $2,185.20 | $970.97 | 2.27 | 0.9322 | 52.4% |
| Markov | $2,092.40 | $938.12 | 2.22 | 0.9219 | 52.4% |
| Hybrid (POG headline) | $2,361.43 | $1,755.78 | 2.43 | 0.9246 | 67.0% |
Read the medians, not just the means. They diverge far more than the means do, and the reason is structural rather than a defect: the survival and Markov families assign exactly zero to a lapsed customer, and 30.8% of the book has lapsed. That pushes their medians down while leaving their means close to the BTYD families. A reviewer who compares only medians will conclude the models disagree; they do not.
Round 1 produced two written reviews and 24 comments, all incorporated. AJ Robinson's 13 August review left 18, of which four found defects rather than ambiguities; Mark Mondello's 17 August review left 6, which drove the v3.1 and v3.2 exposition changes and surfaced one genuine error and one genuine loose end (§4.5). Because v2.0 had explicitly changed no reported figures, v3.0 is the version where numbers moved — so the changes are itemised rather than summarised.
cross_sell_value was folded into the per-term margin and multiplied by DERT,
valuing a $150 cross-sale at $150 × 2.7736 per customer. Cross-sell is now a
one-time event with an annual hazard, credited once at the modelled time of sale, with
first-occurrence weights summing to at most 1. Upsell is separated as what it actually is: a
permanent step-up in annual margin from the upgrade year onward.
| Growth component | Contribution | Treatment |
|---|---|---|
| Cross-sell | $98.85 | One-time event, annual hazard cross_sell_prob |
| Upsell | $277.36 | Permanent margin step-up |
| Total growth gap (v3.0) | $376.22 | clv_with_growth − clv_base |
| Pre-v3.0 figure | $693.40 | Overstated by the whole annuity |
Credibility was applied to customer counts against the 1,082-claim classical standard. Because the CLV segments are equal-size quintiles, this returned 0.9614 in all five rows — a column carrying no information whatsoever. On claim counts the worst-loss-ratio quintile (2,433 claims against 1,000 customers) correctly reaches full credibility. Indicated relativities did not change; the exhibit's honesty did.
The retention-adjusted loss ratio grouped by tenure only, so it could not support the decision it exists for. Grouped by rating segment instead:
| Segment | Customers | Retention | Expected years | First-term LR | Lifetime LR | Relativity |
|---|---|---|---|---|---|---|
| 1 product | 3,205 | 0.8375 | 5.11 | 0.6515 | 0.5928 | 1.0992 |
| 2 products | 1,545 | 0.8929 | 6.33 | 0.5939 | 0.5132 | 0.9516 |
| 3+ products | 250 | 0.9665 | 8.62 | 0.6652 | 0.5202 | 0.9646 |
Mark Mondello built a discrete CLV worksheet reconstructing the calculation and asked for it to be corrected where his reading missed the intent. It reproduces in our code to the cent — $1,759.73 with growth, $1,610.15 without, gap $149.58 — before any correction. It also reproduced the c16 defect independently: built from the v2.0 text, it credits cross-sell every term, with per-term probabilities summing to 1.15. Two corrections follow: growth shape (+$101.23) and a multiplicative rather than additive tenure credit (−$199.76), for a v3.2-consistent $1,661.20. His acquisition-vs- renewal expense split was already exactly right.
Renamed clv_base / clv_with_growth. Actuaries read "constrained" as
the regulatorily constrained variant, which is the opposite of the sense the July 2
decision intended. Old names ship as deprecated aliases for one release. This amends a decision
this group made, and is presented as an amendment adopted on review.
The Aug 6 poll's reserve question asked the POG to help design the NAIC five-year cross-carrier reproducibility test for the tenure gradient. That test has been attempted. It failed, and the failure is reported here rather than buried.
| Evidence | Result | What it means |
|---|---|---|
| Designed cross-carrier test | Not executable | No public source carries carrier identity and policyholder tenure in the same record. Not a data-access problem another researcher would solve. |
| Wisconsin LGPIF panel, tenures 0–3 | +19.93% | R² 0.949 — apparently strong support… |
| Wisconsin LGPIF, full window | −7.93% | …until one fund-year of catastrophe losses reverses the sign. R² 0.058. |
| Spanish multi-product panel | +1.80% | Not monotone. |
| Schedule P between-carrier dispersion | 24.4% CV | Against a 35.0% ten-year curve effect — same order of magnitude as the effect itself. |
loss_trend remains 0.0 by default and
an exact no-op, asserted per model family by tests — nothing previously published moved.
The paper's §3.5 "provisional" note is replaced: the POG has not ratified the treatment, but the
evidence is no longer provisional.A related caution stands from Call #4 and is worth repeating, because it inverts the obvious reading: on a revenue-neutral filing basis, recognising that loss improves with tenure makes long-tenured business carry relatively more loss. The 10+ year cohort's mean CLV falls 7.0% while younger cohorts rise 1–3%. "Loss ratio improves with tenure" and "give loyal policyholders a discount" are not the same proposition. A tenure relativity built naively from a loss-ratio-by-tenure exhibit is likely backwards.
The reason to trust the numbers is not that they were checked once, but that they cannot drift without the build failing.
| Check | What it prevents |
|---|---|
scripts/check_paper_numbers.py |
Any figure in the paper or executive summary that no longer matches the seeded run. Must report 0 unmatched or the build is not shippable. |
Seeded generation (seed 42) |
Non-reproducible results. Every figure traces to one run. |
| Disjoint-split assertion | Train/test leakage in the ML families — asserted, not assumed. |
| Elasticity invariant test | features.SURVIVAL_COVARIATES gaining a price or rate-change covariate. This
is the invariant on which §7 argues the exhibits are admissible where price optimization is
prohibited. |
loss_trend no-op test, per family |
The unvalidated tenure curve silently affecting a published number. |
| Release curation audit + secret scan | Internal files or credentials reaching the public repository. The release is an explicit allowlist, not an exclusion list. |
| Screen | Groups | Min adverse-impact ratio | Flagged |
|---|---|---|---|
| Geography | 5 | 1.0000 | None |
| Age bands | 5 | 0.9729 | None |
Both clear the four-fifths threshold of 0.80 on every group. The age-band screen was added after Call #4, where it was flagged as a gap before any filing use.
§3 showed six families agreeing on the means and disagreeing on the medians. This is the panel that says how we choose between them, because "which model is best" has no single winning number in ratemaking — raw predictive accuracy and fitness for a rate filing are different questions, and the toolkit measures both.
Computed on a 20% holdout, disjoint by customer, targeting each customer's realized discounted renewals over five years. Bold = best on that metric.
| Metric | What it asks | Better | Where the families land |
|---|---|---|---|
| Holdout MAE | Average size of the miss, in discounted renewal-years | lower | Ensemble 0.632 · Hybrid 0.655 |
| Validation R² | Share of variance in realized value explained | higher | Ensemble 0.617 · Hybrid 0.560 |
| 1-yr renewal calibration | Does predicted near-term retention match the empirical 0.9221? | closest | Ensemble 0.923 · Hybrid 0.9246 · Cox 0.9555 (over) · BG/NBD 0.8052 (under) |
| Rank correlation (Spearman) | Do the families agree on who is high- vs low-value? | higher | BG/NBD vs Pareto 0.946 · vs Hybrid 0.895 · vs Cox 0.820 |
Accuracy does not pick the winner. These three override it, and they are the same three that decide what belongs in a public release.
| Factor | Why it decides | Evidence in this package |
|---|---|---|
| Filing defensibility | Can you show a regulator the mechanism? BTYD closed forms and Cox hazard ratios are transparent; a pure gradient-boosted model is not. This is why the POG chose the Hybrid as the headline — a probabilistic base for the filing exhibits, with ML correcting only the residual. | §1's single identity; the deprecated-alias rename (§4, c14) |
| Loss-correlation | Does the CLV ranking track loss experience rather than willingness to pay? A segmentation that correlates with loss cost is filing-relevant; one that ranks on price response is not, and is banned outright in several states. | Loss-ratio gradient 1.68 in the bottom quintile → 0.16 in the top; the
features.SURVIVAL_COVARIATES invariant test in §6 |
| Reproducibility and governance | Deterministic under set_seed(42), pinned dependencies, CI, and fast enough
to refit — which is itself the answer to the NAIC AI Model Bulletin and Colorado
SB21-169. |
§6 in full: 122 tests inside the release subset itself, drift gate 0 unmatched, curation audit, secret scan |
dist_release/
rather than only in the working repository — which is the specific reason the researcher
recommendation on Q3 is to provision now and let POG review run in parallel, rather than gate
publication on it. The one criterion still carrying open risk is loss-correlation, and that is
carried Aug-6 poll Q3: whether a tenure relativity may enter a filed exhibit at
all. Every point either reviewer raised, with its disposition, is on the
consolidated review-response register.Accuracy metrics from EnsembleCLV.params_ and paper §6.6; the
loss-ratio gradient from cas_clv.ratemaking.clv_rate_relativities().
0.0 and is an exact no-op.python docs/calls/phase6_package_review_data.py, which reads the same two seeded
payloads the drift gate validates the manuscript against — paper/paper_numbers.json
(seed 42) and paper/tenure_validation_numbers.json. If a figure is not in that
output, it is not in the paper either. The §2 anchor uses only published industry figures and is
hand-checkable with a calculator.ruff clean · mypy
clean · drift check 0 unmatched across 289 tracked figures · six notebooks execute clean ·
release subset independently gated at 96 files.