Superseded, 27 September 2026. This page is the record of Call #6 (21 Aug) as presented, and keeps the figures and vocabulary of that date. In paper v5.2 (27 September 2026): the 3+ product relativity is 0.9683 on 9.44 expected years against 6.53 for mono-line (was 0.9646 on 8.62 against 5.11); filing / planning basis are now the revenue-neutral / level-effect basis. The current state is on the project hub.
← Project hub

The package, and how to check it

CAS Ratemaking Working Group · Bi-Weekly Call #6 · Friday, August 21, 2026 · Phase 6 — paper v3.2, executive summary, toolkit

§1One formula, and every symbol in it

Everything in the paper — six model families, four ratemaking exhibits, five regulatory packages — reduces to one identity. The families differ only in how they produce nt.

CLV = Σt   nt  ×  margin(t) where nt = DERT(t) − DERT(t−1)  ·  margin(t) = Pt − Lt − Et
TermNameWhat it means, in words
ntDiscounted renewal increment The discounted probability that the customer is still on the books for term t specifically — differenced from the model's own cumulative expectation, so it sums back exactly. Max absolute reconstruction error across all six families: 0.0.
DERT(t)Discounted Expected Residual Transactions Cumulative discounted renewals expected through term t. This is the quantity each model family estimates differently.
PtPremiumWritten premium for the term.
LtLoss Projected loss including ALAE (Economics.lae_ratio). May vary with tenure; by default it does not — see §4.
EtExpense Variable expense, including ULAE. Split acquisition (0.40) vs renewal (0.20): residual CLV charges renewal expense only, because acquisition is sunk for an in-force customer. Fixed expense is excluded from a marginal CLV by design.
Why this shape matters for filing. A single DERT × margin product cannot express a margin that varies by year. Writing it as a sum over terms can — which is what let the tenure-varying loss work land inside all six families without re-estimating anything.

§2Hand-check it on public data, with a calculator

Before any synthetic book, the identity should survive a back-of-envelope test on published industry figures. Every input below is a public number.

InputValueSource
Full-coverage personal auto premium$2,300 Industry average, full-coverage annual
Loss ratio66%NAIC personal auto, incurred
Variable expense ratio20%Renewal-expense convention
Annual retention88%Industry personal-lines retention
Discount rate8%Cost of capital (Damodaran)

Step 1 — the margin on one term.
Loss = $2,300 × 0.66 = $1,518 · Expense = $2,300 × 0.20 = $460
margin = $2,300 − $1,518 − $460 = $322

Step 2 — how many renewal terms, discounted.
v = r / (1 + d) = 0.88 / 1.08 = 0.81481
DERT(5y) = v(1 − v5) / (1 − v) = 2.8197

Step 3 — multiply.
CLVbase = $322 × 2.8197 = $907.93

The point of this exhibit. A reviewer with a calculator can reproduce $908 in about ninety seconds, from figures none of which come from us. Everything else in the paper is that calculation carried out more carefully — with retention estimated per customer rather than assumed flat at 88%, and with the horizon taken to the model's own limit rather than truncated at five years.
What is deliberately not on this page. Earlier versions of this anchor also quoted a "planning" figure that added cross-sell and upsell. It has been removed rather than updated. Review round 1 established that the old growth arithmetic was wrong (§4, c16), and a hand-check is worth nothing if it anchors to a superseded convention. The corrected growth figures are book-level and appear in §4.

§3Six families, one contract — and they agree

This is the paper's central credibility claim. Six structurally different model families, one predict_clv() contract, fitted to the same seeded synthetic book (5,000 customers · 32,728 policy terms · 5,617 claims · $103.3M written premium).

Model familyMean CLVMedian DERT (5y)1-yr in-force renewal % CLV positive
BG/NBD$2,236.27$3,004.27 2.770.805275.7%
Pareto/NBD$2,232.75$2,558.86 2.540.854675.7%
Cox PH$2,340.37$1,098.80 2.480.955552.4%
Weibull AFT$2,185.20$970.97 2.270.932252.4%
Markov$2,092.40$938.12 2.220.921952.4%
Hybrid (POG headline)$2,361.43 $1,755.782.430.9246 67.0%
12.9%
Mean CLV spread
$2,092 to $2,361 across six families
0.9221
Empirical annual renewal
The number every family is trying to match
4.65
Mean active years
69.2% of the book still in force

Read the medians, not just the means. They diverge far more than the means do, and the reason is structural rather than a defect: the survival and Markov families assign exactly zero to a lapsed customer, and 30.8% of the book has lapsed. That pushes their medians down while leaving their means close to the BTYD families. A reviewer who compares only medians will conclude the models disagree; they do not.

§4What review round 1 changed

Round 1 produced two written reviews and 24 comments, all incorporated. AJ Robinson's 13 August review left 18, of which four found defects rather than ambiguities; Mark Mondello's 17 August review left 6, which drove the v3.1 and v3.2 exposition changes and surfaced one genuine error and one genuine loose end (§4.5). Because v2.0 had explicitly changed no reported figures, v3.0 is the version where numbers moved — so the changes are itemised rather than summarised.

c16 — cross-sell was an annuity

cross_sell_value was folded into the per-term margin and multiplied by DERT, valuing a $150 cross-sale at $150 × 2.7736 per customer. Cross-sell is now a one-time event with an annual hazard, credited once at the modelled time of sale, with first-occurrence weights summing to at most 1. Upsell is separated as what it actually is: a permanent step-up in annual margin from the upgrade year onward.

Growth componentContributionTreatment
Cross-sell$98.85 One-time event, annual hazard cross_sell_prob
Upsell$277.36Permanent margin step-up
Total growth gap (v3.0)$376.22 clv_with_growth − clv_base
Pre-v3.0 figure$693.40 Overstated by the whole annuity

c39 — credibility on the wrong exposure base

Credibility was applied to customer counts against the 1,082-claim classical standard. Because the CLV segments are equal-size quintiles, this returned 0.9614 in all five rows — a column carrying no information whatsoever. On claim counts the worst-loss-ratio quintile (2,433 claims against 1,000 customers) correctly reaches full credibility. Indicated relativities did not change; the exhibit's honesty did.

c41 — the strongest new result in the paper

The retention-adjusted loss ratio grouped by tenure only, so it could not support the decision it exists for. Grouped by rating segment instead:

SegmentCustomersRetention Expected yearsFirst-term LR Lifetime LRRelativity
1 product3,2050.8375 5.110.65150.5928 1.0992
2 products1,5450.8929 6.330.59390.5132 0.9516
3+ products2500.9665 8.620.66520.5202 0.9646
The 3+ product segment has the worst first-term loss ratio of the three — and still indicates a credit. On a single-term view it ranks last; on a lifetime view it ranks second. It stays on the books 8.62 expected years against 5.11 for mono-line, and that difference outweighs a 1.4-point worse first-term loss ratio. This inversion is invisible until retention enters the exhibit, and it is the clearest thing in the paper about what lifetime value actually adds to a rate indication.

§4.5 — what the second review caught

Mark Mondello built a discrete CLV worksheet reconstructing the calculation and asked for it to be corrected where his reading missed the intent. It reproduces in our code to the cent — $1,759.73 with growth, $1,610.15 without, gap $149.58 — before any correction. It also reproduced the c16 defect independently: built from the v2.0 text, it credits cross-sell every term, with per-term probabilities summing to 1.15. Two corrections follow: growth shape (+$101.23) and a multiplicative rather than additive tenure credit (−$199.76), for a v3.2-consistent $1,661.20. His acquisition-vs- renewal expense split was already exactly right.

One red-flagged cell found a real loose end. Cash flows are quoted in day-0 real terms but discounted at a nominal cost of equity. The mismatch understates CLV, so it is conservative in the filing direction, but it should be stated rather than left implicit. Now done at v3.2 — §3.1 states the convention and limitation L14 records it. A second comment surfaced an outright error: the v3.0 global rename had overwritten the old variant names inside the paragraph explaining the rename, leaving it self-contradictory. Both are corrected or scheduled.

c14 — the variants were named backwards

Renamed clv_base / clv_with_growth. Actuaries read "constrained" as the regulatorily constrained variant, which is the opposite of the sense the July 2 decision intended. Old names ship as deprecated aliases for one release. This amends a decision this group made, and is presented as an amendment adopted on review.

§5The negative result, stated plainly

The Aug 6 poll's reserve question asked the POG to help design the NAIC five-year cross-carrier reproducibility test for the tenure gradient. That test has been attempted. It failed, and the failure is reported here rather than buried.

EvidenceResultWhat it means
Designed cross-carrier testNot executable No public source carries carrier identity and policyholder tenure in the same record. Not a data-access problem another researcher would solve.
Wisconsin LGPIF panel, tenures 0–3+19.93% R² 0.949 — apparently strong support…
Wisconsin LGPIF, full window−7.93% …until one fund-year of catastrophe losses reverses the sign. R² 0.058.
Spanish multi-product panel+1.80% Not monotone.
Schedule P between-carrier dispersion24.4% CV Against a 35.0% ten-year curve effect — same order of magnitude as the effect itself.
Consequence, already implemented. The tenure curve ships as a mechanism, not a parameter. loss_trend remains 0.0 by default and an exact no-op, asserted per model family by tests — nothing previously published moved. The paper's §3.5 "provisional" note is replaced: the POG has not ratified the treatment, but the evidence is no longer provisional.

A related caution stands from Call #4 and is worth repeating, because it inverts the obvious reading: on a revenue-neutral filing basis, recognising that loss improves with tenure makes long-tenured business carry relatively more loss. The 10+ year cohort's mean CLV falls 7.0% while younger cohorts rise 1–3%. "Loss ratio improves with tenure" and "give loyal policyholders a discount" are not the same proposition. A tenure relativity built naively from a loss-ratio-by-tenure exhibit is likely backwards.

§6How the package verifies itself

The reason to trust the numbers is not that they were checked once, but that they cannot drift without the build failing.

289
Tracked figures
274 reproduced from the toolkit · 15 cited to primary sources
0
Unmatched
Drift gate, across both documents
122
Tests passing
Also inside the release subset itself
CheckWhat it prevents
scripts/check_paper_numbers.py Any figure in the paper or executive summary that no longer matches the seeded run. Must report 0 unmatched or the build is not shippable.
Seeded generation (seed 42) Non-reproducible results. Every figure traces to one run.
Disjoint-split assertion Train/test leakage in the ML families — asserted, not assumed.
Elasticity invariant test features.SURVIVAL_COVARIATES gaining a price or rate-change covariate. This is the invariant on which §7 argues the exhibits are admissible where price optimization is prohibited.
loss_trend no-op test, per family The unvalidated tenure curve silently affecting a published number.
Release curation audit + secret scan Internal files or credentials reaching the public repository. The release is an explicit allowlist, not an exclusion list.

Regulatory screens

ScreenGroupsMin adverse-impact ratio Flagged
Geography51.0000None
Age bands50.9729None

Both clear the four-fifths threshold of 0.80 on every group. The age-band screen was added after Call #4, where it was flagged as a gap before any filing use.

§7How we judge — what "best" means, and what it means for the release

§3 showed six families agreeing on the means and disagreeing on the medians. This is the panel that says how we choose between them, because "which model is best" has no single winning number in ratemaking — raw predictive accuracy and fitness for a rate filing are different questions, and the toolkit measures both.

A. Accuracy scorecard — backtest on a customer-disjoint holdout

Computed on a 20% holdout, disjoint by customer, targeting each customer's realized discounted renewals over five years. Bold = best on that metric.

MetricWhat it asksBetterWhere the families land
Holdout MAE Average size of the miss, in discounted renewal-years lower Ensemble 0.632 · Hybrid 0.655
Validation R² Share of variance in realized value explained higher Ensemble 0.617 · Hybrid 0.560
1-yr renewal calibration Does predicted near-term retention match the empirical 0.9221? closest Ensemble 0.923 · Hybrid 0.9246 · Cox 0.9555 (over) · BG/NBD 0.8052 (under)
Rank correlation (Spearman) Do the families agree on who is high- vs low-value? higher BG/NBD vs Pareto 0.946 · vs Hybrid 0.895 · vs Cox 0.820
The subtlety that matters for pricing. For ratemaking, ranking matters more than level — you are segmenting customers, not quoting one number. A family can have modest R² and still rank customers correctly, and it is the ranking that drives the rate relativities in §3 of the paper's Chapter 6. This is also why the §5 negative result is not fatal: it says five-year CLV adds little to the current-term margin for ranking in-force customers, which is a statement about one use, not about the lens.

B. Fitness for a rate filing — the factors that actually decide

Accuracy does not pick the winner. These three override it, and they are the same three that decide what belongs in a public release.

FactorWhy it decidesEvidence in this package
Filing defensibility Can you show a regulator the mechanism? BTYD closed forms and Cox hazard ratios are transparent; a pure gradient-boosted model is not. This is why the POG chose the Hybrid as the headline — a probabilistic base for the filing exhibits, with ML correcting only the residual. §1's single identity; the deprecated-alias rename (§4, c14)
Loss-correlation Does the CLV ranking track loss experience rather than willingness to pay? A segmentation that correlates with loss cost is filing-relevant; one that ranks on price response is not, and is banned outright in several states. Loss-ratio gradient 1.68 in the bottom quintile → 0.16 in the top; the features.SURVIVAL_COVARIATES invariant test in §6
Reproducibility and governance Deterministic under set_seed(42), pinned dependencies, CI, and fast enough to refit — which is itself the answer to the NAIC AI Model Bulletin and Colorado SB21-169. §6 in full: 122 tests inside the release subset itself, drift gate 0 unmatched, curation audit, secret scan
What this means for the decision in front of you. Poll Q3 asks how the toolkit release should proceed. The three factors above are the criteria the release was built to satisfy, and all three are already evidenced inside dist_release/ rather than only in the working repository — which is the specific reason the researcher recommendation on Q3 is to provision now and let POG review run in parallel, rather than gate publication on it. The one criterion still carrying open risk is loss-correlation, and that is carried Aug-6 poll Q3: whether a tenure relativity may enter a filed exhibit at all. Every point either reviewer raised, with its disposition, is on the consolidated review-response register.

Accuracy metrics from EnsembleCLV.params_ and paper §6.6; the loss-ratio gradient from cas_clv.ratemaking.clv_rate_relativities().

§8Glossary — every term used on this page

ALAE
Allocated Loss Adjustment Expense. Carried in the loss projection here.
Adverse-impact ratio
Favorable-outcome rate for a group divided by the reference group's. Below 0.80 fails the four-fifths rule.
BG/NBD
Beta-Geometric / Negative Binomial Distribution. Buy-till-you-die model; renewal is the purchase occasion.
clv_base
Premium − loss − expense on the renewal book. The filing-defensible variant.
clv_with_growth
clv_base plus cross-sell and upsell value. Business-planning view.
Cox PH
Cox proportional hazards survival model.
Credibility
Weight given to a segment's own experience against a complement. Square-root rule against a full-credibility standard of 1,082 claims.
Cross-sell
Sale of an additional product. A one-time event; its value is the CLV of the secondary product.
CV
Coefficient of variation — standard deviation over mean.
DERT
Discounted Expected Residual Transactions. Discounted renewals expected from a customer from now on.
Filing basis
Tenure curve rescaled to a loss-weighted mean multiplier of 1.000000, so the book total is unchanged and only relativities move.
Four-fifths rule
Disparate-impact screen: a group's favorable rate below 80% of the reference group's is flagged.
Hybrid
BG/NBD base plus a machine-learning residual. The POG headline model (kickoff Q3).
In-force
Still an active customer at the observation date. 69.2% of this book.
LAE
Loss Adjustment Expense. ALAE in the loss projection, ULAE in the expense ratio.
LGPIF
Wisconsin Local Government Property Insurance Fund — a public panel used for external validation.
Loss ratio
Incurred loss over earned premium.
loss_trend
Toolkit parameter for the tenure-varying loss curve. Defaults to 0.0 and is an exact no-op.
Markov CLV
Product-portfolio states with an absorbing lapse state.
Pareto/NBD
Schmittlein (1987) buy-till-you-die model with continuous dropout.
Planning basis
Tenure curve applied unrescaled, so the book total moves. Business planning, not filing.
Relativity
A segment's indicated rate factor against the book average.
Residual CLV
Value from here forward for an in-force customer; charges renewal expense only, acquisition being sunk.
Retention
Probability a customer renews at the next term.
Schedule P
NAIC annual-statement schedule of loss development by line and year.
ULAE
Unallocated Loss Adjustment Expense. Carried in the expense ratio here.
Upsell
Coverage upgrade. A permanent step-up in annual margin, not a one-time credit.
Weibull AFT
Accelerated failure time survival model with Weibull baseline.
Reproducibility & provenance. Every toolkit figure on this page is printed by python docs/calls/phase6_package_review_data.py, which reads the same two seeded payloads the drift gate validates the manuscript against — paper/paper_numbers.json (seed 42) and paper/tenure_validation_numbers.json. If a figure is not in that output, it is not in the paper either. The §2 anchor uses only published industry figures and is hand-checkable with a calculator.

Gate as of this call: 122 tests passing · ruff clean · mypy clean · drift check 0 unmatched across 289 tracked figures · six notebooks execute clean · release subset independently gated at 96 files.

Companion materials: call brief · decision poll · RPM session deck · review response memo · L4 validation working
This research project has been funded by the Casualty Actuarial Society.