← Project hub

Superseded, 27 September 2026. This page is the record of review response to R. Kozlowski (9 Sep, paper v5.0) as presented, and keeps the figures and vocabulary of that date. In paper v5.2 (27 September 2026): the ten-year renewer is $1,093.62, breaking even in year 2 (was $782.14, year 4); the fully-allocated ten-year renewer is $590.36 (was $278.88); filing / planning basis are now the revenue-neutral / level-effect basis. The current state is on the project hub.

Response to POG Review — Ron Kozlowski's marked-up manuscript v2.0

To: Ron Kozlowski (RTK Services), Project Oversight Group cc: Geoff Werner, Mark Mondello, AJ Robinson, Aran Paik (POG); , From: Pramod Misra, Principal Researcher Review received: 7 September 2026 (CAS_CLV_Paper_v2.0_AJR+MM+RTK.docx — 93 comments dated 24–30 August, Abstract–§5.3, plus 30 tracked insertions and 15 deletions) Response: 9 September 2026 · for Call #8, 10 September Manuscript: v5.0 — the paper (59 pp) and its technical supplement (51 pp) — accompanies this memo, and every row below is checked against it as written. Register ids RK-2 … RK-91 in docs/calls/review_responses.html §3 map to your Word comment ids in order; every row below carries both.


Opening

Thank you for this. Ninety-three comments over a week is a real review, and it is the first one the paper has had from a reader who said, in the margin, exactly where he stopped being able to follow. That is the most useful thing a reviewer can do for a paper that is meant to be handed from an actuary to a modelling team, and your review is the reason v5.0 is two documents — a conceptual paper a P&C actuary can read end to end without the code or the formulas, and a technical supplement with the formulas, the full tables and the code, under the paper's own section numbers. Geoff Werner and Mark Mondello asked for the same thing on the 27 August call; your comments are what showed, sentence by sentence, where the v2.0 text failed the first reader.

Every comment has a row in the register and in the table at the end of this memo. Three of your questions changed what the paper says rather than how it says it, and I answer those in full below (#24, #95, #169). One (#112, fixed expense) is a disagreement between you and AJ Robinson, and I set out how v5.0 carries both positions and why it goes to the group on 10 September.

What changed because of your review

Global rules, applied across the whole paper, not only where you marked them:

The three framing answers, each a change to what the paper claims:

  1. The abstract states the finding positively (#24) — below.
  2. The paper names two use cases — rating decisions and valuing a book of business — and adds a subsection on the second (#95) — below.
  3. The synthetic-data paragraph leads with the reader's prior and says exactly what synthetic data permits and nothing more (#167–#171) — below.

Corrections of fact or emphasis: Proposition 103 is 1988 and is no longer "meanwhile" (#57); the recent items are Colorado SB21-169 (2021) and the NAIC model bulletin (2023). "Price optimization" is removed from the paper except in the §2 literature citations, and §3.6 is retitled The bright line: cost-based factors versus demand-based factors (#125) — you and Geoff both said on the call that this is not price optimization, and you are right. The annuity metaphor is replaced (#101) with "a stream of yearly margins, each weighted by the chance the customer is still here and discounted"; you are right that the payments are not level. Table 7, the calibration check, is redesigned so every line shows the same three metrics and the book-level statistics are labelled as such (#164/165). The 8% discount rate gets a plain paragraph: for a value estimate a higher rate is the conservative choice, three anchors are given, and the sensitivity grid reports results at 6% / 8% / 12% (#103).

#24 — "Then why write this paper?"

You read the abstract's sentence that "five-year CLV adds little to the current-term margin" as saying CLV is not worthwhile, and then asked, reasonably, why the paper exists. The sentence was reporting a finding as if it were a deficiency, and the abstract is rewritten to say what the finding actually means.

The finding is this. If the question is which of my in-force customers is worth more than which, the current-term margin already answers most of it, and a five-year CLV re-orders very few customers. That is a negative result about one use of CLV — re-ranking customers you already have — and the paper reports it because it is true and because the CLV literature outside insurance tends to assume the opposite. It is not a negative result about CLV. The economic content of a multi-year view is somewhere else, in three places the current-term margin cannot reach: acquisition, where a customer written at a first-term loss is or is not worth writing depending on what follows; retention allocation, where the value of keeping a customer is the stream of future margins the current term does not show; and valuing a book, where the present value of the in-force customers is the quantity and there is no single-term substitute. In each of those the decision changes when you look forward, and §9 shows where. The example Geoff Werner asked for on the call — two customers identical this year and different over five — is the same point in one picture, and v5.0 opens with it.

So the abstract now says: CLV's economic content is in acquisition, retention allocation and book valuation, not in re-ranking in-force customers. That is the thesis, and the earlier sentence was undercutting it by stating the negative result first and alone. Thank you for reading it the way a reader would.

#95 — valuing a book of business

You said you think of this as a valuation exercise — a book has a value for acquisition purposes, one company buying another's book — more than as a ratemaking input, and on the call you asked whether it could give the present or economic value of a book. It can, the toolkit already computes it, and the paper did not say so. v5.0 fixes that in two places.

First, early in §1, the paper names both use cases. Rating decisions are the primary one because that is what the CAS request for proposals asked for and what the filing-side exhibits serve. Valuing a book of business is the second, and it is not a stretch of the method: the same quantity, summed over the in-force customers rather than compared between them, is the present value of the book's expected future underwriting margin. On the reference book, 3,458 of the 5,000 customers are in force at the close of the observation window. Their five-year base value on the marginal view (renewal expense only, no fixed overhead; hybrid model, 8%) sums to \$13,842,714.22, or \$4,003.10 per in-force customer. That is the book valued as a going concern by its own retention, before anything a buyer would change.

Second, a new subsection, §9.6 Valuing a book of business, sets out what a buyer would change and what the toolkit exposes for it. The expense view matters more here than anywhere else in the paper: a buyer inherits the whole cost of running the book, so the fully-allocated view from #112 below is the right lens for a valuation, not the marginal view used for rating. §9.6 shows the same in-force customers on that view, charging an illustrative \$75.00 of fixed expense per policy-term through fixed_expense_per_term: \$12,973,898.83, or \$3,751.85 each, which is 93.72% of the marginal figure. The difference, \$868,815.39, is the present value of the overhead a buyer of the renewal book would inherit. The \$75.00 is not a recommendation; a real valuation takes the fixed expense per policy from the target's expense study and replaces the five-year horizon with the run-off horizon the transaction assumes, and both are single parameters. The scenario planner's book-average mode computes both views from a handful of book-level inputs, and §9.6 says so.

Two limits stated in the same subsection. The value is of the underwriting margin stream, not of the reserves or the investment income on them; a full appraisal adds those. And the synthetic book has no loss development, so the ultimate-versus-reported question you raised at #138–#140 matters more for a valuation than for a relativity, and is flagged there.

#169 — "synthetic data is better than real data?"

No, and the sentence said so by implication, which is worse than saying it outright. You and Geoff both flagged it on the call as well; it is rewritten, and the reasoning is this.

Your premise is that synthetic data is normally generated from real data, and so cannot know more than its source. That is true of synthetic data made by resampling or fitting a proprietary dataset. This book is not made that way, and the paper now says so before anything else in §4: it is generated from a structural model — each simulated customer carries a hidden retention propensity, a hidden cross-sell propensity and a hidden risk level, and premium, renewal, product additions and claims are simulated from those — and the model's outputs are calibrated to public statistics (the ISO frequency and premium bands in Table 7), not fitted to anyone's book. That design has one consequence a real dataset can never have: the true retention behind every customer is known to us, because we set it. A real book shows you whether a customer lapsed; it cannot show you the probability with which they were going to.

That single fact permits one specific check, and only that one: fit each model family to the simulated history as if it were real, ask it for the retention rate, and score the answer against the truth. That is what §4.4 does. The result is the one you then asked about at #172 and #173: the survival and Markov families recover the true renewal rate; the BG/NBD family — the reference model of the CLV literature, built for non-contractual purchasing where a customer can only drop out right after a transaction — under-predicts annual renewal by about 12 points on a contractual, annual-renewal book. It stays in the comparison because it is the standard, because its customer ranking still agrees closely with the others (0.9491 rank agreement with Pareto/NBD), and because the miss is the evidence that the choice of family matters; it is not the recommended production family, and §5 now says that plainly. On real data, that comparison would have been a difference of opinion between six models. Here it has an answer.

In every other respect synthetic data is the weaker evidence, and the rewritten paragraph says that too: the calibration bands are the only contact with the world, the tenure loss curve is a mechanism illustrated rather than an industry parameter (limitation L4), and no result in the paper is a claim about any real book. The paragraph now runs: your prior, the one thing this permits, the result, the limits — in that order.

#112 — fixed expense, and where you and AJ disagree

AJ's comment (#111) was that fixed expenses should be ignored in a customer-level CLV; yours is that someone has to pay for them, that leaving them out overstates values, and that in reality renewal customers pay the acquisition cost of newer ones. Both are right about different questions, and v5.0 carries two views rather than choosing.

Marginal CLV stays the default. A CLV supports a decision about one customer, and fixed overhead does not change with that decision; charging it would burden the decision with a cost the decision does not alter, and would make every marginal customer look worse than they are. That was the binding decision of 16 August (framework doc §2.1a), and the expense loads in the toolkit are variable expense for that reason.

A fully-allocated view is added, named, and shown beside it. For valuing a book — your #95 — or for asking whether a segment pays its way, the fixed cost has to be in, and you are right that a marginal CLV overstates the value of the book as a whole. The toolkit already has the parameter (Economics.fixed_expense_per_term, for avoidable per-policy dollars); v5.0 uses it for a named fully-allocated view, with a worked example in §3.4 valuing the same ten-year renewer both ways (\$782.14 marginal against \$278.88 fully allocated, net of acquisition, at an illustrative \$75.00 of fixed expense per term; the hand-worked version is the supplement's Appendix F.4), and one sentence of guidance: marginal for a decision about one customer or segment, fully-allocated for book valuation. Your observation that renewal customers carry the acquisition cost of new ones is exactly what the fully-allocated view makes visible, and the expense-recovery exhibit (from your July question, RK-1) is where the timing of that cross-subsidy shows.

What the choice does and does not move. The two views differ by the fixed charge in every policy term, so over the relationship they differ by that charge times the customer's expected discounted renewals. That is a level effect that grows with expected tenure. Two customers of equal expected tenure keep their order under both views; two customers of similar value can swap places where their expected tenures differ, because the longer-lived one is charged overhead on every extra term. An earlier draft of this memo said the views rank customers identically, which is true only for customers of equal expected tenure, and the paper (§3.4, §9.6) says it correctly. The relativity and segment exhibits (§6) are computed on the marginal view and are unaffected by the choice.

Loss adjustment expense. Agreed on all three points, and the toolkit has followed this convention since 16 August: allocated LAE sits inside the loss projection (Economics.lae_ratio), so the loss ratio is L&ALAE; unallocated LAE sits in the expense ratio. On whether ULAE is a fixed expense: it is treated as a variable load, because it scales with claim volume rather than with the policy count, and the fixed-expense parameter is reserved for the avoidable per-policy dollars. v5.0 defines ALAE and ULAE inline where they first appear.

Because the fully-allocated view amends a binding decision, it goes to the group on 10 September as an amendment adopted on review, not as a change made unilaterally; the register carries it as shipped-pending-ratification.

#153 / #154 — the 22% underwriting margin, stated rather than asserted

You asked what a 22% underwriting margin means and whether it is profit. It is premium minus ultimate losses minus expenses, as a share of premium, before investment income and income tax: an underwriting result, not a profit after all costs. On the flat expense convention of basis A, expenses are 0.25 of premium and the book loss ratio is 0.5246, so the margin is one minus 0.5246 minus 0.25, which is 0.2254. §4.3 of v5.0 now says what that includes — the loss amount is loss plus ALAE, ULAE sits inside the 0.25 load, and no fixed overhead is carried beyond the load — and whether it is realistic. It is not typical. The value sits near the top of its 0.02–0.30 band because the line loss ratios (0.5024 to 0.5589) run below the NAIC pure loss ratios that anchor them, so the reference book is more profitable than the industry, whose underwriting margin has been a few points of premium and negative in some years. The consequence is one of level: CLV dollar amounts on this book are higher than a typical carrier would see; the rankings, relativities and model comparisons the conclusions rest on do not depend on the level. Limitation L1 carries the margin as a calibration candidate for v4.1, with notice to the group if it moves.

One item I am not closing yet

Point by point

Ninety-three comments in ninety rows; three pairs that mark the same phrase (#66/67, #81/82,

164/165) share a row. Status vocabulary as in the register: manuscript — the paper changed;

answered — addressed without a change, reason stated; ratify — shipped, awaiting POG ratification because it amends a binding decision; open — unresolved. Everything marked manuscript lands in v5.0 unless the row names an earlier version. Where a row says Appendix, the paper's appendices are now A (the two policyholders, term by term), B (sample filing exhibit) and C (glossary); the sensitivity grid, the hand-worked exhibits and the full glossary are the technical supplement's Appendices C, F and G, and the Where column says which document.

Word # § You asked What we did Where Status
#12 (RK-2) Abstract “which” — not sure if you are talking about the toolkit or the paper or both? Rewritten so every sentence names its subject: the paper presents the framework and the toolkit implements it. No pronoun in the abstract now points at both. Abstract manuscript
#16 (RK-3) Abstract “Three design choices address regulatory constraint structurally rather than procedurally” — not sure what you are saying here? Replaced. The abstract now says the three design choices “keep the framework usable in a regulated market” and names them (two reported values; expense split by timing; a year-by-year sum with tenure-varying losses). §1.2 says why they work: the demand-based content is kept out of the filing-side exhibits by how the toolkit is built, “by construction rather than by policy”, not by a review step after the fact. Abstract; §1.2 What this paper is, and is not manuscript
#18 (RK-4) Abstract “unconstrained” — what do you mean by constrained and unconstrained? Closed at v3.0, where the pair was renamed clv_base / clv_with_growth (AJ c14, MM 22). v4.0 adds a plain gloss — “the customer as they are today” versus “plus expected future purchases”. AJ’s single-line / multi-line was not adopted because a base-variant customer may already hold three products; the gloss goes to the 10 September poll. Abstract; §3.2 Two CLV variants; Call #7 poll Q2 manuscript
#21 (RK-5) Abstract “list of ratemaking/regulatory outputs” — I am not sure how to react to this list as some seem to be assumptions while others are exhibits. Agreed — the list mixed inputs and outputs. The abstract list now names only what the toolkit produces (rate relativities, retention-adjusted lifetime loss ratios, expense-recovery exhibits, a disparate-impact screen, filing exhibits for five jurisdictions). The assumptions a practitioner sets — expense ratios, growth values, discount rate, horizon — are introduced in §3 where each is set, and the three named bases of Table 4 fix which set every exhibit uses. Abstract; §3.3 Three reporting bases manuscript
#22 (RK-6) Abstract “run against expectation” — what does this mean? The phrase is gone. The abstract states the result directly: applied so that total premium is unchanged, a measured loss-ratio improvement with tenure lowers the reported value of long-tenured customers rather than crediting it; §3.5 says why that runs against the intuition. Abstract; §3.5 Loss timing manuscript
#23 (RK-7) Abstract “applied revenue-neutrally” — applied revenue neutrality and a measured tenure loss improvement that reduces long tenured CLV rather than crediting it sound like two different things. Is this age dependent - say a 20 year old versus a 50 year old? They are two things. The abstract now says the first in plain words (“applied so that total premium is unchanged”) and §3.5 explains the second: because a long-tenured customer has already banked the improvement, the redistribution runs against them. On age: §3.5 has a paragraph “Does the curve depend on the age of the insured?” — no; the curve is fitted on tenure alone and age enters only the retention models — and the point is recorded among the limitations. Abstract; §3.5 Loss timing; §11 Limitations manuscript
#24 (RK-8) Abstract “five-year CLV adds little to the current-term margin…” — This part makes in sound like CLV isn't worthwhile but making decisions on how to allocate expenses is. Then why write this paper? The abstract was reporting the finding as a deficiency. It is rewritten to state it positively: CLV’s economic content is in acquisition, in retention allocation and in valuing a book — not in re-ranking in-force customers, where the current-term margin already carries the order. Full answer in the response memo. Abstract; §1.3 Contributions; response memo manuscript
#25 (RK-9) Abstract “All results are deterministic under seed 42.” — why is this important? A fixed random seed makes every re-run of the toolkit reproduce the identical tables; that is what the sentence was for. A different seed gives statistically equivalent, not identical, numbers, and the supplement's Appendix C shows the spread. Nobody needs to reproduce it constantly; the mention is cut from the abstract and reduced to one sentence in the reproducibility section. §12 Where the numbers come from answered + manuscript
#31 (RK-10) §1 Introduction “Pricing” — I think we should say property & casualty actuaries or does this apply to all actuaries? Property-casualty. The first sentence now says so, and “actuaries” is qualified as P&C pricing actuaries at first use. §1 Introduction manuscript
#38 (RK-11) §1 Introduction “Firestone/Hindawi five CLV measures” — assuming we will define all these terms somewhere. Each of the five measures is defined in one line where first cited, in §1, and again in the glossary. §1 Introduction; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#51 (RK-12) §1 Introduction “long “best read as the implementation…” sentence” — difficult to read … Keep sentences simple. I'd review all sentences over a certain length and work to make them simpler to read and understand. Adopted, and applied as a global rule: his tracked splits are taken as written, long sentences are split throughout, and run-on lists use 1) 2) 3) enumerators. Whole paper manuscript
#52 (RK-13) §1 Introduction “filing-aware column” — I am not sure what a "filing-aware column" is? Replaced with “a CLV column built only from cost-based inputs, so it can be shown in a rate filing” — which is what clv_base is. §1 Introduction manuscript
#53 (RK-14) §1 Introduction “§9.4” — what is this? Does this refer to section number? Yes, a section number. Adopted as a v4.0 convention across the paper: the concept is named first and the section number follows in parentheses, so a reader who has not reached that section still knows what is being pointed to. Whole paper manuscript
#54 (RK-15) §1 Introduction “literature-gap sentence” — Make this separate sentences Split into three sentences. §1 Introduction manuscript
#55 (RK-16) §1 Introduction “treatment” — be more specific The lead sentence keeps the word (“What has been missing is an implementation-ready treatment”); the four sentences that follow now say specifically what each literature supplies and lacks — valuation machinery with no link to losses (marketing), churn dynamics with no filing pipeline (actuarial retention), deployable models without regulatory defensibility (applied machine learning), and a taxonomy with no code (the 2013 CAS talk). §1 Introduction manuscript
#56 (RK-17) §1 Introduction “paper, no” — should the "," be replaced with "has"? Yes; corrected. §1 Introduction manuscript
#57 (RK-18) §1 Introduction “Meanwhile the regulatory environment has hardened: California's prior-approval regime” — What do you mean by "Meanwhile"? … Prop 103 … happened ages ago not "meanwhile"? Corrected. Proposition 103 is 1988 and is now cited as the long-standing regime; the recent items are Colorado SB21-169 (2021) and the NAIC model bulletin on AI systems (2023). “Meanwhile” is gone. §1 Introduction manuscript
#59 (RK-19) §1.3 Contributions “an actuarial annuity reading” — what is an "actuarial annuity reading" mean, what does "cohort-level reporting" mean? Did the CAS Project Oversight Group assist in the develop the three binding design decisions? I'd leave this credit to the author not the POG. Three fixes. “Annuity reading” is gone from the contributions (see RK-36 on Word #101 for where the analogy survives). Cohort-level reporting is spelled out where it appears: “results are reported at the cohort and segment level with sensitivity ranges, never as per-customer point estimates to be acted on alone”. Attribution: §1.3 says the three design decisions were “developed by the author and adopted on review by the CAS Project Oversight Group”. §1.3 Contributions; §3.1 Formal definition; §15 Acknowledgments manuscript
#60 (RK-20) §1.3 Contributions “causal structure” — is this phrase needed? No. The phrase is gone from the contributions; §1.3 says the generator mirrors an agency data model and that its realism is checked automatically, and §4.2 says in plain terms what the structure is: renewal, product growth and risk depend on each household’s own characteristics, rather than only matching averages. §1.3 Contributions; §4.2 Causal structure manuscript
#61 (RK-21) §1.3 Contributions “continuous-integration realism gate … parameter-recovery validation … only synthetic ground truth permits” — Can we use simpler terms? Yes. “An automated check that fails the build if the synthetic book leaves its calibration bands”; and “because the true retention is known, the fitted retention can be scored against it”. §1.3 Contributions; §4.4 Parameter-recovery validation manuscript
#62 (RK-22) §1.3 Contributions “Six … under one contract” — Aren't these just six different assumptions? Why is "under one contract" necessary? “Under one contract” becomes “share one interface: same inputs, same output columns”. The six are model families for retention, each with its own assumptions about how a customer leaves — that distinction is now said in the sentence. §1.3 Contributions; §5 Models manuscript
#63 (RK-23) §1.3 Contributions “native, auditable” — what is meant by "native"? “Built in.” §1.3 Contributions manuscript
#64 (RK-24) §1.3 Contributions “GDPR” — What is GDPR? Spelled out at first use — the EU General Data Protection Regulation — and added to the glossary. §1.3 Contributions; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#65 (RK-25) §1.3 Contributions “honest negative” — what does "honest negative" mean? Replaced with “a negative result we report rather than omit”. §1.3 Contributions; §9 Economic value manuscript
#66/67 (RK-26) §1.3 Contributions “scipy” — Is this a word? Are you referring to "SciPy"? Is this a commonly known term? It is the SciPy scientific-computing library, and it is now called that at first use. Lower-case package names appear only inside code listings. §1.3 Contributions; §12 Reproducibility manuscript
#68 (RK-27) §1.3 Contributions “deterministic seed-42” — Is the seed number important? what if we used a different seed - would results be the same? Do we constantly need this reproduced? A fixed random seed makes every re-run of the toolkit reproduce the identical tables; that is what the sentence was for. A different seed gives statistically equivalent, not identical, numbers, and the supplement's Appendix C shows the spread. Nobody needs to reproduce it constantly; the mention is cut from the abstract and reduced to one sentence in the reproducibility section. §12 Where the numbers come from; supplement Appendix C answered + manuscript
#72 (RK-28) §2 Literature review “direct CAS practitioner lineage” — what do you mean? Softened, using his tracked change: “a number of CAS and other articles have touched this concept before”, followed by the specific CAS items cited by name (Boucek and Conway 2003; Firestone and Hindawi 2013). §1 Introduction manuscript
#73 (RK-29) §2 Literature review “treatment” — maybe use "document" or "publication" “Publication.” §2 Literature review manuscript
#74 (RK-30) §2 Literature review “Firestone/Hindawi long sentence” — consider breaking up … I tried to change "," to ";" Split into separate sentences, which supersedes the semicolon experiment. §2 Literature review manuscript
#81/82 (RK-31) §2 Literature review “§3.5” — ??? / If I haven't read sections 3.5 and 3.1, then how do I know what this is referring to? The forward reference now names the concept — “the tenure-varying loss ratio, developed later (§3.5)” — so it reads without the section. Adopted as a v4.0 convention across the paper: the concept is named first and the section number follows in parentheses, so a reader who has not reached that section still knows what is being pointed to. §2 Literature review manuscript
#83 (RK-32) §2 Literature review “§3.1's” — ??? Adopted as a v4.0 convention across the paper: the concept is named first and the section number follows in parentheses, so a reader who has not reached that section still knows what is being pointed to. §2 Literature review manuscript
#92 (RK-33) §2 Literature review “BTYD / Pareto-NBD / BG-NBD paragraph” — Yes, please define (reply endorsing MM #91) Closed at v3.1, which defined BG/NBD, Pareto/NBD, BTYD, GBM, GLM, ML and the m/r/d notation at first use. v4.0 adds a glossary so the definitions are also in one place. §2 Literature review; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#95 (RK-34) §2.1 Taxonomy for P&C “wired into ratemaking exhibits and state-specific regulatory tests” — I understand the concept of using this into ratemaking. However I was thinking that this is more of a valuation exercise - that there is a value to the book of business for acquisition purposes (e.g., one company buying another's book of business). Adopted. v4.0 names both use cases early: rating decisions (the primary one, per the RFP) and valuing a book of business — the present or economic value of the in-force customers. A new subsection shows the book-value reading from the existing results, and the scenario planner’s book-average mode already computes it. Full answer in the response memo. §1.2 What this paper is, and is not; §9.6 Valuing a book of business; response memo manuscript
#100 (RK-35) §3.1 Formal definition “formula” — I am having a hard time understanding what this is? Is this survival probabilities, discounted? Yes, exactly that: a sum over future years of the year’s margin, each weighted by the probability the customer is still here and discounted. v4.0 puts a term-by-term reading under the formula and a two-customer numerical example in front of it. §3.1 Formal definition; §1.1 Two customers, two horizons manuscript
#101 (RK-36) §3.1 Formal definition “annuity” — When I think of an annuity I think of a payment of a given amount … for this the annuity payments may change over time … I never thought of this as an annuity. He is right that the payments vary. The formula is now introduced as “a stream of yearly margins”, each weighted by the chance the customer is still there and discounted. The annuity survives only in a paragraph headed “A loose analogy to an annuity”, which says what he said: an annuity pays a fixed amount on a schedule, these margins vary by year and by customer, and the weights are an estimated retention rather than a validated mortality table. The word is gone from the contributions. §3.1 Formal definition; §1.3 Contributions manuscript
#102 (RK-37) §3.1 Formal definition “bounded at 500 combinations” — why bounded at 500 combinations? Is this something we decide or was it a function of the parameter options? Something we decided. The 500 is the cap on the sensitivity grid (sensitivity_analysis(), eight dimensions), not the fitting routine and not a property of the data. §3.1 now says: “That bound is a design choice, not a property of the parameters. It caps the number of parameter combinations a single search will evaluate.” §3.1 Formal definition manuscript
#103 (RK-38) §3.1 Formal definition “discount rate 8%” — As a reserving actuary I think higher rates are bad and lower rates are good … for pricing/valuation a larger discount rate means less credit given. Why use 8% over a 12% in this case? A plain paragraph is added: for a value estimate a higher rate is the conservative choice, because it gives less credit to distant renewals, so a reserving actuary’s instinct runs the other way here. Three anchors bracket the 8% (Damodaran’s P&C cost of equity ≈6.1%; embedded-value practice; the CLV-literature convention of 12%), results are recommended at 6% / 8% / 12%, and the sensitivity grid produces all three (the supplement's Appendix C carries the discount-rate row), so a reader who prefers 12% has that result. §3.1 Formal definition; §10 Sensitivity; supplement Appendix C manuscript
#112 (RK-39) §3.4 Expense timing “renewal (reply to AJ Robinson, Word #111)” — Not sure I agree with AJ as someone needs to pay for fixed expenses. Leaving it out would overstate values. In reality renewal customers pay for acquisition costs for newer customers. Loss Adjustment Expenses should be considered inside the "Losses", often referred to as L&ALAE. Should we be making a distinction between Allocated Loss Adjustment Expenses and Unallocated Loss Adjustment Expenses. Would the Unallocated Loss Adjustment Expenses be considered fixed expenses? Two views rather than a winner. The marginal CLV (variable expense only) stays the default, per the binding 2026-08-16 decision, because a CLV supports a decision about one customer and fixed overhead does not change with that decision. v4.0 adds a named fully-allocated view on the existing Economics.fixed_expense_per_term parameter, with a worked example in §3.4 valuing the same ten-year renewer both ways (\$782.14 marginal against \$278.88 fully allocated, net of acquisition, at an illustrative \$75.00 of fixed expense per term), and guidance: marginal for a decision about one customer or segment, fully-allocated for valuing a book (§9.6). The two views differ by the fixed charge times expected discounted renewals — a level effect that grows with expected tenure, so customers of equal expected tenure keep their order and customers of similar value can swap places where their expected tenures differ; the relativity and segment exhibits (§6) are computed on the marginal view and are unaffected. The choice of default is before the POG on 10 September. LAE: ALAE sits in the loss projection (Economics.lae_ratio) and ULAE in the expense ratio — the toolkit convention since 2026-08-16; §3.4 defines LAE, ALAE and ULAE inline and says the fixed share of ULAE belongs in the fully-allocated view. §3.4 Expense timing; §9.6 Valuing a book of business; framework doc §2.1a; Call #8 poll, 10 September manuscript + ratify
#114 (RK-40) §3.5 Loss timing “premium-weighted log-linear” — Why "log-linear"? What would you extrapolate here? a further improvement? Explained in words: the loss-ratio multiplier is fitted as a straight line in log terms, so each tenure year applies the same percentage change (4.22%/yr in the reference book). Nothing is extrapolated beyond the observed tenure window, and the multiplier is floored at 0.70. §3.5 Loss timing manuscript
#115 (RK-41) §3.5 Loss timing “toolkit ships” — Can we use "uses"? Yes — “uses”, here and everywhere “ships” appeared in prose. Whole paper manuscript
#116 (RK-42) §3.5 Loss timing “per-year-sum paragraph” — Can you say this in simpler terms - maybe make this more intuitive to the simple reader. Rewritten as one plain statement — CLV is the sum, year by year, of (chance of still being here) × (that year’s margin), discounted — with one hand-worked line in the supplement's Appendix F. §3.5 Loss timing; supplement Appendix F Every exhibit, worked by hand manuscript
#117 (RK-43) §3.5 Loss timing “…its observed 0.42” — Does the .42 depend on age of the insured? No. The synthetic book carries no age covariate in the loss curve (age enters only as a survival covariate); this is now stated where the figure appears and listed as a limitation. §3.5 Loss timing; §11 Limitations manuscript
#118 (RK-44) §3.5 Loss timing “§3.2” — Best to refer to a term not a section number Adopted as a v4.0 convention across the paper: the concept is named first and the section number follows in parentheses, so a reader who has not reached that section still knows what is being pointed to. Whole paper manuscript
#121 (RK-45) §3.5 Loss timing “bases paragraph” — How would I convert the filing basis to the planning basis? Not sure I am following this. Both bases use the same fitted curve with a different normalisation: the revenue-neutral basis divides every multiplier by the book’s loss-weighted mean so the book total is unchanged; the level-effect basis applies the curve as fitted. Converting is one switch (tenure_loss_normalize), and §3.5 now has a paragraph “Converting one basis to the other”: divide by the mean to go from level-effect to revenue-neutral, multiply to go back, with the two book totals (\$16,111,118 and \$18.73M) as the check. Both bases by cohort are Figure 7 and the supplement's Appendix F.6. The bases were renamed at v3.3 (R2-1). §3.5 Loss timing; Figure 7; supplement Appendix F.6 manuscript
#122 (RK-46) §3.5 Loss timing “put to the Project Oversight Group and not yet ratified.” — Do Mark and AJ follow this? Open. The basis names (revenue-neutral / level-effect) are Call #7 poll Q1, carried to the 10 September call for ratification; §3.5 states that the names “were put to the Project Oversight Group and are on the poll for the next call”. §3.5 Loss timing; Call #7 poll Q1; Call #8, 10 September ratify
#125 (RK-47) §3.6 Bright line “price optimization” — Is this really "price-optimization"? No, and Werner and Kozlowski said the same on Call #7. v4.0 removes the term everywhere except the §2 literature citations (CAS Price Optimization Working Party 2014; NAIC 2015); §3.6 is retitled “The bright line: cost-based factors versus demand-based factors”, and a boxed “What this paper is, and is not” is added early in §1. §3.6 The bright line; §1.2 What this paper is, and is not manuscript
#128 (RK-48) §4.1 Design goal and schema “for all published results” — can we drop this? Dropped. §4.1 Design goal and schema manuscript
#129 (RK-49) §4.1 Design goal and schema “sync” — Is there another word besides "sync"? Maybe "uses"? “Uses.” §4.1 Design goal and schema manuscript
#132 (RK-50) §4.1 Design goal and schema “Reference-book grain” — Is "grain" another word for "detail"? Yes — replaced with “level of detail”. §4.1 Design goal and schema manuscript
#133 (RK-51) §4.1 Design goal and schema “behavioral segment” — what is meant by "behavioral segment"? Defined in §4.1: a label for the kind of customer a household stands for (price-sensitive, loyal multi-product, or growing business), grouping customers meant to behave alike. In the reference generator the label is descriptive — the household’s outcomes are driven by the three latent fields (retention propensity, cross-sell propensity, risk factor), not by the label — and it is never a rating variable. §4.1 Design goal and schema; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#134 (RK-52) §4.1 Design goal and schema “latent retention/cross-sell propensity, latent risk factor” — Can you give more details on what this is? Explained: each synthetic customer carries hidden values for how likely they are to renew, to buy another product, and how risky they are. Those values drive the simulated outcomes but never appear in the output tables — exactly as in a real book. §4.1 Design goal and schema; §4.2 Causal structure manuscript
#135 (RK-53) §4.1 Design goal and schema “coverage tier 1–3” — I was thinking years 1-3 but these are three different grouping. Can we give example like we do for product line? Example added in §4.1: a coverage tier is a level of coverage richness, not a year — tier 1 is the basic coverage a household starts with, tiers 2 and 3 are progressively richer coverage, each step up carrying a higher premium; the three product lines are named as the analogous example. §4.1 Design goal and schema manuscript
#136 (RK-54) §4.1 Design goal and schema “one per claim” — We have "one per household" and "one per claim". Should then we have "one per policy"? … a record for each customer, LOB and year? Yes. There are three tables and the text now names all three: one row per customer, one row per policy term (customer × product × year), and one row per claim. §4.1 Design goal and schema manuscript
#137 (RK-55) §4.1 Design goal and schema “maturity” — Maturity means age since reporting? Why are some fields in a different font? Defined: maturity is the number of years since the claim occurred, measured at the observation date — not since reporting. The monospace font marks a field name in the data schema; §4.1 states that convention once, up front. §4.1 Design goal and schema manuscript
#138 (RK-56) §4.1 Design goal and schema “incurred vs reported” — Isn't the reported losses the given and then incurred losses the calculated? Is "ultimate" a better term as incurred can be mistaken as "reported"? His terminology is adopted: “ultimate” for the calculated, fully developed amount and “reported” for the as-of case amount. “Incurred” is no longer used as a basis name; the column keeps its schema name incurred_loss for compatibility, and the text says so. §4.1 Design goal and schema; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#139 (RK-57) §4.1 Design goal and schema “loss_basis="ultimate"|"reported"” — Isn't this an LDF factor? Yes. The divisor that links the two is a loss development factor and is now called that. §4.1 Design goal and schema manuscript
#140 (RK-58) §4.1 Design goal and schema “parameter choice, not a rework” — The ultimate is a calculation, not a given - unless we are tracking as of dates and showing the development over time. Agreed. §4.1 now says the reported amount is the observed quantity and the ultimate is the calculation; the synthetic book runs it in reverse (draws the ultimate, divides by the factor at the claim’s maturity), holds one observation date and no history of development — recorded with the book’s other calibration limits (L1). §4.1 Design goal and schema; §11 Limitations (L1) manuscript
#141 (RK-59) §4.1 Design goal and schema “personal and commercial lines” — "what"? data? records? … written more clearly. Rewritten: “the book contains policy records for three product lines”. §4.1 Design goal and schema manuscript
#142 (RK-60) §4.1 Design goal and schema “cover” — Is the POG and RFP two separate groups? Can we just say that the personal lines are composed of personal auto and homeowners and the commercial account is the business-owners line. His sentence is adopted as written. The RFP is the CAS request for proposals and the POG is the oversight group; neither needed to be in that sentence. §4.1 Design goal and schema manuscript
#144 (RK-61) §4.2 Causal structure “A generator that merely matches marginal distributions produces CLV variation that is pure claim noise.” — How about "The synthetic data is generated to match the mean with some variation from account to account." His sentence opens the paragraph; the next sentence then says why matching the mean alone is not enough (the customers would differ only by claim luck). §4.2 Causal structure manuscript
#145 (RK-62) §4.2 Causal structure “Risk-based rating with partial rating efficiency” — ??? Replaced with “premium tracks the customer’s true risk, but not perfectly — as in a real rating plan”. §4.2 Causal structure manuscript
#146 (RK-63) §4.2 Causal structure “∝ risk^0.8” — ??? Explained in words: premium is proportional to the risk factor raised to the power 0.8, so it rises with risk but less than proportionally — a household ten percent riskier than average pays about eight percent more — and the exponent is named in the same sentence as the rating efficiency (no footnote). §4.2 Causal structure manuscript
#147 (RK-64) §4.2 Causal structure “Premium–risk correlation 0.869–0.878” — ??? Explained: how closely premium tracks true risk within each line, where 1.0 would be perfect rating. The three values are labelled by line: 0.8728 personal auto, 0.8776 homeowners, 0.8685 commercial. §4.2 Causal structure manuscript
#148 (RK-65) §4.2 Causal structure “Renewal 0.9569 at tenure 4–8 vs 0.8923 at tenure ≤1” — What is it from tenure of <1 and </=4 The two figures were the generator’s two asserted checks, not a full table, and §4.2 now says what each band is: tenure 0–1 is the first two policy years (renewal 0.8923); tenure 4–8 is the fifth through ninth policy years (0.9569). Tenures 2–3 and 9+ are not separately reported; because the mechanism is a straight line (+1.2 points per year of tenure, level after ten years), their renewal rates lie between the two figures. §4.2 Causal structure manuscript
#149 (RK-66) §4.2 Causal structure “annual product-addition hazard 0.12 × propensity; tier-upgrade hazard 0.07” — Explain more clearly Rewritten: each year a customer has a 12% chance, scaled by their own propensity, of adding a product, and a 7% chance of moving up a coverage tier. §4.2 Causal structure manuscript
#150 (RK-67) §4.2 Causal structure “raise coverage” — you mean "increases the coverage amount"? or … raise coverage amounts and purchase new products? “Increase the coverage amount (a tier upgrade)”. Buying a new product is the separate cross-sell event, and the two are now named separately every time. §4.2 Causal structure manuscript
#151 (RK-68) §4.2 Causal structure “mid-relationship additions present” — ??? Replaced with “customers add products part-way through the relationship, not only at inception”. §4.2 Causal structure manuscript
#152 (RK-69) §4.2 Causal structure “renewal probability rises 3.5 pts per product beyond the first” — Does this mean that the multi-product is 3.5% higher than the monoline Yes, in percentage points and per additional product: a two-product customer renews 3.5 points more often than a mono-line one, before other effects. The observed book figures (0.9097 mono-line vs 0.9443 multi-product) are given alongside. §4.2 Causal structure manuscript
#153 (RK-70) §4.2 Causal structure “underwriting margin ≈22% of premium” — can you explain this? Does this mean profit? Stated in the manuscript. The underwriting margin is premium minus ultimate losses minus expenses, as a share of premium, before investment income and income tax — an underwriting result, not a profit after all costs. §4.3 says what it includes (loss is loss plus ALAE; ULAE sits inside the 0.25 load; no fixed overhead beyond the load) and whether it is realistic: the value sits near the top of its 0.02–0.30 band because the line loss ratios (0.5024–0.5589) run below the NAIC pure loss ratios that anchor them, so the reference book is more profitable than the industry, whose underwriting margin has been a few points of premium and negative in some years. The consequence is level, not order: CLV dollar amounts on this book are high; rankings, relativities and model comparisons are unaffected. L1 carries it as a v4.1 calibration candidate, with POG notice if it moves. §4.2 Causal structure; §4.3 Calibration bands; §11 Limitations (L1) manuscript
#154 (RK-71) §4.2 Causal structure “portfolio margin 0.2254 at flat 0.25” — What does this mean? Rewritten in words in §4.2: on the flat expense convention of basis A, expenses are 0.25 of premium; the book loss ratio is 0.5246, so the margin is one minus 0.5246 minus 0.25, which is 0.2254, about 22% of premium. Composition and realism as RK-70 (Word #153); the v4.1 calibration note stands. §4.2 Causal structure manuscript
#156 (RK-72) §4.3 Calibration bands “Realism is not a documentation claim; it is a test suite.” — Not sure what this means? Replaced: “Realism is checked automatically, not asserted” — a test generates the book, compares every statistic against its band, and the build fails if any leaves it. §4.3 Calibration bands manuscript
#157 (RK-73) §4.3 Calibration bands “CALIBRATION_TARGETS sentence” — This is an example of a hard to read sentence. Split into three sentences. §4.3 Calibration bands manuscript
#158 (RK-74) §4.3 Calibration bands “liability-mixed” — what is "liability-mixed?" Defined where it appears: the NAIC 2022 personal-auto expenditure (\$1,127) is liability-mixed — it averages liability-only policies with full-coverage ones, which is why it sits below a full-coverage band. §4.3 Calibration bands manuscript
#159 (RK-75) §4.3 Calibration bands “ISO all-coverage ≈0.114/vehicle-year × household vehicle counts. Implies 1.768 cars per household.” — what if household has fewer or more cars? Does analysis still work? The generator does not model vehicles, and the paper no longer implies it does. Claims are drawn per household policy-year at a rate scaled by the household’s risk factor, so frequency is household-level; the ISO anchor is per vehicle, and a household frequency above it reflects households that insure more than one car on average. The tracked “1.768 cars” insert is not adopted, because no vehicle count exists to average. A one-car or three-car household is handled the same way — its value is computed from its own claim history — so the analysis holds for any count. §4.3 Calibration bands manuscript
#161 (RK-76) §4.3 Calibration bands “current market $2,490” — why so much higher? implications? The \$2,490 is the current-market homeowners figure. The band is anchored to the NAIC 2022 average (\$1,569) and the book’s homeowners premium (\$1,598.49) sits at that level; the current-market figure reflects recent rate increases the book does not model. Implication stated: CLV dollar levels move with premium; the rankings and relativities the ratemaking exhibits use do not. §4.3 Calibration bands manuscript
#162 (RK-77) §4.3 Calibration bands “0.2435” — is this so much higher than personal because there are multiple vehicles or higher frequency/miles? Because commercial frequency is per policy-year at the account level, and a business-owners account covers several exposures at once — premises, vehicles and liability — where homeowners covers one dwelling. The commercial rate (0.2435) is therefore close to the household auto rate (0.2016), not to homeowners (0.0777). Stated next to the figure. §4.3 Calibration bands manuscript
#163 (RK-78) §4.3 Calibration bands “multi-product share” — what is multi-product share? Defined: the share of households that held more than one product line at any time in the window (0.3590 in the reference book). §4.3 Calibration bands; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#164/165 (RK-79) §4.3 Calibration bands “underwriting margin / premium” — Why does personal auto, homeowners and commercial have avg premium and claim frequency while "portfolio" has annual retention, multi-product share and u/w margin? Why does this exist for "Portfolio" but not other lines? Table 7 is redesigned so every line row shows the same three metrics (average premium, claim frequency, loss ratio) and the whole-book column carries the loss ratio plus the three statistics that exist only at book level (retention, multi-product share, underwriting margin), labelled as such; a paragraph “On the dashes” says why retention is not per line and why no whole-book premium or frequency is computed. §4.3 Calibration bands, Table 7 manuscript
#167 (RK-80) §4.4 Parameter-recovery validation “ground truth is latent-but-known” — what does this mean? Replaced: “the true retention rate is hidden from the model but known to us, because we generated it”. §4.4 Parameter-recovery validation manuscript
#168 (RK-81) §4.4 Parameter-recovery validation “synthetic data permits a validation real data cannot…” — I have issues with the sentence structure and meaning. The paragraph is rewritten to lead with the reader’s prior (synthetic data is usually the weaker evidence), then the one thing it permits (true retention is known, so recovered retention can be scored), then the result. §4.4 Parameter-recovery validation manuscript
#169 (RK-82) §4.4 Parameter-recovery validation “same sentence” — synthetic data is better than real data? How can it be when synthetic data is normally generated from real data? It is not better, and the sentence implied it was. Synthetic data permits one specific check that real data cannot: the true retention behind each customer is known, so the model’s recovered retention can be scored against it. Everything else about it is weaker. Rewritten as above; full answer in the response memo. §4.4 Parameter-recovery validation; response memo manuscript
#170 (RK-83) §4.4 Parameter-recovery validation “CI-gate BG/NBD sentence” — Can you write this more clearly? Rewritten in two short sentences (see RK-84). §4.4 Parameter-recovery validation manuscript
#171 (RK-84) §4.4 Parameter-recovery validation “CI gate” — What is "CI gate"? Defined at first use — “continuous integration, the automated test run on every code change”; if any statistic leaves its band, the build fails — and the abbreviation is dropped from prose. §4.4 Parameter-recovery validation; Glossary (supplement Appendix G; core terms in paper Appendix C) manuscript
#172 (RK-85) §4.4 Parameter-recovery validation “0.9246” — Why are the renewals so different between these families? Answered in the section: the families assume different things about how a customer leaves. BG/NBD is a non-contractual model — it assumes a customer can only drop out right after a purchase — so on an annual-renewal book it under-reads retention; the survival and Markov families read the contract structure directly. §4.4 Parameter-recovery validation manuscript
#173 (RK-86) §4.4 Parameter-recovery validation “BG/NBD under-predicts by about 12 points” — Doesn't this imply that this would not be a good family to use? Yes for retention level, and the paper now says so plainly: BG/NBD is not the recommended family for estimating the level of retention (L5). It stays in the comparison because it is the reference model of the CLV literature, its customer ranking still agrees closely with the others (0.9491 rank correlation with Pareto/NBD), the 12-point miss is the evidence that the family choice matters, and it is the base the headline hybrid corrects. §4.4 Parameter-recovery validation; §5 Models; §11 Limitations (L5) manuscript
#175 (RK-87) §4.5 Public datasets “(~13 MB committed)” — what does MB relate to? Is records a better term? Yes. File sizes are replaced with record counts. §4.5 Public datasets manuscript
#177 (RK-88) §5 Models “heading” — This section is beyond my abilities to review. Noted, and it shaped v5.0’s structure: the model families themselves (§5.1.1–§5.1.4) are in the technical supplement, and the paper’s §5 keeps one plain subsection, 5.1 Four ways to estimate how long a customer stays, then the hybrid (5.2) and the model comparison (5.3); §1.5 Where the detail lives says which document holds what. §1.5 Where the detail lives; §5 Models; supplement §5.1 answered + manuscript
#178 (RK-89) §5 Models “All six families implement one contract” — Maybe "All six families (of assumption?) are run through the model." Rewritten: “All six families share one interface: the same inputs and the same output columns” (see RK-22 on Word #62 for why they are families, not assumptions). §5 Models manuscript
#180 (RK-90) §5 Models “numpy/scipy” — ???? “The NumPy and SciPy scientific-computing libraries” at first use. §5 Models manuscript
#184 (RK-91) §5.3 Survival models “coefficient list” — should these terms be shown consistently for all families or do the different families require different assumptions? The families genuinely differ — the buy-till-you-die models have four shape parameters, the survival models a coefficient per covariate, Markov a transition matrix, the hybrid a validation score — so the lists cannot be identical. §5 says so in a paragraph “How the fitted values are reported”, and each subsection reports its values in the same order: the parameters, their meaning, then the value on the reference book. Table 10 explains the two survival scales (Cox on the log-hazard, AFT on log survival time). §5 Models; supplement §5.1.3 Survival models, Table 10 manuscript

What you did not review, and an invitation

Your last note, on §5, was that the section was beyond your abilities to review, and §6–§15 are unmarked. I do not read that as a gap in the review; I read it as the clearest evidence the paper has had that the v2.0 register was wrong for its first audience, and v5.0 moves the model families themselves into the technical supplement for exactly that reason; the paper's §5 is one plain subsection on the four ways to estimate how long a customer stays, then the hybrid and the comparison.

The unreviewed sections are where a second pass from you would be worth the most, because they are the ones closest to your questions: §6 (the rate relativities and the retention-adjusted loss ratio — the ratemaking payoff of the expense split you asked for in July), §7 (the regulatory tests and the bright line, now without the phrase "price optimization"), §8 (cohort results), §9 (economic value, including the new §9.6 on valuing a book), §11 (limitations, including the one your ex-ante question on the call produced, L16) and §12 (Where the numbers come from, where the seed sentence now lives). If you read only two, read §9.6 and §6.2 — the first is your use case and the second is the exhibit your July question built.

The v5.0 paper and technical supplement come by email with the planner link and Mark's reconciled worksheet in one message, as the group asked. Comments on any of it, in any form, go on the register with the same treatment as these ninety rows.

Pramod

This research project has been funded by the Casualty Actuarial Society.