Machine learning for IBNR

Choosing reserving
methods by how well
they predict

Several methods, and many reasonable ways to apply each, give different reserves for the same triangle. The development that would settle the choice has not happened yet.

Learn how machine learning uses past claims development to test familiar methods for IBNR reserving in property and casualty (P&C), or non-life, insurance.

This guide explains published research. The paper is not ReserveAI: its out-of-time AvE and CDR concepts are stepping stones to a much broader platform. See what ReserveAI adds.

Start with the example Read the executive summary
Cumulative claims by accident and development period
Claims experience by calendar periodEach row is an accident period and each column is a development period. Filled cells are known. The orange diagonal is the next period to predict. Outlined cells are later, unobserved development.Development period →Accident period →1122334455667788991010
Period 10

The orange diagonal is the next period to forecast.

KnownNext to predictStill unknown

Executive summary

What a chief actuary needs to know

The opportunity is a more consistent, evidence-based way to select reserving methods. Familiar actuarial techniques remain at the centre; historical prediction provides a repeatable basis for choosing how to apply them.

Make method selection repeatable
Treat the method and its assumptions as candidates to be tested. Replay earlier valuations using only information available then, score each candidate’s subsequent predictions, and use those scores to inform today’s selection. This makes the reasons for a choice easier to compare, document and challenge.
Agree what success means
Actual versus expected (AvE) measures next-period prediction error. The claims development result (CDR) measures the revision in estimated ultimate claims. They can select different methods. Decide which outcome matters for the portfolio, and review both: a stable ultimate estimate can still be wrong.See AvE and CDR explained
Use the evidence in context
The paper tests three triangles. Some selections improve ultimate-claims prediction; others make little difference or perform worse. The results support testing this approach on your own portfolio. They do not establish a universal best method or quantify a capital release, reserve reduction or operating-cost saving.
Keep actuarial judgement central
The actuary still owns the data, candidate assumptions and final reserve. Changes in inflation, claims handling or business mix can weaken the relevance of historical scores. Review those changes, investigate material differences from the current approach, and record the reasons for any override.

A practical next step

Commission a controlled pilot on one portfolio. Agree the candidates and scoring criteria in advance, keep later periods separate for evaluation, and compare with the existing process. Review prediction error, reserve movements and the work required before deciding whether to extend the approach.

Review the evidence

A reserve is also
a prediction

Depending on the method, choices include the history used for development factors, the loss-ratio assumption and how much credibility to give the data. Each choice changes what the method expects to happen next.

We can test those choices by replaying earlier valuations. Fit using only the information available at the time, freeze the prediction, then compare it with the claims development that followed. This is out-of-time validation.

The first prediction

Forecasting the next diagonal
before it is known

The example below fits a chosen method to the known cells of a triangle, freezes its forecast for the next diagonal, and then reveals what was actually observed. Nothing from the future goes into the fit.

A working example · illustrative data
1Make a forecast
2Reveal experience
3Refit & compare
What is known at the end of2021
Development period1234567820171002017, development 1: observed 100.001552017, development 2: observed 155.001812017, development 3: observed 181.001942017, development 4: observed 194.002012017, development 5: observed 201.002012017, development 6: forecast 201.00; actual still hidden2017, development 7: not revealed2017, development 8: not revealed20181152018, development 1: observed 115.001742018, development 2: observed 174.002042018, development 3: observed 204.002192018, development 4: observed 219.002272018, development 5: forecast 226.90; actual still hidden2018, development 6: not revealed2018, development 7: not revealed2018, development 8: not revealed20191292019, development 1: observed 129.001902019, development 2: observed 190.002202019, development 3: observed 220.002362019, development 4: forecast 236.00; actual still hidden2019, development 5: not revealed2019, development 6: not revealed2019, development 7: not revealed2019, development 8: not revealed20201452020, development 1: observed 145.002142020, development 2: observed 214.002492020, development 3: forecast 249.46; actual still hidden2020, development 4: not revealed2020, development 5: not revealed2020, development 6: not revealed2020, development 7: not revealed2020, development 8: not revealed20211572021, development 1: observed 157.002352021, development 2: forecast 235.34; actual still hidden2021, development 3: not revealed2021, development 4: not revealed2021, development 5: not revealed2021, development 6: not revealed2021, development 7: not revealed2021, development 8: not revealed20222022, development 1: not revealed2022, development 2: not revealed2022, development 3: not revealed2022, development 4: not revealed2022, development 5: not revealed2022, development 6: not revealed2022, development 7: not revealed2022, development 8: not revealed20232023, development 1: not revealed2023, development 2: not revealed2023, development 3: not revealed2023, development 4: not revealed2023, development 5: not revealed2023, development 6: not revealed2023, development 7: not revealed2023, development 8: not revealed20242024, development 1: not revealed2024, development 2: not revealed2024, development 3: not revealed2024, development 4: not revealed2024, development 5: not revealed2024, development 6: not revealed2024, development 7: not revealed2024, development 8: not revealed
Known experienceFrozen next forecastUnknown

Cumulative paid claims · illustrative units. Cells beyond the valuation date are hidden from fitting.

The unknown
next diagonal

The chosen method is fitted to the teal cells, and its predictions for the next period are outlined in orange. These predictions are now frozen.

Expected next development137.7

Changing the method restarts the example.

Inspect the calculation and assumptions

Scores use the paper's absolute-increment-weighted RMSE, across the same existing origin periods. The signed total is shown separately; it is not the RMSE. Development ages with no observed pairs use a factor of 1 in this teaching implementation. There is no added tail beyond the final displayed age.

Next-diagonal calculation, by origin year. Future observations appear after reveal.
OriginExpectedActualAvECDR after refit
20170.00Hidden
20187.90Hidden
201916.00Hidden
202035.46Hidden
202178.34Hidden

Machine learning applied
to method selection

Chain Ladder, Bornhuetter–Ferguson and Generalised Cape Cod remain recognisable. Machine learning gives us a systematic way to choose between their many possible applications.

Read about the three methods
1

The choices behind a candidate

A method, a history window and a loss-ratio assumption together define one candidate.

2

The same test for every candidate

Earlier valuations are replayed, and each candidate predicts the next diagonal before it is revealed.

3

A choice made on the historical scores

The historical scores are compared, and the selected method is then checked against the development that came later.

AvE & CDR · A year in the life of a reserve

The year that happened.
The future we now expect.

A new diagonal tells us what happened. It also changes what we expect to happen next. AvE isolates the first effect. CDR brings both into the same result.

AvE · Actual versus expected

Did this year’s claims match our forecast?

Compare the new experience with the prediction frozen at the previous valuation.

CDR · Claims development result

How has the total estimated cost changed?

Include both the new experience and the revised estimate of everything still to come.

  1. 01 · At the opening valuationFreeze the forecast

    Split the reserve into what is expected next year and what is expected after that.

  2. 02 · Over the next yearObserve the payments

    The difference from the frozen next-year forecast is the AvE deviation.

  3. 03 · At the closing valuationRe-estimate the remaining claims

    Use the enlarged triangle to revise the future outlook for those same accident years.

Try the two effects independently

Two questions about one year of development

An illustrative paid-claims example, in undiscounted claims units. Track the same prior accident years at both valuations; exclude the new accident year.

Start from a scenario
Opening reserve · the frozen forecast100
30
70
Expected payments this yearExpected payments after this year
One year passes. Payments arrive. The remaining claims are re-estimated. ↓
Payments + closing reserve96
36
60
Actual payments this yearNew estimate of later payments

Earlier payments are common to both valuations, so they cancel. Compare the opening reserve with payments plus the closing reserve, not with the closing reserve alone.

Change payments, then change the future estimate. These controls isolate the arithmetic; they are not fitting a reserving model.

The bridge below uses this guide’s sign: positive CDR means a higher ultimate.

AvE+63630
+
Remaining-reserve revision−106070
=
CDR−4Updated − previous ultimate

The signed components offset: +6 and −10 combine to −4. Square -4 when scoring CDR, not the two components separately.

Favourable development

The total estimated cost falls by 4. That movement includes the payment surprise and the change in the future outlook.

Same development. Two sign conventions.

Switch the convention. The underlying claims and the economic outcome stay the same.

Show CDR using
Opening reserve − payments − closing reserve100 − 36 − 60 = +4

Favourable developmentPrevious ultimate − revised ultimate

Merz–Wüthrich: observable CDR, Definition 2.5, equation (2.19), printed page 552. This guide follows the opposite sign in Balona–Richman §2.2. AvE here remains actual minus expected, in claims units rather than a ratio.

The one-year insight

One year of learning can change many years of claims.

Merz and Wüthrich’s 2008 paper studies the change between successive ultimate estimates. A one-year view includes the revaluation of the remaining run-off, even when settlement is many years away.

Read the two signals together

A quiet CDR can hide a noisy year.

Higher payments can be offset by a lower future estimate. Or payments can match the forecast while the future estimate rises. Check the components before interpreting a small total as evidence of a good forecast.

From a realised CDR to one-year reserve risk

After the year, CDR is an observed result. Before the year, it is uncertain. Under distribution-free chain-ladder assumptions, Merz and Wüthrich estimate the conditional mean squared error of predicting the observable CDR by zero.

That prospective MSEP differs from the historical weighted RMSE used to select candidates in this guide. A realised CDR, or its square, is not by itself a capital requirement or a full predictive distribution.

Read §3, equation (3.1) · printed page 554
The evidence

Results on
three published triangles

The paper tests selection on three triangles with different development patterns and levels of volatility. In some cases selection reduced the error of the forecast ultimate, and in others it made little difference or made the forecast worse. Both kinds of result are shown below.

Ultimate prediction RMSE · lower is betterBasic GCC3,170.88Selected by AvE2,552.39Selected by CDR2,893.2301,0002,0003,000
Ultimate-claims RMSE on later development

Both selections gave a lower RMSE than the baseline

AvE-selected BF improves ultimate RMSE by 19.5%; CDR-selected BF improves it by 8.8%.

−19.5%Selected by AvERMSE versus baseline
−8.8%Selected by CDRRMSE versus baseline
Read the reported values and their scope

These are transcribed paper results, not outputs from this website's replay. RMSE is measured against the paper's final development values, in source triangle units. A smaller reserve is not itself a better prediction.

Quarterly long-tail liability
SelectionMethodUltimate RMSEChange from baseline
Basic GCCGCC3,170.880.00%
Selected by AvEBF2,552.39−19.51%
Selected by CDRBF2,893.23−8.76%

Source Balona and Richman (2021) manuscript, table lt_tuning_scores. How this site uses the paper

Read the evidence and its limits
The chapters

The guide
in seven chapters

The chapters can be read in order, or entered at the one that answers a particular question. Each chapter gives the intuition, the calculation and the passage of the paper it rests on.

First page of the paper, The Actuary and IBNR Techniques: A Machine Learning Approach
The original paper

The paper,
section by section

The walkthrough states what each section contributes, links its notation to the visual explanations on this site, and gives access to the original documents.

Caesar Balona & Ronald Richman

Paper awarded the IFoA’s 2020 Brian Hey PrizeWalk through the paper Explore the early sources
The original early draft, headed The Actuary and IBNR Techniques, opening with the problem of choosing a reserving methodOpen the original
Original early draft · Undated supplied copy · 7 pages

Before the full paper

The early draft

The early draft puts the question plainly: with so many reserving techniques available, how should an actuary choose? Its answer begins with replaying earlier decisions and examining the experience that followed.

The idea dates to 2017. The early draft is undated; the paper is the joint work of Caesar Balona and Ronald Richman.

ReserveAI · by insureAI

Making machine-learning reserving
work in practice.

The principle is published. The practical advantage is built into ReserveAI.

ReserveAI’s patent-pending technology makes these ideas work in day-to-day reserving. Our additions are central to the methodology and dramatically improve its practical usefulness.

Discover ReserveAI