Rebuilding a Published Budget Impact Model: Venetoclax in Acute Myeloid Leukaemia

A published budget impact model in AML, rebuilt from its parameters and audited against its own sources

Author

Xiaoge Zhang, PhD

Published

September 17, 2026

WarningThis is a demonstration, not a deliverable

A real budget impact model is an Excel workbook — that is what a payer or an HTA agency receives, and what I would build. This page is not one.

It shows the structure instead: every parameter names its source, every calculation shows its working, every assumption is labelled.

It uses no confidential or commercial data. Every parameter comes from one published, openly licensed study:

Palacios A, Espinola N, Gonzalez JM, Rojas-Roque C, Rivas MM, Kanevski D, Morisset P, Augustovski F, Pichon-Riviere A, Bardach A. Budget impact analysis of venetoclax for the management of acute myeloid leukemia from the perspective of the social security and the private sector in Argentina. PLoS One. 2024;19(1):e0295798. doi:10.1371/journal.pone.0295798 — CC-BY 4.0.

It is not a validated reimbursement model and not a pricing recommendation. Market shares in the source are projections supplied by the manufacturer and validated by a Delphi panel; they are not observed market data.

What this rebuild found

Rebuilding the model from its parameters — rather than restating its results — surfaced three things the paper does not report.

  1. The supplement contradicts itself twice — two market-share rows are transposed between tables, and two tables disagree on dosing days. Resolved by reconciling each reading against published drug costs, and against the label and the trial.
  2. The mechanism driving hospitalisation is not in the paper — it is in the Delphi questionnaire in the supplementary material. Resolved by reading the questionnaire’s own wording.
  3. Two cost components do not reconcile — administration ≈45% low, monitoring ≈12% low, against unit costs the source never publishes. Left open in the reconciliation panel, not closed by fitting.
NoteVerdict

The published headline holds — year-3 budget impact rebuilds to within 0.1% under both payer perspectives. But it holds for a reason the paper does not state: the model keeps every patient alive for a full year, and the comparator arm has the worse survival, so it overstates the comparator’s cost more than venetoclax’s — compressing the difference between the two worlds. That limitation runs against the sponsor, not for it.

How this was audited

Five steps, in order. The same sequence works on any model you did not build.

  1. Trace every parameter back to a specific table, and label its source type. Inputs with no provenance are invisible until you try to write the citation down.
  2. Reconcile against every published total, not just the headline. Errors that cancel in the aggregate survive a headline check.
  3. Adjudicate contradictions with a third source, and quantify both readings — otherwise you pick the more plausible one and have no way of knowing you picked wrong.
  4. Trace mechanisms the text omits into the appendices and questionnaires. Undocumented links from a clinical result to a cost live there, not in the Methods.
  5. Hold the calibration line: estimate unpublished parameters; never scale whole components to match a total. Fitting hides the very quantity the model exists to report.

The same five steps are applied to a different model type — a four-state Markov cost-effectiveness model — in Rebuilding a Published Hypertension Markov Model, there set against the TECH-VER verification checklist.

The question

A cost-effectiveness model asks whether a therapy is worth its price. A budget impact model asks a narrower and more immediate question: can this payer afford to adopt it, and when does the money move?

The answer is a subtraction — the world with the new therapy, minus the world without it — and that subtraction is the whole model. Both worlds must be built on the same population, the same horizon and the same cost categories, and only then differenced. Without an explicit market share layer there is only one world and no budget impact at all.

The structure

Target population A hypothetical health plan covering 1,000,000 people Newly diagnosed adults with acute myeloid leukaemia Aged over 75, or unable to receive intensive chemotherapy because of comorbidities Current world — without venetoclax Best supportive care Azacitidine Low-dose cytarabine Decitabine New world — with venetoclax The same four options, plus Venetoclax + azacitidine Venetoclax + low-dose cytarabine Venetoclax + decitabine Market share — how patients divide between the options, year by year Total cost, current world Drug acquisition · Administration · Monitoring Adverse events · Hospitalisation · Transfusions Total cost, new world Drug acquisition · Administration · Monitoring Adverse events · Hospitalisation · Transfusions difference Budget impact — absolute, and per member per month

The layers below follow this diagram, with the budget impact first.

The working model

Each block below is one layer of that diagram, in order. The budget impact comes first because that is the answer; everything under it is the justification.

How to use it

A model handed to someone else needs a note saying how to drive it, so here is one.

The payer-perspective switch is the only control. It moves the model between the social security and private-sector settings the source reports. Everything else on the page is output: the panels recompute together, so a switch flipped at the top changes every figure below it.

Seven panels, in the order the calculation runs. Funnel narrows a million covered lives to the treated population. Market share splits that population between the two worlds being compared. Per-patient cost builds one patient’s annual cost from its components. Budget impact is the difference between the two worlds — the number the model exists to report. OWSA varies the seven parameters the source varies, one at a time.

Two panels are the audit rather than the analysis. Reconciliation sets every rebuilt component against the published figure and states the gap where there is one — it is the record of what agrees with the source and what does not, including the two components that do not. Provenance is the data extraction document: every parameter, the upstream source it comes from, and whether that source is published, derived, calibrated or assumed. Hovering any parameter elsewhere on the page shows the same provenance without leaving the panel you are in.

What not to read into it. The figures are a rebuild of one published analysis under that paper’s assumptions, at 2020 US dollars, over three years. They are not a forecast for any real plan, and the market shares are manufacturer projections rather than observed data.

How it is put together

The model reconstructs a published budget impact analysis from its parameters rather than restating its results, so that every input is visible and every calculation can be checked against the source. Three files, no server, no build step, no dependencies: bim_params.js holds every parameter with its source and no logic, bim_engine.js holds the calculation with no interface code, and bim_ui.js renders. A reviewer auditing the model opens the parameter file and checks it against the paper — the same thing they would ask to do with a spreadsheet, and rarely can.

Three things surfaced during the reconstruction that a spreadsheet would have hidden.

The published supplement contradicts itself, in two places. S1 Table and S7 Table assign the venetoclax + decitabine and venetoclax + LDAC market shares in opposite order, and Table 2 and S2 Table disagree on the dosing days for decitabine and low-dose cytarabine. Neither can be resolved by choosing the more plausible reading; both were settled against a third source.

The mechanism behind hospitalisation is in the questionnaire, not the paper. The main text never says how hospital days are derived. The Delphi questionnaire in the supplementary material asks the panel for days per cycle separately for patients who do and do not achieve complete remission, on the stated assumption that failing to reach remission means progression, and progression means admissions. That is how the headline clinical result reaches the budget.

Two kinds of calibration, and only one of them is legitimate. Scaling whole cost components to match published levels distorts the difference between the two worlds, which is the only quantity a budget impact model reports. Estimating individual parameters that the source uses but never publishes — the panel’s returned hospital days, the effective duration of transfusion independence — against all the published totals including the difference is a different thing, and it is what this model does. The estimates land close to the questionnaire’s own suggested values.

Settling the market share rows. S1 Table gives venetoclax + decitabine 14.5% of the market and venetoclax + LDAC 9.6%; S7 Table gives those same two regimens 12 and 19 patients, the opposite assignment. Reconciled against the drug costs in Table 4, S7’s version reproduces published spend to within 0.35–1.3% and S1’s is out by 3.7–5.2%. The two rows of S1 Table appear to be transposed.

Settling the dosing days. The EMA label for decitabine gives five consecutive days per 28-day cycle, and VIALE-C gave low-dose cytarabine on days 1–10 — matching S2 Table, not Table 2.

Limitations

Each budget year is an independent cohort of 129 newly diagnosed patients, each costed for exactly one year. Nobody carries over. Three consequences, and a payer will find all three.

Incidence rather than prevalence is right here, but the source says the wrong word. This is first-line treatment of newly diagnosed disease, so only incident patients are eligible; counting prevalent patients would double-count people already treated. The source nevertheless writes that it assumes “the prevalence of AML is constant”, while the parameter it uses is an incidence rate. The distinction matters elsewhere: for a chronic therapy the treated pool accumulates year on year, and an incidence-only model would badly understate later years. In unfit AML, with median survival around ten to fifteen months, almost nothing accumulates.

Survival gain does not feed back into the population. Venetoclax combinations extend overall survival materially. Both worlds nonetheless contain the same 129 patients, each alive and incurring a full year of costs. Mortality is absent from both arms, so the model cannot express the argument a payer will certainly make — that a therapy which keeps people alive creates patients who go on costing money. Because the comparator arm has the worse survival, holding everyone alive for a full year overstates the comparator’s cost more than venetoclax’s, which compresses the difference between the two worlds.

Costs are assigned to the year of diagnosis, not the year they fall. The source acknowledges this one: in practice, patients responding well continue therapy beyond the budget period. Venetoclax regimens run longer than their comparators — 10.98 cycles against 8.8 for azacitidine and 3.76 for low-dose cytarabine — so they spill across year boundaries more. Charging each cohort’s whole cost to its year of diagnosis therefore flattens the ramp: the real budget pressure is more back-loaded than the model shows. This is the difference between costing each incident cohort in full and costing only what falls inside the budget year, and it is the second that answers a payer’s cash-flow question.

Scope held to the source

Three years, 28-day cycles, deterministic, no probabilistic sensitivity analysis. No QALYs or ICERs — this is a budget impact model, and mixing the two questions is a common way of answering neither. Costs are US dollars at 2020 prices with no inflation adjustment, following the source’s stated reasoning for excluding it under high inflation.

What reconciles, and what does not

Drug acquisition, adverse events, hospitalisation and transfusions rebuild to within 0.3% of the published figures, and the year-3 budget impact lands within 0.1% under both payer perspectives. Two components do not. Administration is about 45% low, because the source costs later-cycle administration at a daily hospital stay whose unit cost it never reports. Monitoring is about 12% low for reasons not traceable to any published quantity. Together they are 0.5% of total spend. Both are shown in the reconciliation panel rather than closed by fitting — the distinction being that estimating an unpublished parameter inside a documented structure is legitimate, while scaling a component until the total agrees is not.