This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.
Specification conformance, grid resolution and inter-code spread in an independent simulation of the SPE11B CO2 storage benchmark
Downloads
Authors
Abstract
Confidence in geological CO2 storage forecasts rests on simulators whose discrepancies are, for the most part, characterised against each other rather than against a truth. We wrote a two-phase, two-component, fully implicit CO2 storage simulator from scratch, ran it on the SPE11B benchmark to the specification's own reporting resolution (840 × 120 cells, 1,000 yr), and used the reimplementation as an instrument for asking where the discrepancy against the 32 archived submissions comes from. A clause-by-clause reading of the specification against the code found 7 conformance defects, none of which the verification suite—465 assertions including analytical solutions checked to machine precision—could have detected, because every test encoded the same reading as the code. Repairing 6 of them at fixed grids moved the run 9 and 10 of 54 sampled quantity–time pairs further inside the peer envelope, a measure that cannot be improved by becoming more typical; grid refinement at matched conformance moves that count by less. Because both conformance stages were run at all four grids, the stage change and the resolution change are separated rather than confounded: the largest per-quantity stage effect decays monotonically from 0.88 to 0.03 MAD as the grid refines, and its aggregate, which changes sign twice across the ladder, ends at -0.01 MAD at the reporting grid. Two measured floors calibrate that—a change of time-step policy alone moves the aggregate by 0.46 MAD on identical physics, and a tenfold-finer injection step by 0.06 MAD, more than the finest refinement buys. Separately, a silent clamp on the dissolved-CO2 mass fraction inside the Newton loop created a false convergence wall: it capped the time step at 0.095 yr and consumed 52% of wall-clock time in discarded iterations while every gate stayed green. Of 18 performance interventions measured by isolated paired A/B, 8 survived and 4 first reported the wrong sign. On the Sleipner benchmark model, at a single fixed resolution, five reduced models scored against the observed top-sand plume outline exceed a matched-area baseline by at most 0.035 in overlap—so which mechanism a model represents separates them less than an explicit baseline reveals. On the benchmark, where both axes were varied, the larger discrepancy shifts came from what the model was asked to represent rather than from how finely it was resolved. Benchmark reporting should require conformance to be demonstrated rather than assumed.
DOI
https://doi.org/10.31223/X5DB8M
Subjects
Geology
Keywords
CO2 storage, SPE11, benchmarking, inter-code comparison, code verification, specification conformance, grid convergence, Sleipner
Dates
Published: 2026-08-12 22:55
Last Updated: 2026-08-12 22:55
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None.
Metrics
Views: 19
Downloads: 0
There are no comments or no comments have been made public for this article.