Skip to main content
Back to Blog

Mining & Resource Modelling

Reproducibility as a Feature of Geoscience Software

Content-addressed results, recorded provenance, and what it takes for a reviewer to verify a number instead of trusting it.

SpatialTechSolutions Team2026-07-225 min read

Identity by content, not by filename

When a result is identified by the hash of its own content, re-running the same job returns the same identity — and any change to the inputs or parameters, however small, returns a different one. That turns 'is this the same model we signed off?' from a matter of file naming discipline into something a machine can answer.

Provenance belongs on the artifact

A model file in a shared drive knows nothing about itself. A manifest that records the kernel and its digest, the input identities, the parameter hash, the randomness root, the person who ran it, and the toolchain it ran under answers the whole first page of a due-diligence questionnaire without anybody opening a spreadsheet.

Determinism has to be declared and checked

Reproducibility is not a property you hope for. Kernels declare a determinism class and an execution class, jobs declare the policy they expect, and a mismatch is refused before anything runs. Worker count and restarts must not change results — and that is a testable claim, not a reassurance.

Where provenance legitimately stops

A content hash covers bytes, not their origin, so a lineage chain can honestly say which inputs produced a result but not that they came from a particular database. Recording that separately, and labelling it as capture evidence rather than a claim the computation makes, is the difference between provenance and marketing.

Request demo

Turn geospatial insight into an operational workflow

Explore SpatialTechSolutions products, services, and AI-enabled GIS architecture for your organization.