Reproducibility and provenance¶
A result is reproducible only when its scientific inputs and execution identity can be reconstructed.
Record for every analysis¶
- Altar and binding package versions;
- genome build and reference FASTA identity;
- canonicalization and validation version;
- model architecture and model-instance identifier;
- checkpoint, weights, fold set, or hosted API configuration;
- immutable container digest or external API/model version;
- declared score schema;
- prioritization rule and thresholds;
- annotation and relation-source releases;
- local asset manifests and checksums;
- execution date and relevant parameters.
Prefer immutable references¶
Use digest-qualified container images, pinned dataset revisions, checksummed local assets, and explicit model configuration. A floating image tag or service default can change while the Python integration remains the same version.
Hosted models require special care: record the provider, model or API version when available, selected output types and ontology terms, and request date. If a provider does not expose immutable model versions, document that limitation rather than implying byte-for-byte reproducibility.
Keep transformations visible¶
Bindings that reduce rich output to headline scalar columns must document the reduction. One-to-many source records should remain recoverable in detail output or a repeated field when scientifically material.