Skip to content

Identity package

altar-identity (import altar_identity) is the dependency-free package that defines variant keys, the variant file Altar stages for a model container, and SHA-256 file identity. Bindings import its variant names through altar.models, altar.sources, or altar.variants; model runtimes that cannot install Altar core import this package directly. See the identity package for its compatibility rules.

altar_identity

Variant and content identity shared by Altar core and its model runtimes.

This package has no dependencies and supports Python 3.9, so a runtime pinned to an older scientific stack builds variant IDs, reads Altar's variant files, and verifies staged files with the same code as Altar core.

  • VariantKey, canonical_chromosome, canonical_variant_id, parse_position: the canonical chr:pos:REF:ALT key and its textual position rule.
  • read_variants, read_variant_rows, batched: the headerless variant file Altar hands to a model container.
  • sha256_file, verify_file, parse_sha256_digest: the sha256:<hex> identity of one regular file, and the verification records that let many tasks share one hash of a staged file. SHA256_DIGEST_PATTERN is the one spelling of the digest rule, for schemas and patterns that embed it.

SHA256_DIGEST_PATTERN module-attribute

SHA256_DIGEST_PATTERN = f'{SHA256_PREFIX}[0-9a-f]{{64}}'

Unanchored regular expression for a canonical digest. Anchor it (re.fullmatch, or ^...$ in a schema) or embed it in a larger pattern, such as an image reference, instead of spelling the digest rule again.

DigestMismatchError

Bases: ValueError

A path does not hold the single regular file whose bytes a digest names.

VariantFileError

VariantFileError(
    message: str,
    *,
    line: int | None = None,
    reason: str | None = None,
)

Bases: ValueError

A variant file row is malformed, has no canonical identity, or repeats a variant.

The message is path:line: reason. line and reason are also attributes, so a caller that reports errors in its own words (for example, without a temporary local path) need not parse the message.

VariantRow

Bases: NamedTuple

One canonical variant and the one-based file line it was read from.

VariantIdentityError

Bases: ValueError

A variant key is malformed, non-canonical, or disagrees with its locus.

VariantKey dataclass

VariantKey(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
)

Canonical one-based biallelic key serialized as chr:pos:ref:alt.

Fields follow the rules in this module's documentation. The class normalizes spelling but does not normalize biological representation. Callers must left-align and trim indels against the selected reference before constructing a key when equivalence across representations matters.

variant_id property

variant_id: str

Return the portable score, detail, and annotation join key.

parse classmethod

parse(variant_id: str) -> VariantKey

Parse and normalize a four-field textual key.

require_canonical classmethod

require_canonical(variant_id: str) -> VariantKey

Parse variant_id and reject aliases instead of silently rewriting them.

from_fields classmethod

from_fields(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
) -> VariantKey

Construct a canonical key from decomposed one-based locus fields.

has_verification_record

has_verification_record(
    path: str | PathLike[str], digest: str
) -> bool

Return whether path is a regular file with a current verification record for digest.

The record must name digest and match the file's current size, modification time, status-change time, and inode number. A missing, unreadable, or stale record returns False. Checking a record reads two small pieces of metadata and never reads the file's bytes. Symlinks are followed.

is_sha256_digest

is_sha256_digest(value: object) -> bool

Return whether value is exactly sha256: followed by 64 lowercase hexadecimal digits.

parse_sha256_digest

parse_sha256_digest(value: str) -> str

Return value unchanged if it is a canonical SHA-256 digest; otherwise raise ValueError.

The check is exact: no surrounding whitespace, no trailing newline, and no uppercase hexadecimal.

sha256_file

sha256_file(path: str | PathLike[str]) -> str

Return the sha256:<hex> digest of the bytes at path, reading it in bounded chunks.

verification_record_line

verification_record_line(
    digest: str, status: stat_result
) -> str

Return the verification-record line for digest and a file's os.stat result.

int() of the float times matches the whole seconds stat -c %Y and %Z print for any file written after 1970, so a shell verifier renders the same line.

verify_file

verify_file(
    path: str | PathLike[str],
    expected: str,
    *,
    label: str = "resource",
    trust_record: bool = False,
    record: bool = False,
) -> bool

Raise DigestMismatchError unless path is one regular file whose bytes have digest expected.

expected must be a canonical SHA-256 digest; a malformed one raises ValueError. label names the file in the error message.

By default the file is always hashed. With trust_record, a current verification record accepts the file without reading it, so a file staged once and read by many tasks is hashed once. With record, a successful hash writes a record for later checks. Anyone who can write the storage can also write a record, so trust records only on storage whose writers you trust, and never where bytes first arrive (a download).

Returns True when the file was hashed and False when a record accepted it.

write_verification_record

write_verification_record(
    path: str | PathLike[str],
    digest: str,
    before: stat_result,
) -> bool

Write a record for bytes that just matched digest. Returns whether a record was written.

before is the file's status from before it was hashed. No record is written if the file changed while it was being hashed. No record is written either while the file's status-change second is still the current second: a later change within that same second would leave every recorded value unchanged. The next check after that second writes the record instead. The write is atomic and best effort. A directory the caller cannot write only means the next check hashes the file again.

batched

batched(
    values: Iterable[T], size: int
) -> Iterator[tuple[T, ...]]

Yield consecutive tuples of at most size items without reading ahead of the current batch.

read_variant_rows

read_variant_rows(
    path: str | PathLike[str],
    *,
    duplicates: DuplicatePolicy = "error",
) -> Iterator[VariantRow]

Stream VariantRow(line, key) records, as read_variants does, for callers that report lines.

A runtime that applies its own policy to each variant (for example, SNVs only) uses the line to name the offending row in its error, in the same path:line form as this reader's errors.

read_variants

read_variants(
    path: str | PathLike[str],
    *,
    duplicates: DuplicatePolicy = "error",
) -> Iterator[VariantKey]

Stream canonical VariantKey values from a headerless chr pos ref alt [label] file.

duplicates sets what happens when two rows have the same canonical variant_id: "error" (the default) raises VariantFileError, "skip" yields only the first, and "allow" yields every row. Core's scoring preparation never writes a variant twice. Errors name the file and line.

canonical_chromosome

canonical_chromosome(chromosome: str) -> str

Return the canonical chr-prefixed spelling used in portable keys.

Primary chromosomes accept common aliases (1/chr1/CHR01 and M/MT/chrM) in any case. Other reference contigs keep their exact spelling after an optional case-insensitive chr prefix is normalized, so CHRUn_KI270302v1 becomes chrUn_KI270302v1 but chrun_KI270302v1 stays as written. Non-ASCII text, whitespace inside the name, and characters outside the VCF contig-name set (including :) raise VariantIdentityError. This is a naming rule only; it does not claim that contigs from different assemblies are interchangeable.

canonical_variant_id

canonical_variant_id(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
) -> str

Return the canonical text key for decomposed one-based locus fields.

parse_position

parse_position(text: str) -> int

Parse a position written as text, such as a VCF POS or a TSV cell.

Surrounding ASCII whitespace is trimmed and the rest must be ASCII digits, so +5, 1_000, 1e3 and non-ASCII digits raise VariantIdentityError although int() accepts some of them. The value is not range-checked; VariantKey requires at least 1.