Identity package¶
altar-identity (import altar_identity) is the dependency-free package that defines variant keys, the
variant file Altar stages for a model container, and SHA-256 file identity. Bindings import its variant names
through altar.models, altar.sources, or altar.variants; model runtimes that cannot install Altar core
import this package directly. See the identity package for its
compatibility rules.
altar_identity
¶
Variant and content identity shared by Altar core and its model runtimes.
This package has no dependencies and supports Python 3.9, so a runtime pinned to an older scientific stack builds variant IDs, reads Altar's variant files, and verifies staged files with the same code as Altar core.
VariantKey,canonical_chromosome,canonical_variant_id,parse_position: the canonicalchr:pos:REF:ALTkey and its textual position rule.read_variants,read_variant_rows,batched: the headerless variant file Altar hands to a model container.sha256_file,verify_file,parse_sha256_digest: thesha256:<hex>identity of one regular file, and the verification records that let many tasks share one hash of a staged file.SHA256_DIGEST_PATTERNis the one spelling of the digest rule, for schemas and patterns that embed it.
SHA256_DIGEST_PATTERN
module-attribute
¶
Unanchored regular expression for a canonical digest. Anchor it (re.fullmatch, or ^...$ in a schema)
or embed it in a larger pattern, such as an image reference, instead of spelling the digest rule again.
DigestMismatchError
¶
Bases: ValueError
A path does not hold the single regular file whose bytes a digest names.
VariantFileError
¶
Bases: ValueError
A variant file row is malformed, has no canonical identity, or repeats a variant.
The message is path:line: reason. line and reason are also attributes, so a caller that reports
errors in its own words (for example, without a temporary local path) need not parse the message.
VariantRow
¶
Bases: NamedTuple
One canonical variant and the one-based file line it was read from.
VariantIdentityError
¶
Bases: ValueError
A variant key is malformed, non-canonical, or disagrees with its locus.
VariantKey
dataclass
¶
Canonical one-based biallelic key serialized as chr:pos:ref:alt.
Fields follow the rules in this module's documentation. The class normalizes spelling but does not normalize biological representation. Callers must left-align and trim indels against the selected reference before constructing a key when equivalence across representations matters.
parse
classmethod
¶
parse(variant_id: str) -> VariantKey
Parse and normalize a four-field textual key.
require_canonical
classmethod
¶
require_canonical(variant_id: str) -> VariantKey
Parse variant_id and reject aliases instead of silently rewriting them.
from_fields
classmethod
¶
from_fields(
chromosome: str,
position: int,
reference_allele: str,
alternate_allele: str,
) -> VariantKey
Construct a canonical key from decomposed one-based locus fields.
has_verification_record
¶
Return whether path is a regular file with a current verification record for digest.
The record must name digest and match the file's current size, modification time, status-change time,
and inode number. A missing, unreadable, or stale record returns False. Checking a record reads two
small pieces of metadata and never reads the file's bytes. Symlinks are followed.
is_sha256_digest
¶
Return whether value is exactly sha256: followed by 64 lowercase hexadecimal digits.
parse_sha256_digest
¶
Return value unchanged if it is a canonical SHA-256 digest; otherwise raise ValueError.
The check is exact: no surrounding whitespace, no trailing newline, and no uppercase hexadecimal.
sha256_file
¶
Return the sha256:<hex> digest of the bytes at path, reading it in bounded chunks.
verification_record_line
¶
Return the verification-record line for digest and a file's os.stat result.
int() of the float times matches the whole seconds stat -c %Y and %Z print for any file written
after 1970, so a shell verifier renders the same line.
verify_file
¶
verify_file(
path: str | PathLike[str],
expected: str,
*,
label: str = "resource",
trust_record: bool = False,
record: bool = False,
) -> bool
Raise DigestMismatchError unless path is one regular file whose bytes have digest expected.
expected must be a canonical SHA-256 digest; a malformed one raises ValueError. label names
the file in the error message.
By default the file is always hashed. With trust_record, a current verification record accepts the file
without reading it, so a file staged once and read by many tasks is hashed once. With record, a
successful hash writes a record for later checks. Anyone who can write the storage can also write a record,
so trust records only on storage whose writers you trust, and never where bytes first arrive (a download).
Returns True when the file was hashed and False when a record accepted it.
write_verification_record
¶
Write a record for bytes that just matched digest. Returns whether a record was written.
before is the file's status from before it was hashed. No record is written if the file changed while it
was being hashed. No record is written either while the file's status-change second is still the current
second: a later change within that same second would leave every recorded value unchanged. The next check
after that second writes the record instead. The write is atomic and best effort. A directory the caller
cannot write only means the next check hashes the file again.
batched
¶
Yield consecutive tuples of at most size items without reading ahead of the current batch.
read_variant_rows
¶
read_variant_rows(
path: str | PathLike[str],
*,
duplicates: DuplicatePolicy = "error",
) -> Iterator[VariantRow]
Stream VariantRow(line, key) records, as read_variants does, for callers that report lines.
A runtime that applies its own policy to each variant (for example, SNVs only) uses the line to name the
offending row in its error, in the same path:line form as this reader's errors.
read_variants
¶
read_variants(
path: str | PathLike[str],
*,
duplicates: DuplicatePolicy = "error",
) -> Iterator[VariantKey]
Stream canonical VariantKey values from a headerless chr pos ref alt [label] file.
duplicates sets what happens when two rows have the same canonical variant_id: "error"
(the default) raises VariantFileError, "skip" yields only the first, and "allow" yields every
row. Core's scoring preparation never writes a variant twice. Errors name the file and line.
canonical_chromosome
¶
Return the canonical chr-prefixed spelling used in portable keys.
Primary chromosomes accept common aliases (1/chr1/CHR01 and
M/MT/chrM) in any case. Other reference contigs keep their exact
spelling after an optional case-insensitive chr prefix is normalized, so
CHRUn_KI270302v1 becomes chrUn_KI270302v1 but chrun_KI270302v1
stays as written. Non-ASCII text, whitespace inside the name, and characters
outside the VCF contig-name set (including :) raise
VariantIdentityError. This is a naming rule only; it does not claim that
contigs from different assemblies are interchangeable.
canonical_variant_id
¶
canonical_variant_id(
chromosome: str,
position: int,
reference_allele: str,
alternate_allele: str,
) -> str
Return the canonical text key for decomposed one-based locus fields.
parse_position
¶
Parse a position written as text, such as a VCF POS or a TSV cell.
Surrounding ASCII whitespace is trimmed and the rest must be ASCII digits, so +5, 1_000, 1e3 and
non-ASCII digits raise VariantIdentityError although int() accepts some of them. The value is not
range-checked; VariantKey requires at least 1.