VCF input¶
The preprocessing pipeline accepts coordinate-sorted VCF and BGZF-compressed
VCF as alternatives to the canonical headerless variant TSV
(chr, pos, ref, alt[, variant_id]). Binary BCF and gVCF are not supported.
The production path uses htslib through pysam and processes a configurable
number of ALT occurrences at a time. It does not load the input file into one
DataFrame. TSV rows pass through the same occurrence contract.
Usage¶
python -m altar.variants.preprocessing.pipeline \
-i input.vcf.gz \
-g genome.fa \
-f auto \
-o valid.tsv \
--invalid-out invalid.tsv \
--region-out-dir regions/ \
--occurrence-out-dir ingest/occurrences/ \
--canonical-out-dir ingest/canonical/valid/ \
--invalid-parquet-out-dir ingest/canonical/invalid/ \
--manifest-out ingest/manifest.json \
--header-out ingest/header.vcf
The artifact arguments are optional for older callers. When omitted, the
pipeline writes them under an ingest/ directory beside valid.tsv.
For a case-selected multi-sample VCF, pass --target-samples-file with one
VCF sample identifier per line. An ALT is scoreable when at least one selected
sample carries it. The manifest records a digest and count of the selected
sample set, not the identifiers themselves. This option is mutually exclusive
with --family-phenopacket.
For a sharded callset, use the smaller ingest command on each verified shard; it does not run region filtering:
python -m altar.variants.preprocessing.selected_vcf_ingest \
--input shard-01.vcf.gz --genome genome.fa \
--selected-samples-file case-samples.txt --output-dir shard-01-ingest \
--expected-input-sha256 "$SHARD_SHA256" \
--expected-selected-samples-sha256 "$SAMPLES_SHA256"
After all shard manifests are verified, deduplicate the canonical variants:
python -m altar.variants.preprocessing.callset_union \
--ingest-dir shard-01-ingest --ingest-dir shard-02-ingest \
--output-dir callset-union
The union rejects changed or incompatible shard outputs and publishes its manifest last. Neither command decides which samples are cases; that selection belongs to the caller and is pinned by the sample-set digest.
Outputs¶
The compatibility outputs remain available:
valid.tsv: headerlesschr, pos, ref, alt, source_variant_idrows;invalid.tsv: diagnostic rows with stable error codes;- bounded region-category TSVs for existing scorers.
The durable outputs are:
- occurrence Parquet shards with
(record_ordinal, alt_index)identity and one row for every ALT, including unsupported alleles; - canonical Parquet shards for scoreable alleles;
- invalid Parquet shards;
- the parsed source header;
- a manifest containing input/header digests, counts, shard integrity metadata, reference build, canonicalizer version, and validation version.
Each Parquet shard is atomically renamed into place and has a sibling success marker containing its row count, byte size, and SHA-256 digest. The final manifest is the completion marker for the ingest as a whole.
Semantics¶
- A multi-allelic VCF record produces one occurrence per ALT in original order.
- Duplicate records remain duplicate occurrences. Canonical scoring identity is
chr:pos:ref:alt, and downstream consumers may select it distinctly. - VCF
IDand an optional fifth TSV column becomesource_variant_id. - Spanning-deletion
*, symbolic structural variants, breakends, andALT == REFremain in the occurrence relation withUNSUPPORTEDstatus and a precise error code. They are not sent to scoring. - Bare chromosomes are normalized during reference validation. POS remains one-based.
- Reference validation checks the chromosome, the allele alphabet, the position and the REF bases, and applies no model's sequence window. See Reference validation.
- INFO, FORMAT, and samples are not copied into the canonical relation. The immutable source VCF retains them for occurrence-aware augmented output.
Rejections¶
- gVCF declarations (
##GVCFBlock*or##ALT=<ID=NON_REF,...>) are rejected from the parsed header before output sinks are allocated. - Undeclared gVCF input fails on the first
<NON_REF>, legacy<*>, orEND > POSreference block without a concrete ALT. - Ordinary symbolic structural variants with
ENDand spanning-deletion*are not misclassified as gVCF. - Unsorted VCF records are rejected because augmented output includes a CSI random-access index.
- BCF is rejected with an instruction to convert it first.
The older altar.variants.read_variants_frame and load_variants_vcf APIs
still return a whole DataFrame for compatibility. Their variant_id column is
now the canonical locus-derived score key; VCF IDs remain available only through
occurrence-aware preprocessing. Unlike the pipeline, these loaders (and the
preprocessing.validate and preprocessing.vcf commands built on them) do not
reject gVCF: for VCF input they drop and count symbolic ALTs such as <NON_REF>
reference blocks and alleles that have no variant key, and keep
concrete alleles. TSV input has no such records, so read_variants_frame and
load_variants reject a TSV at the first row that has no key and name that row;
preprocessing.validate reports such a row as invalid instead. Every loader
rejects a position that is not ASCII digits, naming its row or line. Large-file
callers must use preprocessing.streaming.iter_input_occurrence_batches or the
preprocessing pipeline.
Augmented VCF export¶
python -m altar.variants.export.augmented_vcf consumes ordinal-partitioned
Parquet buckets, streams the immutable input through htslib, and adds
versioned ALTAR1_* INFO fields aligned with ALT order. It writes a BGZF VCF,
CSI index, and checksum manifest. The default profile emits bounded headline
annotations; full per-model INFO entries require --include-model-scores.
The caller chooses which annotation columns become INFO fields; see
Annotation INFO fields.
The writer names no model architecture. ALTAR1_MAX_ABS_LOGFC reads each
model's logfc result field, a semantic score shared by several models.
Per-model ALTAR1_MODEL_SCORE entries contain exactly the result fields listed
in the models JSON:
{
"score_fields": ["logfc", "jsd", "active_allele_quantile", "in_peak"],
"models": [{"model_id": "...", "name": "...", "column_prefix": "..."}]
}
Use the model manifest's result-schema field names. Each entry is
ALT|model|<score_fields...>|prioritized, and the INFO header description
records that order. --include-model-scores fails without score_fields.
Input buckets¶
The --bucket-dir input follows the altar-augmentation-buckets-v1 schema,
which the output manifest records as bucket_schema_version.
BigQueryScoreStore.export_occurrence_buckets writes it. The directory holds
bucket=<n>/part-*.parquet files with one row per source ALT occurrence. The
writer reads these columns, plus the annotation columns that the
annotation INFO fields name, and ignores all
others.
| Column | Type | Fills | Source |
|---|---|---|---|
record_ordinal |
int | alignment | ingest occurrences (zero-based record number) |
alt_index |
int | alignment | ingest occurrences (one-based ALT number) |
status, error_code |
str | ALTAR1_STATUS, ALTAR1_ERROR |
ingest occurrences |
result_variant_id |
str | presence of a result | joined canonical result, null when none joined |
prioritized |
bool | ALTAR1_PRIORITIZED |
materialized results |
<column_prefix>_logfc |
float | ALTAR1_MAX_ABS_LOGFC |
each model's logfc result field |
<column_prefix>_<score_field>, <column_prefix>_prioritized |
any | ALTAR1_MODEL_SCORE |
with --include-model-scores |
most_active_celltype |
str | ALTAR1_TOP_MODEL |
optional display-rider column holding the headline model's name or column_prefix |
The writer reads the other fields only when result_variant_id is set.
PENDING status becomes SCORED when a result joined and NO_RESULT
otherwise.
export_occurrence_buckets writes every annotation column of every fused
annotation source into the buckets, with its native type: a repeated column is
a Parquet list and a struct column is a Parquet struct. A column from a
host-defined source therefore reaches the writer without any change to Altar.
Annotation INFO fields¶
The core fields ALTAR1_STATUS, ALTAR1_ERROR, ALTAR1_PRIORITIZED,
ALTAR1_MAX_ABS_LOGFC, ALTAR1_TOP_MODEL, and ALTAR1_MODEL_SCORE are fixed.
Annotation fields are declared by the caller. Each declaration maps one or more
bucket columns to one per-ALT (Number=A) INFO field:
| Key | Required | Meaning |
|---|---|---|
id |
yes | The INFO ID, in VCF syntax ([A-Za-z_][A-Za-z0-9_.]*). It cannot be a core field ID, END, or repeat another declaration. |
columns |
yes | Bucket columns tried in order. Each column is extracted on its own, and the first value that is not null or empty after extraction wins. |
type |
yes | Integer, Float, or String. A boolean renders as 1 or 0. A value that does not convert, or an Integer outside -2147483640 to 2147483647, fails the writer. |
description |
yes | The INFO header description: one line, without double quotes or backslashes. |
first_element |
no | Take the first element of a repeated column. An empty list is missing. Default false. |
fields |
no | Struct fields tried in order, after first_element. A struct whose listed fields are all null is missing. Default none. |
required |
no | Fail when none of columns is in any bucket part, or when a listed struct field is in none of those columns' struct types. Default true. |
Because each column is extracted before the next is tried, an empty list or an all-null struct falls back to the next column just as a null does. BigQuery stores a NULL array as an empty one, so this matters for exported buckets.
When first_element or fields is set, a string value is parsed as JSON
first, the form a JSON-backed store such as SQLite returns. A string that is not
JSON is used as it is. The value that remains must be a scalar; a list or
struct fails the writer and names the field. Like the other result fields,
annotation fields are filled only for an ALT whose result joined.
The struct-field check reads the Parquet column types. It cannot check a
column stored as a JSON string, and it is skipped for a field with
required: false, such as the default ALTAR1_NEAREST_GENE, which tries
several gene keys.
String values are percent-encoded so that a value cannot break the VCF line or
split the per-ALT list. %, ,, ;, =, tab, CR, and LF become %25,
%2C, %3B, %3D, %09, %0D, and %0A, the VCF 4.3 encoding. A space
also becomes %20. That is an Altar addition, not part of the VCF 4.3 list:
many inputs declare VCFv4.2, which forbids whitespace in INFO values. : is
left as it is, because it has no meaning inside INFO.
A declared ID that the input VCF already defines, such as AF, fails the
writer. Writing it would replace the input's own values and drop them for an
ALT without a result. IDs in Altar's ALTAR1_ namespace are the exception:
an input that an earlier augmentation wrote already carries them, so they
follow the core-field rule. A header line with the same Number and Type is
kept and its values are replaced, and a different Number or Type fails. An
ID may use the ALTAR1_ prefix, for example to keep an INFO ID such as
ALTAR1_GNOMAD_AF that an earlier writer version emitted and a downstream
consumer still reads.
Without a declaration, the writer emits these defaults, which do not set
required, so a job without a region or AlphaMissense source leaves them
missing:
| INFO ID | Type | Columns | Extraction | Source |
|---|---|---|---|---|
ALTAR1_REGION |
String | region_type |
altar.variants.annotation.annotate |
|
ALTAR1_NEAREST_GENE |
String | nearest_genes |
first element, then gene_name, symbol, name, or gene_id |
altar.variants.annotation.annotate |
ALTAR1_CCRE |
String | ccre_id, then the legacy ccre |
altar.variants.annotation.annotate |
|
ALTAR1_AM_SCORE |
Float | am_pathogenicity |
altar-alphamissense |
|
ALTAR1_AM_CLASS |
String | am_class |
altar-alphamissense |
A declaration file must set include_defaults. true keeps the defaults and
adds its fields after them; false writes only its fields. Requiring the
key means adding one field cannot silently drop the defaults. This file maps a
host's own gnomAD source column
af_global and the phred field of a struct column cadd to INFO IDs the
host chooses:
{
"include_defaults": true,
"fields": [
{
"id": "HOST_GNOMAD_AF",
"columns": ["af_global"],
"type": "Float",
"description": "gnomAD v4.1.1 genomes global allele frequency per ALT"
},
{
"id": "HOST_CADD_PHRED",
"columns": ["cadd"],
"type": "Float",
"description": "CADD v1.7 PHRED score per ALT",
"fields": ["phred"]
}
]
}
Pass it with --annotation-fields-json annotation-fields.json. The command's
write_augmented_vcf function takes the same declarations as its
annotation_fields argument, a sequence of AnnotationInfoField values from
the command module: None writes the defaults, and
(*DEFAULT_ANNOTATION_FIELDS, field) keeps them and adds field.
load_annotation_fields(path) reads the JSON file into that sequence.
The output manifest records the declarations that produced the VCF under
annotation_fields, in the JSON form above with every key filled in. The
manifest schema_version stays altar-augmented-vcf-v1: the key is an
addition, and a run without a declaration writes the same VCF as before.
Custom variant IDs¶
The source identifier is preserved through compatibility outputs and the occurrence relation. It is not the canonical scoring key. See variant-id.md.