Skip to content

VCF input

The preprocessing pipeline accepts coordinate-sorted VCF and BGZF-compressed VCF as alternatives to the canonical headerless variant TSV (chr, pos, ref, alt[, variant_id]). Binary BCF and gVCF are not supported.

The production path uses htslib through pysam and processes a configurable number of ALT occurrences at a time. It does not load the input file into one DataFrame. TSV rows pass through the same occurrence contract.

Usage

python -m altar.variants.preprocessing.pipeline \
    -i input.vcf.gz \
    -g genome.fa \
    -f auto \
    -o valid.tsv \
    --invalid-out invalid.tsv \
    --region-out-dir regions/ \
    --occurrence-out-dir ingest/occurrences/ \
    --canonical-out-dir ingest/canonical/valid/ \
    --invalid-parquet-out-dir ingest/canonical/invalid/ \
    --manifest-out ingest/manifest.json \
    --header-out ingest/header.vcf

The artifact arguments are optional for older callers. When omitted, the pipeline writes them under an ingest/ directory beside valid.tsv.

For a case-selected multi-sample VCF, pass --target-samples-file with one VCF sample identifier per line. An ALT is scoreable when at least one selected sample carries it. The manifest records a digest and count of the selected sample set, not the identifiers themselves. This option is mutually exclusive with --family-phenopacket.

For a sharded callset, use the smaller ingest command on each verified shard; it does not run region filtering:

python -m altar.variants.preprocessing.selected_vcf_ingest \
    --input shard-01.vcf.gz --genome genome.fa \
    --selected-samples-file case-samples.txt --output-dir shard-01-ingest \
    --expected-input-sha256 "$SHARD_SHA256" \
    --expected-selected-samples-sha256 "$SAMPLES_SHA256"

After all shard manifests are verified, deduplicate the canonical variants:

python -m altar.variants.preprocessing.callset_union \
    --ingest-dir shard-01-ingest --ingest-dir shard-02-ingest \
    --output-dir callset-union

The union rejects changed or incompatible shard outputs and publishes its manifest last. Neither command decides which samples are cases; that selection belongs to the caller and is pinned by the sample-set digest.

Outputs

The compatibility outputs remain available:

  • valid.tsv: headerless chr, pos, ref, alt, source_variant_id rows;
  • invalid.tsv: diagnostic rows with stable error codes;
  • bounded region-category TSVs for existing scorers.

The durable outputs are:

  • occurrence Parquet shards with (record_ordinal, alt_index) identity and one row for every ALT, including unsupported alleles;
  • canonical Parquet shards for scoreable alleles;
  • invalid Parquet shards;
  • the parsed source header;
  • a manifest containing input/header digests, counts, shard integrity metadata, reference build, canonicalizer version, and validation version.

Each Parquet shard is atomically renamed into place and has a sibling success marker containing its row count, byte size, and SHA-256 digest. The final manifest is the completion marker for the ingest as a whole.

Semantics

  • A multi-allelic VCF record produces one occurrence per ALT in original order.
  • Duplicate records remain duplicate occurrences. Canonical scoring identity is chr:pos:ref:alt, and downstream consumers may select it distinctly.
  • VCF ID and an optional fifth TSV column become source_variant_id.
  • Spanning-deletion *, symbolic structural variants, breakends, and ALT == REF remain in the occurrence relation with UNSUPPORTED status and a precise error code. They are not sent to scoring.
  • Bare chromosomes are normalized during reference validation. POS remains one-based.
  • Reference validation checks the chromosome, the allele alphabet, the position and the REF bases, and applies no model's sequence window. See Reference validation.
  • INFO, FORMAT, and samples are not copied into the canonical relation. The immutable source VCF retains them for occurrence-aware augmented output.

Rejections

  • gVCF declarations (##GVCFBlock* or ##ALT=<ID=NON_REF,...>) are rejected from the parsed header before output sinks are allocated.
  • Undeclared gVCF input fails on the first <NON_REF>, legacy <*>, or END > POS reference block without a concrete ALT.
  • Ordinary symbolic structural variants with END and spanning-deletion * are not misclassified as gVCF.
  • Unsorted VCF records are rejected because augmented output includes a CSI random-access index.
  • BCF is rejected with an instruction to convert it first.

The older altar.variants.read_variants_frame and load_variants_vcf APIs still return a whole DataFrame for compatibility. Their variant_id column is now the canonical locus-derived score key; VCF IDs remain available only through occurrence-aware preprocessing. Unlike the pipeline, these loaders (and the preprocessing.validate and preprocessing.vcf commands built on them) do not reject gVCF: for VCF input they drop and count symbolic ALTs such as <NON_REF> reference blocks and alleles that have no variant key, and keep concrete alleles. TSV input has no such records, so read_variants_frame and load_variants reject a TSV at the first row that has no key and name that row; preprocessing.validate reports such a row as invalid instead. Every loader rejects a position that is not ASCII digits, naming its row or line. Large-file callers must use preprocessing.streaming.iter_input_occurrence_batches or the preprocessing pipeline.

Augmented VCF export

python -m altar.variants.export.augmented_vcf consumes ordinal-partitioned Parquet buckets, streams the immutable input through htslib, and adds versioned ALTAR1_* INFO fields aligned with ALT order. It writes a BGZF VCF, CSI index, and checksum manifest. The default profile emits bounded headline annotations; full per-model INFO entries require --include-model-scores. The caller chooses which annotation columns become INFO fields; see Annotation INFO fields.

The writer names no model architecture. ALTAR1_MAX_ABS_LOGFC reads each model's logfc result field, a semantic score shared by several models. Per-model ALTAR1_MODEL_SCORE entries contain exactly the result fields listed in the models JSON:

{
  "score_fields": ["logfc", "jsd", "active_allele_quantile", "in_peak"],
  "models": [{"model_id": "...", "name": "...", "column_prefix": "..."}]
}

Use the model manifest's result-schema field names. Each entry is ALT|model|<score_fields...>|prioritized, and the INFO header description records that order. --include-model-scores fails without score_fields.

Input buckets

The --bucket-dir input follows the altar-augmentation-buckets-v1 schema, which the output manifest records as bucket_schema_version. BigQueryScoreStore.export_occurrence_buckets writes it. The directory holds bucket=<n>/part-*.parquet files with one row per source ALT occurrence. The writer reads these columns, plus the annotation columns that the annotation INFO fields name, and ignores all others.

Column Type Fills Source
record_ordinal int alignment ingest occurrences (zero-based record number)
alt_index int alignment ingest occurrences (one-based ALT number)
status, error_code str ALTAR1_STATUS, ALTAR1_ERROR ingest occurrences
result_variant_id str presence of a result joined canonical result, null when none joined
prioritized bool ALTAR1_PRIORITIZED materialized results
<column_prefix>_logfc float ALTAR1_MAX_ABS_LOGFC each model's logfc result field
<column_prefix>_<score_field>, <column_prefix>_prioritized any ALTAR1_MODEL_SCORE with --include-model-scores
most_active_celltype str ALTAR1_TOP_MODEL optional display-rider column holding the headline model's name or column_prefix

The writer reads the other fields only when result_variant_id is set. PENDING status becomes SCORED when a result joined and NO_RESULT otherwise.

export_occurrence_buckets writes every annotation column of every fused annotation source into the buckets, with its native type: a repeated column is a Parquet list and a struct column is a Parquet struct. A column from a host-defined source therefore reaches the writer without any change to Altar.

Annotation INFO fields

The core fields ALTAR1_STATUS, ALTAR1_ERROR, ALTAR1_PRIORITIZED, ALTAR1_MAX_ABS_LOGFC, ALTAR1_TOP_MODEL, and ALTAR1_MODEL_SCORE are fixed. Annotation fields are declared by the caller. Each declaration maps one or more bucket columns to one per-ALT (Number=A) INFO field:

Key Required Meaning
id yes The INFO ID, in VCF syntax ([A-Za-z_][A-Za-z0-9_.]*). It cannot be a core field ID, END, or repeat another declaration.
columns yes Bucket columns tried in order. Each column is extracted on its own, and the first value that is not null or empty after extraction wins.
type yes Integer, Float, or String. A boolean renders as 1 or 0. A value that does not convert, or an Integer outside -2147483640 to 2147483647, fails the writer.
description yes The INFO header description: one line, without double quotes or backslashes.
first_element no Take the first element of a repeated column. An empty list is missing. Default false.
fields no Struct fields tried in order, after first_element. A struct whose listed fields are all null is missing. Default none.
required no Fail when none of columns is in any bucket part, or when a listed struct field is in none of those columns' struct types. Default true.

Because each column is extracted before the next is tried, an empty list or an all-null struct falls back to the next column just as a null does. BigQuery stores a NULL array as an empty one, so this matters for exported buckets.

When first_element or fields is set, a string value is parsed as JSON first, the form a JSON-backed store such as SQLite returns. A string that is not JSON is used as it is. The value that remains must be a scalar; a list or struct fails the writer and names the field. Like the other result fields, annotation fields are filled only for an ALT whose result joined.

The struct-field check reads the Parquet column types. It cannot check a column stored as a JSON string, and it is skipped for a field with required: false, such as the default ALTAR1_NEAREST_GENE, which tries several gene keys.

String values are percent-encoded so that a value cannot break the VCF line or split the per-ALT list. %, ,, ;, =, tab, CR, and LF become %25, %2C, %3B, %3D, %09, %0D, and %0A, the VCF 4.3 encoding. A space also becomes %20. That is an Altar addition, not part of the VCF 4.3 list: many inputs declare VCFv4.2, which forbids whitespace in INFO values. : is left as it is, because it has no meaning inside INFO.

A declared ID that the input VCF already defines, such as AF, fails the writer. Writing it would replace the input's own values and drop them for an ALT without a result. IDs in Altar's ALTAR1_ namespace are the exception: an input that an earlier augmentation wrote already carries them, so they follow the core-field rule. A header line with the same Number and Type is kept and its values are replaced, and a different Number or Type fails. An ID may use the ALTAR1_ prefix, for example to keep an INFO ID such as ALTAR1_GNOMAD_AF that an earlier writer version emitted and a downstream consumer still reads.

Without a declaration, the writer emits these defaults, which do not set required, so a job without a region or AlphaMissense source leaves them missing:

INFO ID Type Columns Extraction Source
ALTAR1_REGION String region_type altar.variants.annotation.annotate
ALTAR1_NEAREST_GENE String nearest_genes first element, then gene_name, symbol, name, or gene_id altar.variants.annotation.annotate
ALTAR1_CCRE String ccre_id, then the legacy ccre altar.variants.annotation.annotate
ALTAR1_AM_SCORE Float am_pathogenicity altar-alphamissense
ALTAR1_AM_CLASS String am_class altar-alphamissense

A declaration file must set include_defaults. true keeps the defaults and adds its fields after them; false writes only its fields. Requiring the key means adding one field cannot silently drop the defaults. This file maps a host's own gnomAD source column af_global and the phred field of a struct column cadd to INFO IDs the host chooses:

{
  "include_defaults": true,
  "fields": [
    {
      "id": "HOST_GNOMAD_AF",
      "columns": ["af_global"],
      "type": "Float",
      "description": "gnomAD v4.1.1 genomes global allele frequency per ALT"
    },
    {
      "id": "HOST_CADD_PHRED",
      "columns": ["cadd"],
      "type": "Float",
      "description": "CADD v1.7 PHRED score per ALT",
      "fields": ["phred"]
    }
  ]
}

Pass it with --annotation-fields-json annotation-fields.json. The command's write_augmented_vcf function takes the same declarations as its annotation_fields argument, a sequence of AnnotationInfoField values from the command module: None writes the defaults, and (*DEFAULT_ANNOTATION_FIELDS, field) keeps them and adds field. load_annotation_fields(path) reads the JSON file into that sequence.

The output manifest records the declarations that produced the VCF under annotation_fields, in the JSON form above with every key filled in. The manifest schema_version stays altar-augmented-vcf-v1: the key is an addition, and a run without a declaration writes the same VCF as before.

Custom variant IDs

The source identifier is preserved through compatibility outputs and the occurrence relation. It is not the canonical scoring key. See variant-id.md.