Skip to content

Add an execution backend

An execution backend submits and observes container tasks. It must not interpret a model's command or score schema.

Implement the required operations

Subclass ExecutionBackend from altar.execution and provide:

  • submit(spec) to launch one ContainerTaskSpec and return a durable TaskHandle;
  • poll(handles) to return one current TaskStatus per handle.

Override capacity(), input staging, output collection, cancellation, or cleanup when the provider supports them. Keep polling non-blocking and preserve task labels unchanged.

Honor the task declaration

The backend must map:

  • image and command without model-specific rewrites;
  • CPU, memory, and GPU requests, with ResourceRequest.count as the exact number of devices;
  • logical input and output transfers;
  • environment values and structured labels;
  • timeout_s as a wall-clock limit enforced by the provider or by poll, and None as no limit. If the provider caps task lifetime or applies a default when no limit is passed, request its maximum for None and reject a longer explicit limit;
  • provider state into Altar task states.

If a requested capability cannot be represented, fail explicitly rather than silently degrading it. In particular, a required GPU task must not be converted into CPU work.

Verify content-addressed inputs

An input Transfer with a digest names exact bytes, such as model weights: the SHA-256 of one regular file. A digest never covers a directory, so a digest-bearing input that resolves to a directory or a missing file is rejected like a mismatch. Every backend verifies digests before the task's command runs, and declares where with the class attribute input_digest_verification:

  • "staging": stage_inputs hashes the bytes it materializes and raises TransferIntegrityError on a mismatch, as the local and Modal backends do.
  • "task": each task checks its own inputs first and fails with reason input_digest_mismatch. Use this when inputs already sit on a shared volume the submitting process cannot read; the Kubernetes backend adds an init container that runs verify_files_script() on the mounted paths.

The conformance suite fails a backend that declares neither.

Hash once per volume

A large input such as a reference genome is read by every task of every job. Hashing it in full before each task can take longer than the task itself, so a verified file carries a verification record: a one-line file beside it, at its path plus .sha256-verified.

verified-file/1 sha256:<64 hex> size=<bytes> mtime=<seconds> ctime=<seconds> inode=<number>

The line names the digest the bytes matched, and the file's size, whole-second modification and status-change times, and inode number when they matched. Rewriting, replacing, or truncating the file changes at least one of those values, so the record no longer matches and the next check hashes the file again. A record is only written once the file's status-change second has passed, and only if the file did not change while it was hashed.

  • verify_file_digest(path, digest) accepts a file from a current record, or hashes it and writes one.
  • verify_transfer_digest(transfer, path, record=True) always hashes, then writes a record. Use it where bytes enter storage your tasks read, and only for copies the backend owns.
  • verify_files_script() is the same check as a POSIX shell program, for BusyBox or GNU coreutils.

The local backend records each verified copy and skips copying an input whose destination already holds a recorded copy of the same digest. The Kubernetes init container accepts recorded inputs and records the ones it hashes. Model runtimes check their inputs the same way, so after a host stages a resource and records it, fold tasks read only its metadata.

Anyone who can write the volume can also write a record. Where untrusted parties can write it, construct the Kubernetes backend with trust_verification_records=False, which hashes every input in every task. Its init container still writes a record after each successful hash, so a runtime that trusts records then accepts the file from the record its own task's verifier just wrote instead of hashing it a second time.

Report failures with stable reasons

Use the reason codes documented on TaskFailure for the same events on every backend: exit_<n> with exit_code for an application exit, timeout, oom_killed, input_digest_mismatch, container_missing for a task the provider can no longer find, and job_failed when the provider gives no finer detail. A missing task must be a terminal failure, never a task that stays running forever.

Preserve durable identity

TaskHandle.external_id must be sufficient to poll the provider after the submitting process restarts. Provider display names are not a substitute for structured labels or durable IDs.

Accept resubmission

A host retries a failed or interrupted run by submitting the same specs again, so submit must never raise because a spec was submitted before. A backend whose provider mints a fresh ID per task, as Docker and Modal do, simply launches another task. A backend that names tasks by their spec must settle the name conflict itself:

  • a task that is still pending or running is adopted: submit returns a handle to it, so the same work never runs twice at once;
  • a task that has finished, whether it succeeded or failed, is deleted and submitted again. A spec names its inputs by path, and a new run under the same job can put different content at the same path (a new variant batch, for example), so a finished task's output is never taken as the new run's.

The Kubernetes backend names each Job by a hash of its spec and follows these rules:

  • Its runtime reports a taken name as K8sJobExistsError (the API server's 409 AlreadyExists) and reads the existing Job with read_job. A custom KubernetesRuntime must raise K8sJobExistsError for a name conflict and report each Job's annotations, and should report its uid.
  • Each Job records a hash of its fully rendered manifest in the altar.kundajelab.org/manifest-sha256 annotation. The name covers only the spec, so an active Job whose manifest differs, for example one created by a backend with another PVC, verifier image, or node selector, is neither adopted nor deleted: submit raises K8sJobExistsError. Let it finish, or delete it, before resubmitting.
  • When upgrading Altar: an active Job created by an older Altar (with no annotation), or whose rendered manifest changed across the upgrade, is refused rather than adopted. Wait for it to finish or delete it. A custom KubernetesRuntime must return annotations from list_jobs and read_job.
  • A finished Job is deleted with its pods before it is created again. The delete is conditioned on the Job's uid, so when two drivers resubmit at once, neither deletes the fresh Job the other just created; the slower one adopts it instead.
  • If the existing Job disappears between the conflict and the read, or is still being deleted, the create is retried with backoff for about a minute before submit raises K8sJobExistsError.

An adopted Job runs with whatever its input paths hold. Content-addressed inputs (a Transfer.digest) are verified before it starts, but inputs named only by path are not. So if a driver crashes and a new run writes different content to the same path before the old Job reads it, the adopted Job reads the new content.

Verify portability

Subclass ExecutionBackendContract, provide a backend fixture wired to a fake provider, and implement its scripting hooks: finish_task, expire_task, launched_timeout, and, for "task" digest verification, fail_input_verification. launched_timeout returns the limit the provider would enforce, including any default it substitutes when none is passed. If the provider cannot run a task without a limit, set the contract's unlimited_timeout_s to its maximum. The suite then checks submission, resubmission of the same spec, polling order, success and failure outcomes, failure reasons, missing tasks, timeouts, labels, and input digest verification. Add your own tests for unsupported capabilities. Then run a model plan through both your backend and a reference backend and compare the declared command and outputs.

Register the package under altar.execution_backends.