Add an execution backend¶
An execution backend submits and observes container tasks. It must not interpret a model's command or score schema.
Implement the required operations¶
Subclass ExecutionBackend from altar.execution and provide:
submit(spec)to launch oneContainerTaskSpecand return a durableTaskHandle;poll(handles)to return one currentTaskStatusper handle.
Override capacity(), input staging, output collection, cancellation, or cleanup when the provider supports
them. Keep polling non-blocking and preserve task labels unchanged.
Honor the task declaration¶
The backend must map:
- image and command without model-specific rewrites;
- CPU, memory, and GPU requests, with
ResourceRequest.countas the exact number of devices; - logical input and output transfers;
- environment values and structured labels;
timeout_sas a wall-clock limit enforced by the provider or bypoll, andNoneas no limit. If the provider caps task lifetime or applies a default when no limit is passed, request its maximum forNoneand reject a longer explicit limit;- provider state into Altar task states.
If a requested capability cannot be represented, fail explicitly rather than silently degrading it. In particular, a required GPU task must not be converted into CPU work.
Verify content-addressed inputs¶
An input Transfer with a digest names exact bytes, such as model weights: the SHA-256 of one regular file.
A digest never covers a directory, so a digest-bearing input that resolves to a directory or a missing file is
rejected like a mismatch. Every backend verifies digests before the task's command runs, and declares where
with the class attribute input_digest_verification:
"staging":stage_inputshashes the bytes it materializes and raisesTransferIntegrityErroron a mismatch, as the local and Modal backends do."task": each task checks its own inputs first and fails with reasoninput_digest_mismatch. Use this when inputs already sit on a shared volume the submitting process cannot read; the Kubernetes backend adds an init container that runsverify_files_script()on the mounted paths.
The conformance suite fails a backend that declares neither.
Hash once per volume¶
A large input such as a reference genome is read by every task of every job. Hashing it in full before each
task can take longer than the task itself, so a verified file carries a verification record: a one-line file
beside it, at its path plus .sha256-verified.
The line names the digest the bytes matched, and the file's size, whole-second modification and status-change times, and inode number when they matched. Rewriting, replacing, or truncating the file changes at least one of those values, so the record no longer matches and the next check hashes the file again. A record is only written once the file's status-change second has passed, and only if the file did not change while it was hashed.
verify_file_digest(path, digest)accepts a file from a current record, or hashes it and writes one.verify_transfer_digest(transfer, path, record=True)always hashes, then writes a record. Use it where bytes enter storage your tasks read, and only for copies the backend owns.verify_files_script()is the same check as a POSIX shell program, for BusyBox or GNU coreutils.
The local backend records each verified copy and skips copying an input whose destination already holds a recorded copy of the same digest. The Kubernetes init container accepts recorded inputs and records the ones it hashes. Model runtimes check their inputs the same way, so after a host stages a resource and records it, fold tasks read only its metadata.
Anyone who can write the volume can also write a record. Where untrusted parties can write it, construct the
Kubernetes backend with trust_verification_records=False, which hashes every input in every task. Its init
container still writes a record after each successful hash, so a runtime that trusts records then accepts the
file from the record its own task's verifier just wrote instead of hashing it a second time.
Report failures with stable reasons¶
Use the reason codes documented on TaskFailure for the same events on every backend: exit_<n> with
exit_code for an application exit, timeout, oom_killed, input_digest_mismatch, container_missing
for a task the provider can no longer find, and job_failed when the provider gives no finer detail. A
missing task must be a terminal failure, never a task that stays running forever.
Preserve durable identity¶
TaskHandle.external_id must be sufficient to poll the provider after the submitting process restarts.
Provider display names are not a substitute for structured labels or durable IDs.
Accept resubmission¶
A host retries a failed or interrupted run by submitting the same specs again, so submit must never raise
because a spec was submitted before. A backend whose provider mints a fresh ID per task, as Docker and Modal
do, simply launches another task. A backend that names tasks by their spec must settle the name conflict
itself:
- a task that is still pending or running is adopted:
submitreturns a handle to it, so the same work never runs twice at once; - a task that has finished, whether it succeeded or failed, is deleted and submitted again. A spec names its inputs by path, and a new run under the same job can put different content at the same path (a new variant batch, for example), so a finished task's output is never taken as the new run's.
The Kubernetes backend names each Job by a hash of its spec and follows these rules:
- Its runtime reports a taken name as
K8sJobExistsError(the API server's409 AlreadyExists) and reads the existing Job withread_job. A customKubernetesRuntimemust raiseK8sJobExistsErrorfor a name conflict and report each Job's annotations, and should report its uid. - Each Job records a hash of its fully rendered manifest in the
altar.kundajelab.org/manifest-sha256annotation. The name covers only the spec, so an active Job whose manifest differs, for example one created by a backend with another PVC, verifier image, or node selector, is neither adopted nor deleted:submitraisesK8sJobExistsError. Let it finish, or delete it, before resubmitting. - When upgrading Altar: an active Job created by an older Altar (with no annotation), or whose rendered manifest
changed across the upgrade, is refused rather than adopted. Wait for it to finish or delete it. A custom
KubernetesRuntimemust return annotations fromlist_jobsandread_job. - A finished Job is deleted with its pods before it is created again. The delete is conditioned on the Job's uid, so when two drivers resubmit at once, neither deletes the fresh Job the other just created; the slower one adopts it instead.
- If the existing Job disappears between the conflict and the read, or is still being deleted, the create is
retried with backoff for about a minute before
submitraisesK8sJobExistsError.
An adopted Job runs with whatever its input paths hold. Content-addressed inputs (a Transfer.digest) are
verified before it starts, but inputs named only by path are not. So if a driver crashes and a new run writes
different content to the same path before the old Job reads it, the adopted Job reads the new content.
Verify portability¶
Subclass ExecutionBackendContract, provide a backend fixture wired to a fake provider, and implement its
scripting hooks: finish_task, expire_task, launched_timeout, and, for "task" digest verification,
fail_input_verification. launched_timeout returns the limit the provider would enforce, including any
default it substitutes when none is passed. If the provider cannot run a task without a limit, set the
contract's unlimited_timeout_s to its maximum. The suite then checks submission, resubmission of the same
spec, polling order, success and failure outcomes, failure reasons, missing tasks, timeouts, labels, and input
digest verification. Add your own tests for unsupported capabilities. Then run a model plan through both your
backend and a reference backend and compare the declared command and outputs.
Register the package under altar.execution_backends.