Skip to content

Run container tasks

Execution backends run ContainerTaskSpec values produced by model bindings. The same task declaration can run locally, on Modal, or on Kubernetes.

Backend lifecycle

stage declared inputs
        ↓
submit independent tasks
        ↓
poll task handles
        ↓
submit a dependent task when prerequisites succeed
        ↓
collect declared outputs

The application orchestrates this lifecycle. Altar is a library, not a scheduler service.

Logical paths and transfers

A task command uses paths beneath a common container mount root. PathResolver converts a logical analysis path into that command path. Transfer values tell a backend which objects must be staged in or collected out.

This lets a shared-volume Kubernetes backend and a per-task staging backend consume the same model plan even though their transfer implementations differ.

Task state

Submission returns a TaskHandle. Polling produces TaskStatus values such as pending, running, succeeded, or failed. Persist handles in a TaskLedger when orchestration must survive process restarts.

Some model workflows have dependent steps—for example, score several folds and summarize them. GroupSpec and reconcile_once() describe when the dependent task is ready and ensure it is submitted once. This is low-level orchestration machinery; users scoring through a host application should not need to manage it directly.

Optional model preparation and interpretation use ModelWorkflowPlan, a durable DAG of named WorkflowStage values. run_model_workflow() is the small reference consumer: it stages external inputs, submits every ready stage as a parallel wave, polls for completion, and collects declared terminal outputs. Use the dict or JSON serializers when a host needs to persist the plan before execution. Long-running hosts may instead consume the same plan with their own durable scheduler; the plan contains no provider client or credential.

Scoring deliberately keeps its specialized InlineScoringPlan and ContainerScoringPlan forms. The generic operation DAG does not change cache-miss packing, scoring batches, result codecs, or score-store writes.

Failure handling

Backends report terminal task failure with a stable TaskFailure.reason, such as exit_1, timeout, oom_killed, input_digest_mismatch, or container_missing, so a host can branch on the same code whichever backend ran the task. Every reference backend enforces a task's timeout_s and verifies content-addressed inputs before the task's command runs. A task without timeout_s has no limit on local Docker and Kubernetes. Modal caps a sandbox at 24 hours, so such a task gets that maximum rather than the Modal SDK's 5-minute default, and a timeout_s above 24 hours is rejected at submission. A verified file gets a verification record beside it, so later tasks that read the same unchanged file check its metadata instead of hashing it again (see Hash once per volume). Host applications remain responsible for retry policy, stalled-task detection, and provider-specific remediation; no reference backend implements recover or honors avoid_hosts yet. A container runtime should fail rather than silently fall back when a required accelerator is unavailable.

To retry a run, submit the same specs again; no backend rejects a spec it has seen before. The local and Modal backends launch a new task. The Kubernetes backend names each Job by a hash of its spec, so it returns the existing Job while that Job is still pending or running, and deletes and recreates it once it has finished, whether it succeeded or failed. A later run under the same job_id therefore always computes its own results. It raises instead of adopting an active Job that a differently configured backend created (see Accept resubmission).

Choose a backend

Backend Best for Requirements
LocalDockerExecutionBackend development and a single workstation Docker daemon; local GPU support for GPU tasks
ModalExecutionBackend managed, elastic execution altar[modal], Modal credentials
KubernetesExecutionBackend existing clusters and shared infrastructure altar[kubernetes], cluster credentials and storage design

See execution and storage integrations for scope and installation.