Skip to workflow guide
PipeDream

PipeDream
Workflow Schema

A structured description of your pipeline. Define commands, inputs, outputs, and resource requirements. PipeDream turns your workflow document into a plan it can schedule and track.

A small document.
A plan for every step.

This version 1 example writes a text file for each batch, then gathers both files into one output. Two command definitions expand into three jobs.

fanout-gather.yaml
# Authoring example: replace YOUR_OUTPUT_DATASET_ID with an authorized binding.
# Submission also requires client credentials and granted organization/project context.
version: 1
name: fanout-gather
output_dataset: YOUR_OUTPUT_DATASET_ID

levels:
  batch:
    iter: for_each
    values: [a, b]

files:
  part: {level: batch, type: txt}
  whole: {type: txt}

commands:
  - name: part
    level: batch
    env: _none
    out: part
    run: echo ${batch} > ${part}

  - name: whole
    env: _none
    in: part
    out: whole
    run: cat ${part} > ${whole}
Download YAMLSchema version 1
The expanded planIllustrative · no jobs are submitted
partbatch = abatch=a/part.txt
partbatch = bbatch=b/part.txt
wholeGather both batcheswhole.txt

The two part jobs are independent and can run concurrently. whole waits for both outputs.

PipeDream derives these dependencies from the levels and file declarations.

This is an authoring example. Replace YOUR_OUTPUT_DATASET_ID with an authorized output binding before submitting. A run also needs client credentials, organization and project access, and a configured compute destination.

Schema essentials

YAML is a convenient way to author the document. The API accepts a parsed JSON object inside a submission envelope. The workflow's version: 1 is separate from the API envelope's schemaVersion.

version
The authored schema version. Use 1 for the format shown here.
name / description
Optional human-readable context for your workflow.
vars
Shared values referenced through ${...} expressions. Expressions use a bounded evaluator, not arbitrary Python or JavaScript.
levels
Collections that expand commands and file declarations. Commands without a level run at the implicit root, _main.
files
Logical file names, types, optional sources, and levels. A command's in and out names refer to these declarations and establish dependencies.
commands
Each command has a name and a run string. Add level, in, out, env, modules, and resources as needed. File declarations do not write files; the command must produce its declared outputs.
output_dataset
The authorized destination binding required by the current authored format when it generates outputs. Native API commands can instead declare output paths directly.
environments / default_env
Declare or reference an execution environment and optionally set a default. Commands can override it with env. The example uses _none and shell tools; it installs no packages.
datasets / modules
Reference authorized datasets to mount or code modules to prepare. Their resource metadata and access must be supplied through the client integration.

Request the resources your command needs

resources:
  cores: 1
  ramGb: 2
  timeout: 600

timeout is in seconds. Admission currently caps requests at 8 CPU and 16 GiB; project policies and provider capacity may impose lower limits. Environment preparation also depends on provider support.

The authored validator currently accepts higher CPU/memory values and a diskGb field, but admission uses the limits above and diskGb is not carried into execution allocation. It should not be relied on as a disk reservation.

Work over collections

A level gives repeated work a shared structure. Files and commands attached to it expand together; a root command can gather outputs from its child instances.

for_each
Create an instance for each value, as with batches a and b above.
cross
Combine every instance from each parent level: a Cartesian product.
zip
Pair parent instances by position. Expansion stops at the shortest parent list.
align
Match parent instances by key. Only keys present in every parent are included; duplicate keys within a parent are rejected.

Levels expand before execution. Runtime-generated tasks, workflow-level conditionals and loops, and subworkflow declarations are not currently supported. Whole-dataset mounts are opaque to the plan; their member files do not automatically create jobs or per-file lineage.

From document to run

  1. Describe the work

    Write a document using the schema, or supply an explicit command graph through the API. Reference inputs and environments through authorized client resources.

  2. Inspect the plan

    POST /v1/plans previews the expanded graph. Validation checks dependencies, cycles, duplicate output producers, and graph limits.

  3. Submit and track

    POST /v1/runs admits work under the client's access and resource policies. PipeDream prepares inputs and environments, schedules commands, and tracks outputs and attempts.

One compute destination per run. Each accepted run is pinned to an authorized destination and its revision. Different runs can use your HPC cluster or managed capacity; a single run does not automatically split across providers.

Input references travel in the workflow document. File contents are transferred separately through authorized staging; PipeDream can stage and relay those transfers.

The current API format identifiers retain their original names: akleao-workflow-yaml-v1 for authored documents and akleao-workflow-flow-v1 for expanded flows. Use the authenticated API reference for the complete submission contract and native command examples.

A schema other formats can map to

PipeDream separates the authored document from the execution graph. That provides a foundation for adapters: translate a supported workflow format into commands, dependencies, files, and resource requests that PipeDream understands.

These mappings are a future direction. Importers and exporters for CWL, WDL, Nextflow, and Snakemake are not implemented today. Running on different compute providers is a separate capability from converting workflow formats.

Common Workflow Language (CWL)
Typed tools, inputs, outputs, and workflow steps. A candidate for mapping file dependencies and simple scatter operations into an execution plan.
Workflow Description Language (WDL)
Tasks and calls, with scatter and conditional execution. Task graphs are a useful starting point; types, expressions, and conditional behavior would need explicit compatibility rules.
Nextflow
Processes connected through dataflow channels. Dynamic channel behavior cannot be assumed to translate into a graph expanded before execution.
Snakemake
File-producing rules and wildcard-based dependencies. Rule resolution and runtime-dependent behavior require more than renaming schema fields.

A useful adapter would declare its supported subset and reject unsupported semantics. Export needs its own checks, too: tenant grants, storage permissions, and destination policies remain part of the execution service configuration.

The schema describes the plan.
The run record tells the story.

PipeDream records accepted manifests and digests, input snapshots, command attempts, provider details, output artifacts, checksums, and ordered run events. These records help explain how a result was produced.

W3C PROV describes entities, activities, and agents involved in producing results. CWLProv applies provenance concepts to CWL runs. They complement a workflow definition rather than replace it.

Standard provenance export is not implemented yet. Existing run records provide a foundation; they do not imply PROV or CWLProv compliance, or observation of every file a command accesses.

Connect your application.

Sign in for the API contract, authentication details, and submission examples.

Open API reference