PipeDream
Workflow Schema
A structured description of your pipeline. Define commands, inputs, outputs, and resource requirements. PipeDream turns your workflow document into a plan it can schedule and track.
A small document.
A plan for every step.
This version 1 example writes a text file for each batch, then gathers both files into one output. Two command definitions expand into three jobs.
# Authoring example: replace YOUR_OUTPUT_DATASET_ID with an authorized binding.
# Submission also requires client credentials and granted organization/project context.
version: 1
name: fanout-gather
output_dataset: YOUR_OUTPUT_DATASET_ID
levels:
batch:
iter: for_each
values: [a, b]
files:
part: {level: batch, type: txt}
whole: {type: txt}
commands:
- name: part
level: batch
env: _none
out: part
run: echo ${batch} > ${part}
- name: whole
env: _none
in: part
out: whole
run: cat ${part} > ${whole}
batch=a/ part.txtbatch=b/ part.txtwhole.txtThe two part jobs are independent and can run concurrently. whole waits for both outputs.
PipeDream derives these dependencies from the levels and file declarations.
This is an authoring example. Replace YOUR_OUTPUT_DATASET_ID with an authorized output binding before submitting. A run also needs client credentials, organization and project access, and a configured compute destination.
Schema essentials
YAML is a convenient way to author the document. The API accepts a parsed JSON object inside a submission envelope. The workflow's version: 1 is separate from the API envelope's schemaVersion.
version- The authored schema version. Use
1for the format shown here. name/description- Optional human-readable context for your workflow.
vars- Shared values referenced through
${...}expressions. Expressions use a bounded evaluator, not arbitrary Python or JavaScript. levels- Collections that expand commands and file declarations. Commands without a level run at the implicit root,
_main. files- Logical file names, types, optional sources, and levels. A command's
inandoutnames refer to these declarations and establish dependencies. commands- Each command has a
nameand arunstring. Addlevel,in,out,env,modules, andresourcesas needed. File declarations do not write files; the command must produce its declared outputs. output_dataset- The authorized destination binding required by the current authored format when it generates outputs. Native API commands can instead declare output paths directly.
environments/default_env- Declare or reference an execution environment and optionally set a default. Commands can override it with
env. The example uses_noneand shell tools; it installs no packages. datasets/modules- Reference authorized datasets to mount or code modules to prepare. Their resource metadata and access must be supplied through the client integration.
Request the resources your command needs
resources:
cores: 1
ramGb: 2
timeout: 600timeout is in seconds. Admission currently caps requests at 8 CPU and 16 GiB; project policies and provider capacity may impose lower limits. Environment preparation also depends on provider support.
The authored validator currently accepts higher CPU/memory values and a diskGb field, but admission uses the limits above and diskGb is not carried into execution allocation. It should not be relied on as a disk reservation.
Work over collections
A level gives repeated work a shared structure. Files and commands attached to it expand together; a root command can gather outputs from its child instances.
for_each- Create an instance for each value, as with batches
aandbabove. cross- Combine every instance from each parent level: a Cartesian product.
zip- Pair parent instances by position. Expansion stops at the shortest parent list.
align- Match parent instances by key. Only keys present in every parent are included; duplicate keys within a parent are rejected.
Levels expand before execution. Runtime-generated tasks, workflow-level conditionals and loops, and subworkflow declarations are not currently supported. Whole-dataset mounts are opaque to the plan; their member files do not automatically create jobs or per-file lineage.
From document to run
Describe the work
Write a document using the schema, or supply an explicit command graph through the API. Reference inputs and environments through authorized client resources.
Inspect the plan
POST /v1/planspreviews the expanded graph. Validation checks dependencies, cycles, duplicate output producers, and graph limits.Submit and track
POST /v1/runsadmits work under the client's access and resource policies. PipeDream prepares inputs and environments, schedules commands, and tracks outputs and attempts.
One compute destination per run. Each accepted run is pinned to an authorized destination and its revision. Different runs can use your HPC cluster or managed capacity; a single run does not automatically split across providers.
Input references travel in the workflow document. File contents are transferred separately through authorized staging; PipeDream can stage and relay those transfers.
The current API format identifiers retain their original names: akleao-workflow-yaml-v1 for authored documents and akleao-workflow-flow-v1 for expanded flows. Use the authenticated API reference for the complete submission contract and native command examples.
A schema other formats can map to
PipeDream separates the authored document from the execution graph. That provides a foundation for adapters: translate a supported workflow format into commands, dependencies, files, and resource requests that PipeDream understands.
These mappings are a future direction. Importers and exporters for CWL, WDL, Nextflow, and Snakemake are not implemented today. Running on different compute providers is a separate capability from converting workflow formats.
- Common Workflow Language (CWL)
- Typed tools, inputs, outputs, and workflow steps. A candidate for mapping file dependencies and simple scatter operations into an execution plan.
- Workflow Description Language (WDL)
- Tasks and calls, with scatter and conditional execution. Task graphs are a useful starting point; types, expressions, and conditional behavior would need explicit compatibility rules.
- Nextflow
- Processes connected through dataflow channels. Dynamic channel behavior cannot be assumed to translate into a graph expanded before execution.
- Snakemake
- File-producing rules and wildcard-based dependencies. Rule resolution and runtime-dependent behavior require more than renaming schema fields.
A useful adapter would declare its supported subset and reject unsupported semantics. Export needs its own checks, too: tenant grants, storage permissions, and destination policies remain part of the execution service configuration.
The schema describes the plan.
The run record tells the story.
PipeDream records accepted manifests and digests, input snapshots, command attempts, provider details, output artifacts, checksums, and ordered run events. These records help explain how a result was produced.
W3C PROV describes entities, activities, and agents involved in producing results. CWLProv applies provenance concepts to CWL runs. They complement a workflow definition rather than replace it.
Standard provenance export is not implemented yet. Existing run records provide a foundation; they do not imply PROV or CWLProv compliance, or observation of every file a command accesses.
Connect your application.
Sign in for the API contract, authentication details, and submission examples.
Open API reference