Back to guides
Intermediate20 min

Building Docker-based pipelines

Learn how to create pipelines using custom Docker images, inputs and outputs.

The pipeline definition

Every pipeline is a Docker image plus a contract: what goes in, what comes out, and how much compute it needs.

{
  "imageUri": "myregistry.io/my-pipeline",
  "imageTag": "1.2.0",
  "command": ["python", "main.py"],
  "inputs": [
    { "name": "input_file", "type": "file", "required": true },
    { "name": "threshold", "type": "string", "required": false, "defaultValue": "0.5" }
  ],
  "outputs": [
    { "name": "results", "type": "file", "path": "/outputs/results.csv" },
    { "name": "report", "type": "file", "path": "/outputs/report.json" }
  ]
}

How inputs reach your container

Input values are passed as environment variables, uppercased: input_file becomes INPUT_FILE, threshold becomes THRESHOLD. File inputs are staged locally and the variable holds the path.

Producing outputs

Write your declared outputs to their path. Marathoon collects them as downloadable artifacts. Write a report.json to get a visual report on the job page.

Resources

{ "cpuUnits": 2, "memoryMb": 4096, "timeoutSeconds": 7200 }

cpuUnits maps to CPU cores. Any parameter can be overridden per job submission. Keep secret values out of parameters and environment variables: see Handling secrets and credentials.

Test before publishing

Use test runs to execute an unpublished pipeline, and validate to check a definition without saving. Once it works, publish it.

Next steps