Documentation

Complete reference for building and running pipelines on Marathoon.

Overview

Marathoon is a multi-tenant pipeline orchestration platform. You define pipelines as Docker-based workflows, submit jobs via the UI or API, and get structured results including logs, artifacts, and visual reports.

Projects & Pipelines

Organize work into projects; define reusable pipelines.

Jobs & Batches

Run pipelines on demand or submit hundreds at once.

Visual Reports

Render charts, tables, and metrics directly on job pages.

Base API URL

https://<your-domain>/api/v1/tenants/<tenantId>

Getting Started

1. Create a Project

A Project is the top-level container for pipelines. Go to Projects → New Project in the dashboard.

Projects control access: team members are granted access per-project with roles (Owner, Maintainer, Analyst, Viewer).

2. Define a Pipeline

A Pipeline defines a Docker image, its inputs/outputs, and the command to run. Use the visual pipeline editor or POST via API.

json
{
  "name": "My Pipeline",
  "description": "Processes input files and produces a report",
  "definition": {
    "imageUri": "myregistry.io/my-pipeline",
    "imageTag": "latest",
    "command": ["python", "main.py"],
    "inputs": [
      { "name": "input_file", "type": "file", "required": true },
      { "name": "threshold",  "type": "string", "required": false, "defaultValue": "0.5" }
    ],
    "outputs": [
      { "name": "results", "type": "file", "path": "/outputs/results.csv" },
      { "name": "report",  "type": "file", "path": "/outputs/report.json" }
    ]
  },
  "defaultParameters": {
    "cpuUnits": 2,
    "memoryMb": 2048,
    "timeoutSeconds": 3600,
    "environmentVariables": {},
    "secretRefs": []
  }
}

After creation, Publish the pipeline to make it runnable.

3. Submit a Job

Submit a job via the UI (Run button on the pipeline page) or via API:

bash
curl -X POST \
  https://<domain>/api/v1/tenants/<tenantId>/jobs/run/<pipelineId> \
  -H "X-API-Key: <api-key>" \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": {
      "input_file": "s3://my-bucket/data.csv",
      "threshold": "0.8"
    }
  }'

4. Monitor & Retrieve Results

Track job status in real-time via the Jobs page (SignalR live updates). Each job shows:

  • Status timeline (Queued → Running → Succeeded / Failed)
  • Structured logs (searchable, paginated)
  • Container stdout/stderr
  • Output artifacts (downloadable)
  • Visual report (if report.json produced)

Core Concepts

Projects

Logical grouping for pipelines. Members access projects through roles. A project can be archived, which blocks new job submissions while preserving history.

Pipelines

Versioned Docker-based workflow definitions. Each publish increments the version number. Old versions are preserved and jobs reference the version they ran against. Pipelines must be published before users can submit jobs.

Jobs & Batches

A Job is a single pipeline execution with specific inputs. A Batch groups multiple jobs submitted together (e.g., processing 1000 files). Jobs transition through: Queued → Running → Succeeded | Failed | Cancelled.

Credits

Credits are the billing unit (100 credits = $1). Each job submission deducts credits based on plan limits. Credit balance is visible in the sidebar. Plans: Free (200), Research (1,500), Pro (7,500), Enterprise (custom). Insufficient credits returns 402 Payment Required.

Permissions

Two-layer authorization: JWT (user session) and API keys (per-key permission set). Permissions are scoped to the tenant. Key presets: ReadOnly, Operator, Admin, or custom granular permissions.

Pipeline Definition

The pipeline definition is the core schema stored with each pipeline. It is versioned on every update.

Top-Level Fields

FieldTypeDescription
imageUristring (required)Docker image URI (without tag)
imageTagstringDocker image tag, defaults to latest
commandstring[]Override the container entrypoint command
entryPointstringAlternative entrypoint override
inputsPipelineInput[]Declared input slots
outputsPipelineOutput[]Declared output slots
tasksPipelineTask[]Graph nodes (visual editor)
connectionsPipelineConnection[]Graph edges (visual editor)

Inputs & Outputs

json
{
  "inputs": [
    {
      "name": "input_file",
      "type": "file",
      "required": true,
      "description": "Input CSV to process"
    },
    {
      "name": "mode",
      "type": "string",
      "required": false,
      "defaultValue": "fast"
    }
  ],
  "outputs": [
    {
      "name": "results",
      "type": "file",
      "path": "/outputs/results.csv",
      "description": "Processed output"
    }
  ]
}

Input values are passed as environment variables to the container: INPUT_FILE, MODE, etc. (uppercased).

Parameters (Resources)

json
{
  "cpuUnits": 2,
  "memoryMb": 4096,
  "timeoutSeconds": 7200,
  "environmentVariables": {
    "LOG_LEVEL": "info"
  },
  "secretRefs": []
}

cpuUnits maps to CPU cores. Parameters can be overridden per job submission. secretRefs is reserved for managed pipeline secrets, which are not available yet: never put secret values in parameters or environment variables.

Visual Reports

Any pipeline can produce a rich visual report. Write a file named report.json to the output directory. Marathoon auto-detects it and renders it on the job page.

Top-Level Schema

json
{
  "id": "report-001",
  "jobId": "<jobId>",
  "title": "Analysis Results",
  "description": "Optional subtitle shown below the title",
  "generatedAtUtc": "2026-06-17T12:00:00Z",
  "blocks": [ ... ],
  "tabs": [
    { "label": "Overview", "blocks": [ ... ] },
    { "label": "Details",  "blocks": [ ... ] }
  ]
}

Use tabs for multi-section reports. When present, the renderer shows a tab bar instead of a flat block list. blocks at root level is used when no tabs.

Block Types

text / markdown
json
{ "id":"b1","order":1,"type":"text","data":{ "content":"Hello world","style":"heading1" } }
{ "id":"b2","order":2,"type":"markdown","data":{ "markdown":"**Bold** and _italic_" } }

text styles: paragraph, heading1, heading2, heading3, quote

stats
json
{
  "id": "b3", "order": 3, "type": "stats",
  "data": {
    "stats": [
      { "label": "Accuracy", "value": "94.2", "unit": "%", "change": "+2.1%", "changeDirection": "up" },
      { "label": "Samples",  "value": "10000" }
    ]
  }
}
table
json
{
  "id": "b4", "order": 4, "type": "table",
  "data": {
    "columns": [
      { "key": "name",  "header": "Name",  "type": "text",   "sortable": true },
      { "key": "score", "header": "Score", "type": "number", "sortable": true },
      { "key": "pass",  "header": "Pass",  "type": "boolean" }
    ],
    "rows": [
      { "name": "Sample A", "score": 0.92, "pass": true },
      { "name": "Sample B", "score": 0.51, "pass": false }
    ],
    "sortable": true,
    "filterable": true,
    "paginated": true,
    "pageSize": 20
  }
}

Column types: text number date boolean badge link progress

chart
json
{
  "id": "b5", "order": 5, "type": "chart", "title": "Loss Curve",
  "data": {
    "chartType": "line",
    "series": [
      {
        "name": "Train",
        "data": [{"x":1,"y":0.9},{"x":2,"y":0.7},{"x":3,"y":0.5}]
      },
      {
        "name": "Val",
        "data": [{"x":1,"y":0.95},{"x":2,"y":0.75},{"x":3,"y":0.6}]
      }
    ],
    "options": { "xAxisLabel": "Epoch", "yAxisLabel": "Loss", "showLegend": true }
  }
}

chartType: line bar area pie donut scatter radar

vega
json
{
  "id": "b6", "order": 6, "type": "vega",
  "data": {
    "spec": {
      "$schema": "https://vega.github.io/schema/vega-lite/v5.json",
      "mark": "bar",
      "data": { "values": [{"x":"A","y":3},{"x":"B","y":7}] },
      "encoding": {
        "x": { "field": "x", "type": "ordinal" },
        "y": { "field": "y", "type": "quantitative" }
      }
    }
  }
}

Full Vega-Lite (or Vega) spec. Use for advanced / custom visualizations.

image
json
{
  "id": "b7", "order": 7, "type": "image",
  "data": { "url": "/output/confusion_matrix.png", "altText": "Confusion matrix" }
}

URL references another output artifact by its path. Marathoon resolves it from blob storage.

alert / progress
json
{ "id":"b8","order":8,"type":"alert",   "data":{"message":"Dataset has nulls","severity":"warning"} }
{ "id":"b9","order":9,"type":"progress","data":{"value":750,"max":1000,"unit":"samples"} }

alert severities: info success warning error

Example Pipelines

Four pipelines ready to import from Pipelines → View examples, or via Import JSON. Each pairs with the file agent: drop a sub-folder → a job starts automatically.

Comment Marathoon échange des fichiers avec votre conteneur

Chaque conteneur reçoit ses données via deux répertoires montés :

RépertoireRôle
/inputsFichiers uploadés par l'agent ou soumis manuellement
/outputsTout fichier déposé ici devient un artefact téléchargeable

Variables d'environnement. Les inputs scalaires (type string) sont injectés en majuscules : mode: "fast" → MODE=fast dans l'environnement du conteneur.

Rapport visuel. Si votre script écrit /outputs/report.json au format Marathoon, la page du job l'affiche automatiquement sous forme de graphiques, tableaux et métriques.

Sorties déclarées. Le champ path d'une sortie (ex. /outputs/report.json) indique à Marathoon l'artefact principal à mettre en avant. Les autres fichiers dans /outputs sont quand même capturés.

1. Quick Check

Pipeline minimal — aucune entrée, un job rapide. Commencez ici.

Pipeline le plus simple possible : il démarre, simule une courte analyse et s'arrête avec succès.

Objectif pédagogique : vérifier que votre agent est bien appairé et que déposer un dossier déclenche bien un job de bout en bout, avant de passer aux exemples avec fichiers.

Concept clé : même sans entrées déclarées, le conteneur peut écrire dans /outputs — tout fichier y déposé devient un artefact téléchargeable sur la page du job.

Avec l'agent fichier : déposez n'importe quel sous-dossier dans le répertoire surveillé. Chaque sous-dossier = un job. Le contenu est ignoré.

ADockerfile — how the image is built

dockerfile
# Image de base légère — Alpine Linux avec bash, curl, jq
FROM alpine:3.19
RUN apk add --no-cache bash curl jq

WORKDIR /app
COPY analyze.sh /app/analyze.sh
RUN chmod +x /app/analyze.sh

# Paramètres tunable via environmentVariables à la soumission du job
ENV ANALYSIS_DURATION=30 \
    OUTPUT_DIR=/outputs

ENTRYPOINT ["/app/analyze.sh"]

BEntrypoint script: analyze.sh

bash
#!/bin/bash
set -e

# ── Configuration ──────────────────────────────────────────────────────────
# Marathoon injecte les paramètres de la pipeline en variables d'environnement
# (uppercased). Ici on peut les surcharger via parameterOverrides à la soumission.
DURATION=${ANALYSIS_DURATION:-30}
OUTPUT_DIR=${OUTPUT_DIR:-/outputs}

mkdir -p "$OUTPUT_DIR"

echo "[INFO] Démarrage — durée simulée : ${DURATION}s"

# ── Travail simulé ──────────────────────────────────────────────────────────
for step in "Initialisation" "Chargement" "Traitement" "Agrégation"; do
  echo "[INFO] $step..."
  sleep $((DURATION / 4))
done

# ── Sorties ─────────────────────────────────────────────────────────────────
# Tout fichier écrit dans /outputs devient un artefact téléchargeable.

cat > "$OUTPUT_DIR/summary.json" <<JSON
{
  "status": "completed",
  "duration_seconds": $DURATION,
  "records_processed": 985432,
  "records_skipped": 14568
}
JSON

cat > "$OUTPUT_DIR/metrics.csv" <<CSV
metric,value,unit
records_processed,985432,count
records_skipped,14568,count
execution_time,$DURATION,seconds
CSV

echo "[INFO] Terminé. Résultats dans $OUTPUT_DIR/"

CInputs / Outputs

Entrées : aucune. Le conteneur ignore le contenu du dossier déposé — il a juste besoin d'un déclencheur.

Sorties déclarées :

  • result → /outputs/result.json (déclaré dans la pipeline)

Sorties produites réellement :

  • /outputs/summary.json — résumé JSON de l'exécution
  • /outputs/metrics.csv — métriques au format CSV

Ces deux fichiers apparaissent comme artefacts même s'ils ne sont pas déclarés explicitement dans la définition de la pipeline.

Folder to drop into the watched directory

text
watch/
└── run-001/
    └── anything.txt

Pipeline definition (importable JSON)

json
{
  "name": "Quick Check",
  "description": "Minimal sample pipeline — fast run, no inputs.",
  "definition": {
    "imageUri": "marathoon-analyzer-simple",
    "imageTag": "latest",
    "inputs": [],
    "outputs": [
      {
        "name": "result",
        "type": "file",
        "path": "/outputs/result.json",
        "description": "Run result"
      }
    ],
    "entryPoint": null,
    "command": [],
    "tasks": [],
    "connections": []
  },
  "defaultParameters": {
    "cpuUnits": 1,
    "memoryMb": 256,
    "timeoutSeconds": 120,
    "environmentVariables": {},
    "secretRefs": []
  }
}

2. File Analyzer

Lit tous les fichiers d'un dossier et produit un rapport d'analyse.

Traite l'ensemble des fichiers déposés dans /inputs et génère un rapport d'analyse par fichier, plus un rapport agrégé.

Objectif pédagogique : illustrer le binding mount-all — l'agent uploade tous les fichiers d'un sous-dossier dans /inputs, sans nommage requis. Le script parcourt le répertoire entier.

Concepts clés :

  • /inputs est peuplé automatiquement par Marathoon avec les fichiers uploadés
  • Le script peut traiter n'importe quel type de fichier (texte, JSON, CSV, binaire)
  • Chaque fichier produit son propre rapport JSON dans /outputs/reports/

Avec l'agent fichier : déposez un dossier contenant des fichiers à analyser. Tous les fichiers (n'importe quel type) sont envoyés ensemble dans un seul job.

ADockerfile — how the image is built

dockerfile
FROM alpine:3.19

# python3 pour le script de traitement par type de fichier
RUN apk add --no-cache bash curl jq python3 py3-pip

WORKDIR /app
# Marathoon monte /inputs et /outputs — on les crée pour la clarté
RUN mkdir -p /inputs /outputs

COPY analyze.sh process_file.py /app/
RUN chmod +x /app/analyze.sh

ENV INPUT_DIR=/inputs \
    OUTPUT_DIR=/outputs

# VOLUME déclare les points de montage (optionnel mais explicite)
VOLUME ["/inputs", "/outputs"]

ENTRYPOINT ["/app/analyze.sh"]

BEntrypoint script: analyze.sh

bash
#!/bin/bash
set -e

# ── Configuration ──────────────────────────────────────────────────────────
INPUT_DIR=${INPUT_DIR:-/inputs}
OUTPUT_DIR=${OUTPUT_DIR:-/outputs}

mkdir -p "$OUTPUT_DIR/reports"

# ── Inventaire des entrées ──────────────────────────────────────────────────
echo "[INFO] Scan de $INPUT_DIR..."
FILE_COUNT=$(find "$INPUT_DIR" -type f | wc -l | tr -d ' ')

if [ "$FILE_COUNT" -eq 0 ]; then
  echo "[WARN] Aucun fichier dans $INPUT_DIR"
  exit 0
fi

echo "[INFO] $FILE_COUNT fichier(s) trouvé(s)"

# ── Traitement fichier par fichier ──────────────────────────────────────────
# process_file.py détecte le type (txt, json, csv, binaire) et
# génère un rapport JSON d'analyse pour chaque fichier.
find "$INPUT_DIR" -type f | while read -r file; do
  name=$(basename "$file")
  echo "[INFO] Traitement : $name"
  out="$OUTPUT_DIR/reports/${name%.*}_analysis.json"
  python3 /app/process_file.py "$file" "$out"
done

# ── Rapport agrégé ──────────────────────────────────────────────────────────
DONE=$(find "$OUTPUT_DIR/reports" -name "*.json" | wc -l | tr -d ' ')

cat > "$OUTPUT_DIR/aggregate_report.json" <<JSON
{
  "total_files_processed": $DONE,
  "input_directory": "$INPUT_DIR",
  "output_directory": "$OUTPUT_DIR",
  "status": "completed"
}
JSON

echo "[INFO] Rapport agrégé écrit. $DONE fichier(s) traité(s)."

CEntrypoint script: process_file.py

python
#!/usr/bin/env python3
"""
Analyse un fichier selon son type (texte, JSON, CSV, binaire)
et écrit un rapport JSON dans le répertoire de sortie.

Usage: python3 process_file.py <input_path> <output_path>
"""

import json
import sys
import os
import hashlib
from datetime import datetime
from pathlib import Path


def get_file_info(filepath: str) -> dict:
    path = Path(filepath)
    stat = path.stat()
    with open(filepath, 'rb') as f:
        file_hash = hashlib.sha256(f.read()).hexdigest()
    return {
        "filename": path.name,
        "size_bytes": stat.st_size,
        "extension": path.suffix.lower(),
        "sha256": file_hash[:16] + "...",
    }


def analyze_text(filepath: str) -> dict:
    with open(filepath, 'r', encoding='utf-8', errors='ignore') as f:
        content = f.read()
    lines = content.split('\n')
    return {
        "type": "text",
        "lines": len(lines),
        "words": len(content.split()),
        "characters": len(content),
        "non_empty_lines": len([l for l in lines if l.strip()]),
    }


def analyze_json(filepath: str) -> dict:
    with open(filepath, 'r', encoding='utf-8') as f:
        data = json.load(f)

    def depth(obj, d=0):
        if isinstance(obj, dict):
            return max((depth(v, d + 1) for v in obj.values()), default=d)
        if isinstance(obj, list):
            return max((depth(v, d + 1) for v in obj), default=d)
        return d

    return {
        "type": "json",
        "root_type": type(data).__name__,
        "keys": len(data) if isinstance(data, dict) else None,
        "items": len(data) if isinstance(data, list) else None,
        "max_depth": depth(data),
    }


def analyze_csv(filepath: str) -> dict:
    with open(filepath, 'r', encoding='utf-8', errors='ignore') as f:
        lines = f.readlines()
    if not lines:
        return {"type": "csv", "rows": 0, "columns": 0}
    first = lines[0]
    delimiter = ',' if ',' in first else ('\t' if '\t' in first else ';')
    return {
        "type": "csv",
        "rows": len(lines),
        "columns": len(first.split(delimiter)),
        "delimiter": repr(delimiter),
    }


def analyze_binary(filepath: str) -> dict:
    with open(filepath, 'rb') as f:
        header = f.read(4)
    return {
        "type": "binary",
        "magic": header.hex(),
    }


def process(filepath: str) -> dict:
    info = get_file_info(filepath)
    ext = info["extension"]
    if ext in ('.txt', '.log', '.md', '.py', '.js', '.sh'):
        analysis = analyze_text(filepath)
    elif ext == '.json':
        try:
            analysis = analyze_json(filepath)
        except json.JSONDecodeError:
            analysis = {"type": "json", "error": "Invalid JSON"}
    elif ext in ('.csv', '.tsv'):
        analysis = analyze_csv(filepath)
    else:
        analysis = analyze_binary(filepath)
    return {
        "file_info": info,
        "analysis": analysis,
        "processed_at": datetime.utcnow().isoformat() + "Z",
    }


def main():
    if len(sys.argv) < 2:
        print(json.dumps({"error": "Usage: process_file.py <input> [output]"}))
        sys.exit(1)

    filepath = sys.argv[1]
    output_path = sys.argv[2] if len(sys.argv) > 2 else None

    if not os.path.exists(filepath):
        print(json.dumps({"error": f"File not found: {filepath}"}))
        sys.exit(1)

    result = process(filepath)

    if output_path:
        os.makedirs(os.path.dirname(output_path), exist_ok=True)
        with open(output_path, 'w') as f:
            json.dump(result, f, indent=2)
        print(f"[INFO] Rapport écrit : {output_path}")
    else:
        print(json.dumps(result, indent=2))


if __name__ == "__main__":
    main()

DInputs / Outputs

Entrées : tous les fichiers présents dans /inputs (binding mount-all). Aucun nommage requis — le script parcourt le répertoire entier.

Types supportés : .txt, .log, .md, .py, .js, .sh, .json, .csv, .tsv, binaires.

Sorties produites :

  • /outputs/reports/<nom>_analysis.json — un rapport par fichier
  • /outputs/aggregate_report.json — résumé global

Sortie déclarée : report → /outputs/analysis-report.json

Folder to drop into the watched directory

text
watch/
└── batch-2024-06/
    ├── document-a.txt
    ├── document-b.txt
    └── data.json

Pipeline definition (importable JSON)

json
{
  "name": "File Analyzer",
  "description": "Reads all files in /inputs and emits an analysis report.",
  "definition": {
    "imageUri": "marathoon-analyzer-file",
    "imageTag": "1.0",
    "inputs": [],
    "outputs": [
      {
        "name": "report",
        "type": "file",
        "path": "/outputs/analysis-report.json",
        "description": "Analysis report"
      }
    ],
    "entryPoint": null,
    "command": [],
    "tasks": [],
    "connections": []
  },
  "defaultParameters": {
    "cpuUnits": 2,
    "memoryMb": 512,
    "timeoutSeconds": 300,
    "environmentVariables": {
      "ANALYSIS_DURATION": "5",
      "LOG_INTERVAL": "1"
    },
    "secretRefs": []
  }
}

3. CSV → JSON

Convertit chaque fichier CSV en JSON et affiche un rapport de conversion.

Pour chaque *.csv dans /inputs, écrit un <nom>.json correspondant (tableau d'objets ligne) dans /outputs, plus un report.json rendu automatiquement sur la page du job.

Objectif pédagogique : montrer deux mécanismes de binding en parallèle :

  • Name-match binding : un input nommé data lie le fichier data.csv automatiquement
  • Mount-all : tout autre *.csv est quand même traité via le scan du répertoire

Concept clé : le report.json au format Marathoon est détecté automatiquement par la plateforme et rendu sur la page du job avec stats + tableau.

Avec l'agent fichier : déposez un dossier contenant des fichiers CSV.

ADockerfile — how the image is built

dockerfile
# Python 3.12 alpine — stdlib csv/json suffisent, aucune dépendance externe
FROM python:3.12-alpine

WORKDIR /app
RUN mkdir -p /inputs /outputs

COPY convert.py /app/convert.py

ENV INPUT_DIR=/inputs \
    OUTPUT_DIR=/outputs

VOLUME ["/inputs", "/outputs"]

ENTRYPOINT ["python", "/app/convert.py"]

BEntrypoint script: convert.py

python
#!/usr/bin/env python3
"""
CSV → JSON converter sample.

Lit tous les *.csv sous INPUT_DIR (récursivement), écrit un <nom>.json
(tableau d'objets ligne) pour chacun dans OUTPUT_DIR, puis génère un
report.json au format Marathoon — rendu automatiquement sur la page du job.
"""
import csv
import json
import os
import sys
from datetime import datetime, timezone


INPUT_DIR  = os.environ.get("INPUT_DIR",  "/inputs")
OUTPUT_DIR = os.environ.get("OUTPUT_DIR", "/outputs")


def log(msg: str) -> None:
    ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%S.%fZ")
    print(f"[INFO] {ts} {msg}", flush=True)


def find_csv_files(root: str):
    """Parcourt récursivement root et yield chaque *.csv trouvé."""
    for dirpath, _dirs, files in os.walk(root):
        for name in files:
            if name.lower().endswith(".csv"):
                yield os.path.join(dirpath, name)


def convert_file(path: str) -> dict:
    """Convertit un CSV en liste de dicts et l'écrit dans OUTPUT_DIR."""
    rel = os.path.relpath(path, INPUT_DIR)
    try:
        with open(path, newline="", encoding="utf-8-sig") as fh:
            rows = list(csv.DictReader(fh))
    except Exception as exc:
        log(f"ERREUR lecture {rel}: {exc}")
        return {"file": rel, "rows": 0, "status": f"error: {exc}"}

    out_name = os.path.splitext(os.path.basename(path))[0] + ".json"
    out_path = os.path.join(OUTPUT_DIR, out_name)

    with open(out_path, "w", encoding="utf-8") as out:
        json.dump(rows, out, indent=2, ensure_ascii=False)

    log(f"Converti {rel} → {out_name} ({len(rows)} lignes)")
    return {"file": rel, "rows": len(rows), "status": "ok"}


def write_report(summary: list) -> None:
    """
    Écrit /outputs/report.json au format Marathoon.
    La plateforme détecte ce fichier et le rend sur la page du job
    sous forme de stats + tableau.
    """
    total_rows = sum(s["rows"] for s in summary)
    report = {
        "title": "CSV → JSON conversion",
        "description": "Un fichier JSON écrit par CSV converti.",
        "generatedAtUtc": datetime.now(timezone.utc).isoformat(),
        "blocks": [
            {
                "type": "stats",
                "stats": [
                    {"label": "Fichiers convertis", "value": str(len(summary))},
                    {"label": "Lignes totales",     "value": str(total_rows)},
                ],
            },
            {
                "type": "table",
                "title": "Fichiers convertis",
                "columns": [
                    {"key": "file",   "header": "Fichier"},
                    {"key": "rows",   "header": "Lignes"},
                    {"key": "status", "header": "Statut"},
                ],
                "rows": summary,
            },
        ],
    }
    out = os.path.join(OUTPUT_DIR, "report.json")
    with open(out, "w", encoding="utf-8") as fh:
        json.dump(report, fh, indent=2)
    log(f"report.json écrit ({os.path.getsize(out)} octets)")


def main() -> int:
    os.makedirs(OUTPUT_DIR, exist_ok=True)
    log(f"Scan de {INPUT_DIR} pour les fichiers CSV")

    csv_files = sorted(find_csv_files(INPUT_DIR))
    if not csv_files:
        log("Aucun CSV trouvé — rien à convertir")
        write_report([])
        return 0

    summary = [convert_file(path) for path in csv_files]
    write_report(summary)
    log(f"Terminé : {len(summary)} fichier(s) traité(s)")
    return 0


if __name__ == "__main__":
    sys.exit(main())

CInputs / Outputs

Entrées : tous les *.csv dans /inputs (récursif). L'input nommé data permet de binder data.csv explicitement, mais tous les autres CSV sont quand même traités.

Sorties produites :

  • /outputs/<nom>.json — un tableau JSON de lignes par CSV
  • /outputs/report.json — rapport de conversion (stats + tableau), rendu automatiquement sur la page du job

Sortie déclarée : report → /outputs/report.json

Folder to drop into the watched directory

text
watch/
└── export-001/
    ├── data.csv
    └── more.csv

Pipeline definition (importable JSON)

json
{
  "name": "CSV to JSON",
  "description": "Converts CSV files to JSON with a conversion report.",
  "definition": {
    "imageUri": "marathoon-csv-json",
    "imageTag": "1.0",
    "inputs": [
      {
        "name": "data",
        "type": "file",
        "required": false,
        "defaultValue": null,
        "description": "A CSV file to convert"
      }
    ],
    "outputs": [
      {
        "name": "report",
        "type": "file",
        "path": "/outputs/report.json",
        "description": "Conversion report"
      }
    ],
    "entryPoint": null,
    "command": [],
    "tasks": [],
    "connections": []
  },
  "defaultParameters": {
    "cpuUnits": 1,
    "memoryMb": 256,
    "timeoutSeconds": 120,
    "environmentVariables": {},
    "secretRefs": []
  }
}

4. Visual Report

Émet un report.json qui se rend en graphiques et tableaux sur la page du job.

Génère un report.json couvrant tous les types de blocs Marathoon : stats, alerte, markdown, graphiques (bar, donut), Vega-Lite, tableau.

Objectif pédagogique : servir de référence complète pour le schéma de rapport visuel. Ouvrez le job terminé pour voir le rendu, puis copiez les blocs qui vous intéressent.

Types de blocs disponibles :

TypeDescription
statsMétriques avec valeur, unité, variation
alertBandeau info/success/warning/error
markdownTexte formaté, titres, listes
chartGraphiques (line, bar, area, pie, donut, scatter, radar)
vegaSpec Vega-Lite complète pour visualisations avancées
tableTableau trié/filtré/paginé
imageImage depuis un artefact de sortie
progressBarre de progression

Avec l'agent fichier : déposez n'importe quel dossier pour déclencher un run — le rapport est généré par la pipeline elle-même.

ADockerfile — how the image is built

dockerfile
# BusyBox suffit — le script est en shell pur
FROM busybox:latest

COPY report.sh /report.sh
RUN chmod +x /report.sh

ENTRYPOINT ["/report.sh"]

BEntrypoint script: report.sh

bash
#!/bin/sh
set -e

# Aucune entrée requise — ce pipeline génère ses propres données.
mkdir -p /outputs

# ── Génération du rapport visuel ────────────────────────────────────────────
# Marathoon détecte report.json automatiquement et le rend sur la page du job.
# Utilisez des tabs pour organiser plusieurs sections.
cat > /outputs/report.json <<'JSON'
{
  "title": "Sample Analysis Report",
  "description": "Généré par le pipeline Visual Report pour illustrer le schéma de rapport.",
  "tabs": [
    {
      "label": "Overview",
      "blocks": [
        {
          "type": "stats",
          "title": "Métriques clés",
          "stats": [
            { "label": "Lignes traitées", "value": "12,840" },
            { "label": "Erreurs",         "value": "3",    "changeDirection": "down", "change": "-40%" },
            { "label": "Durée",           "value": "8.2",  "unit": "s" },
            { "label": "Précision",       "value": "98.6", "unit": "%", "changeDirection": "up", "change": "+1.2%" }
          ]
        },
        {
          "type": "alert",
          "severity": "success",
          "message": "Toutes les validations sont passées."
        },
        {
          "type": "markdown",
          "markdown": "## Résumé\n\nLa pipeline a traité le jeu de données avec succès.\n\n- Deux valeurs aberrantes supprimées\n- Tendance **positive** trimestre sur trimestre"
        }
      ]
    },
    {
      "label": "Graphiques",
      "blocks": [
        {
          "type": "chart",
          "title": "Débit mensuel",
          "chartType": "bar",
          "series": [
            { "name": "Jobs", "data": [{"x":0,"y":120},{"x":1,"y":180},{"x":2,"y":150},{"x":3,"y":240}] }
          ],
          "options": { "xAxisLabel": "Mois", "yAxisLabel": "Jobs", "showGrid": true, "showLegend": true }
        },
        {
          "type": "chart",
          "title": "Succès vs Échec",
          "chartType": "donut",
          "series": [
            { "name": "Résultat", "data": [{"label":"Succès","x":0,"y":92},{"label":"Échec","x":1,"y":8}] }
          ],
          "options": { "showLegend": true }
        },
        {
          "type": "vega",
          "title": "Heatmap de corrélation (Vega-Lite)",
          "spec": {
            "$schema": "https://vega.github.io/schema/vega-lite/v5.json",
            "data": { "values": [
              {"x":"A","y":"A","v":1.0},{"x":"A","y":"B","v":0.3},{"x":"A","y":"C","v":0.6},
              {"x":"B","y":"A","v":0.3},{"x":"B","y":"B","v":1.0},{"x":"B","y":"C","v":0.1},
              {"x":"C","y":"A","v":0.6},{"x":"C","y":"B","v":0.1},{"x":"C","y":"C","v":1.0}
            ]},
            "mark": "rect",
            "height": 200,
            "encoding": {
              "x": {"field":"x","type":"nominal"},
              "y": {"field":"y","type":"nominal"},
              "color": {"field":"v","type":"quantitative","scale":{"scheme":"viridis"}}
            }
          }
        }
      ]
    },
    {
      "label": "Données",
      "blocks": [
        {
          "type": "table",
          "title": "Top résultats",
          "sortable": true,
          "filterable": true,
          "columns": [
            { "key": "name",   "header": "Nom",      "type": "text" },
            { "key": "score",  "header": "Score",    "type": "number" },
            { "key": "passed", "header": "Validé",   "type": "boolean" }
          ],
          "rows": [
            { "name": "alpha", "score": 0.98, "passed": true  },
            { "name": "beta",  "score": 0.71, "passed": true  },
            { "name": "gamma", "score": 0.42, "passed": false }
          ]
        }
      ]
    }
  ]
}
JSON

echo "report.json écrit ($(wc -c < /outputs/report.json) octets)"

CInputs / Outputs

Entrées : aucune — le dossier déposé sert uniquement de déclencheur.

Sorties produites :

  • /outputs/report.json — rapport multi-tabs au format Marathoon, rendu automatiquement sur la page du job

Sortie déclarée : report → /outputs/report.json

Pour votre propre pipeline : générez vos graphiques avec matplotlib/seaborn/Vega-Lite, sérialisez le résultat dans le schéma Marathoon, et écrivez-le dans /outputs/report.json.

Folder to drop into the watched directory

text
watch/
└── report-run-1/
    └── trigger.txt

Pipeline definition (importable JSON)

json
{
  "name": "Visual Report",
  "description": "Emits a report.json rendered as charts/tables on the job page.",
  "definition": {
    "imageUri": "marathoon-sample-report",
    "imageTag": "latest",
    "inputs": [],
    "outputs": [
      {
        "name": "report",
        "type": "file",
        "path": "/outputs/report.json",
        "description": "Visual report"
      }
    ],
    "entryPoint": null,
    "command": [],
    "tasks": [],
    "connections": []
  },
  "defaultParameters": {
    "cpuUnits": 1,
    "memoryMb": 256,
    "timeoutSeconds": 120,
    "environmentVariables": {},
    "secretRefs": []
  }
}

Bring Your Own Infrastructure

By default, jobs run on Marathoon's hosted compute and outputs are stored in hosted blob storage. You can instead route a project's jobs to your own Docker host and keep files in your own object storage. When compute is yours, there is no per-run credit charge — only a flat platform fee for orchestration, queueing and the UI. Marketplace pipelines are the exception: they always run on Marathoon's hosted compute.

Connect a runner

Open Settings → Infrastructure → Connect infrastructure → Runner. The guided setup creates a Connected Runner provider, mints a one-time enrollment token (15 minutes, single use) and shows the exact command to run on your machine: a Docker one-liner, a Linux package, or the Windows installer. The machine enrols itself, sends a heartbeat, and the wizard ends with a test job executed end to end on your hardware.

Each project resolves a compute target and a storage target. Override either per project; unset falls back to the tenant default, which falls back to hosted.

How the Connected Runner works

The runner is a small agent that runs next to a Docker engine you own — or inside a Kubernetes cluster. It is pull-only: it polls Marathoon for jobs assigned to it, runs each job in a container (on your Docker, or as a Kubernetes Job), stages inputs and uploads the outputs. Marathoon never connects to your machine.

  • Enrolment exchanges the one-time token for an API key limited to the runner endpoints (it cannot submit jobs or read pipelines) and bound to the machine's agent id.
  • Every job is leased to one runner at a time; a runner that stops heartbeating loses the lease and the job is re-queued.
  • Inputs and outputs are streamed through the Marathoon API to and from your storage provider. Container logs are relayed as they are produced.
  • Revoke the runner at any time by deleting its API key or the agent: it stops receiving work on its next poll.

Network prerequisites

The runner needs outbound HTTPS only. Nothing has to be opened on your side.

DirectionTargetPurpose
OutboundMarathoon API (443)Enrolment, heartbeats, job claims, input download, output upload
OutboundYour image registries (443)Pulling the job images your pipelines reference
InboundnoneNo port, no public IP, no VPN required

Use your own storage

Point a project at your own S3-compatible bucket or Azure Blob container so job inputs and outputs are kept in storage you control. Enter the credentials once in Settings → Infrastructure (Providers tab); they are stored encrypted on Marathoon's side (or in your Key Vault when configured) and are never displayed again. The guided setup also generates a least-privilege IAM policy scoped to the bucket.

Data flow: files are streamed through the Marathoon API when a runner downloads inputs or uploads outputs; Marathoon does not keep a copy of file contents. Metadata (job records, log lines, artifact names) stays in Marathoon.

Managing machines

Settings → Connected machines lists every machine running the agent — file watchers and enrolled runners — with its live status (online, Docker reachable), the providers whose jobs route to it, the API key it authenticates with and the enrollment tokens still waiting for a machine. Revoking a runner deactivates its keys, removes its provider bindings and deletes its record in one action; the machine is refused on its next poll. Revocation applies to file agents too, and is the only way to invalidate a bound key.

Private image registries

A Connected Runner provider can carry credentials for one private registry (host, username, password or access token). They are stored in the platform secret store like any other provider credential and delivered to the enrolled runner only when it claims a job whose image is hosted on that registry: public images pull anonymously, and the credential never appears in any API response or log. Prefer a read-only token scoped to the images you run.

Cloud VM templates

The connect wizard also produces a cloud-init user-data document and Terraform snippets (AWS EC2, Azure VM) that boot a fresh instance straight into an enrolled runner with no inbound rule. The enrollment token inside them is single-use and expires 15 minutes after generation: apply immediately and never bake the document into a machine image.

Running on Kubernetes

The connect wizard's Kubernetes tab gives you a helm install for the marathoon-runner chart. One runner pod per release; every job becomes its own Kubernetes Job in the namespace, with no Docker socket and no privileged container.

  • Each job pod has an emptyDir workspace: a staging init container waits while the runner streams the inputs in through the exec API, the job container sees /inputs and /outputs, and a small sidecar keeps the pod alive so the runner can pull the outputs out the same way. Job pods mount no service-account token and hold no Marathoon credential.
  • The runner's Role is limited to jobs, pods, pods/exec, pods/log and secrets in its namespace; install it in a dedicated namespace. Private registries use a per-job kubernetes.io/dockerconfigjson pull Secret owned by the Job, so it disappears with it.
  • Enrolled credentials (API key, agent id) live in the Secret <release>-credentials, so a restarted pod keeps its identity without a persistent volume. Set marathoon.resetCredentials=true with a fresh token to re-enrol.
  • jobs.resources sets requests/limits on every job pod (required with a ResourceQuota); jobs.nodeSelector, jobs.workspaceSizeLimit and jobs.helperImage cover node placement, workspace size and air-gapped clusters. Cancellation and timeouts delete the Job; ttlSecondsAfterFinished is the safety net.

Direct runner ↔ storage transfers

When a storage provider has "Runners talk to this storage directly" enabled, a Connected Runner asks the API for a short-lived pre-authorized URL (15 minutes, one blob) and downloads inputs or uploads outputs straight from and to your bucket; the API only records the resulting artifact. Job content then never transits through Marathoon. If the runner cannot reach the bucket, it falls back to the API for that file and says so in the job log. Never available for the hosted default storage.

Zero shared secret

Instead of storing a key, you can grant Marathoon’s own cloud identity access: on AWS, a cross-account role Marathoon assumes with a per-provider External ID you put in the trust policy; on Azure, a Storage Blob Data Contributor assignment for Marathoon’s managed identity on your container. Revoking is a change on your side, with nothing to rotate. The identities to trust are shown in the connect form and come from the operator’s configuration.

How billing changes

ModeComputeBilling
HostedMarathoon runs the containerMetered in credits per run
Your infraYour Docker host runs itFlat platform fee, no per-run compute charge

Marketplace

The marketplace lets creators sell pipelines and lets any tenant install them. A listing wraps a global pipeline with pricing and goes through admin review before it becomes purchasable.

As a buyer

Browse the catalog at /marketplace, open a listing, and install it. Installed pipelines become runnable in your projects; they always run on Marathoon's hosted compute, never on your own runners.

As a creator

Become a creator in Dashboard → Marketplace → Creator, list a published pipeline, submit it for review, and earn on every purchase and every run.

Pricing models

ModelCharge
OneTimePurchaseA one-time unlock price in USD
PerExecutionMarkupExtra credits added on top of compute, on every run
BothOne-time unlock plus a per-run credit markup

Revenue split: the platform keeps 20% and the creator 80% of both one-time sales and per-run markup — 15% / 85% once a creator's net earnings pass $10,000 over 12 months. Stripe fees are on the platform. Creators withdraw via Stripe once their balance reaches $50.

Listing lifecycle

Draft → PendingReview → Approved (or Rejected with notes). Approved listings can be Delisted at any time. Only Approved listings are purchasable and appear in the public catalog.

Agent & Batches

Two ways to run at scale without the UI: submit a batch of jobs at once, or install the desktop agent to turn a watched folder into automatic runs.

Batches

A batch groups many jobs submitted together — ideal for processing hundreds or thousands of files against one pipeline. Track them on the Batches page; each child job has its own status, logs and artifacts.

bash
curl -X POST \
  https://<domain>/api/v1/tenants/<tenantId>/jobs/batch/<pipelineId> \
  -H "X-API-Key: <api-key>" \
  -H "Content-Type: application/json" \
  -d '{
    "inputsList": [
      { "input_file": "s3://bucket/a.csv" },
      { "input_file": "s3://bucket/b.csv" }
    ]
  }'

Desktop agent

Pair the desktop agent from Dashboard → Agent → Pair. One click logs you in, mints a pipeline-scoped API key, and writes the local config. The agent then watches a directory:

  • Drop a single file → a one-shot run starts
  • Drop a sub-folder → one run with all files in it as inputs (batch mode)
  • Live log tail and run details are shown in the agent window

The paired key is scoped to a single pipeline, so the agent can only submit jobs for that pipeline.

Authentication

All API requests require authentication. Two methods are supported.

JWT (User Session)

Obtained via POST /api/v1/auth/login. Pass the token as a Bearer header:

bash
Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...

API Keys

Create API keys from Settings → API Keys or via API. Keys are shown once on creation — copy immediately. Send the key in the X-API-Key header (the Authorization header with the ApiKey scheme also works).

bash
X-API-Key: mk_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

Creating a key via API:

bash
curl -X POST \
  https://<domain>/api/v1/tenants/<tenantId>/api-keys \
  -H "Authorization: Bearer <jwt>" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "CI Pipeline Key",
    "preset": "Operator",
    "expiresAtUtc": "2027-01-01T00:00:00Z"
  }'

Permission presets:

PresetCapabilities
ReadOnlyRead jobs, pipelines, projects. No write access.
OperatorSubmit/cancel/retry jobs, download artifacts, read everything.
AdminFull tenant access: manage API keys, settings, team, billing.
CustomGranular permissions specified in customPermissions array.

API Reference

All endpoints are prefixed with /api/v1/tenants/{tenantId}

Jobs

GET/jobsList jobs (filterable by pipeline, project, status, date)
GET/jobs/{jobId}Get job detail
POST/jobs/run/{pipelineId}Submit a single job
POST/jobs/batch/{pipelineId}Submit a batch of jobs
POST/jobs/{jobId}/cancelCancel a running or queued job
POST/jobs/{jobId}/retryRetry a failed job
DELETE/jobs/{jobId}Delete a completed/failed job
GET/jobs/{jobId}/logsGet structured logs (paginated)
GET/jobs/{jobId}/container-logsGet raw container stdout/stderr
GET/jobs/{jobId}/artifacts/{artifactId}/downloadDownload an output artifact

Submit job body:

json
{
  "inputs": { "key": "value" },
  "parameterOverrides": {
    "cpuUnits": 4,
    "memoryMb": 8192,
    "timeoutSeconds": 1800,
    "environmentVariables": { "DEBUG": "true" }
  }
}

Pipelines

Prefix: /projects/{projectId}/pipelines

GET/projects/{projectId}/pipelinesList pipelines in a project
GET/projects/{projectId}/pipelines/{pipelineId}Get pipeline detail
POST/projects/{projectId}/pipelinesCreate a pipeline
PUT/projects/{projectId}/pipelines/{pipelineId}Update definition (creates new version)
DELETE/projects/{projectId}/pipelines/{pipelineId}Delete a pipeline
POST/projects/{projectId}/pipelines/{pipelineId}/publishPublish (make runnable)
POST/projects/{projectId}/pipelines/{pipelineId}/unpublishUnpublish
GET/projects/{projectId}/pipelines/{pipelineId}/versionsList all versions
GET/projects/{projectId}/pipelines/{pipelineId}/versions/{n}Get a specific version
POST/projects/{projectId}/pipelines/{pipelineId}/versions/{n}/revertRevert to a previous version
GET/projects/{projectId}/pipelines/{pipelineId}/versions/{a}/diff/{b}Diff two versions
POST/projects/{projectId}/pipelines/{pipelineId}/validateValidate a definition without saving
POST/projects/{projectId}/pipelines/{pipelineId}/test-runsSubmit a test run (unpublished ok)
PUT/projects/{projectId}/pipelines/{pipelineId}/parametersUpdate default parameters

API Keys

GET/api-keysList all API keys
GET/api-keys/presetsList available permission presets
GET/api-keys/{apiKeyId}Get a specific key
POST/api-keysCreate an API key (key shown once)
PUT/api-keys/{apiKeyId}Update name/description/permissions
POST/api-keys/{apiKeyId}/revokeDeactivate without deleting
DELETE/api-keys/{apiKeyId}Permanently delete

Error Format

All error responses follow this shape:

json
{
  "error": "Human-readable error message",
  "code": "Domain.ErrorCode"
}
HTTPMeaning
400Validation error or bad request
401Missing or invalid authentication
402Insufficient credits
403Insufficient permissions
404Resource not found
409Conflict (e.g., job already cancelled)

Webhooks

Configure webhooks from Settings → Webhooks to receive a signed HTTP POST when a job or batch finishes. Deduplicate on the payload id: retries and replays reuse it.

Payload

json
{
  "id": "6f1c2a4e-8a43-4a0e-9a8e-2d5c3c6b1f00",
  "type": "job.completed",
  "createdAt": "2026-06-17T12:34:56Z",
  "organizationId": "...",
  "data": {
    "jobId": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
    "sequence": 1,
    "executionId": "...",
    "pipelineId": "...",
    "status": "Succeeded",
    "exitCode": 0,
    "startedAtUtc": "2026-06-17T12:30:01Z",
    "endedAtUtc": "2026-06-17T12:34:56Z",
    "durationSeconds": 295
  }
}

Signature

Every request carries X-Marathoon-Signature: t=<unix seconds>,v1=<hex>, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (shown once, at creation or rotation). Recompute it over the raw body, accept the request if any v1 matches (for 24 hours after a rotation there is one per secret), and reject requests whose t is more than 5 minutes old. X-Marathoon-Event carries the event type and X-Marathoon-Delivery the delivery id. Deliveries can arrive out of order: reorder a job's events by data.sequence. Use the Test button to send a signed webhook.ping.

javascript
import { createHmac, timingSafeEqual } from 'node:crypto';

function verify(rawBody, header, secret) {
  const parts = header.split(',').map((p) => p.split('='));
  const t = parts.find(([k]) => k === 't')?.[1];
  if (!t || Math.abs(Date.now() / 1000 - Number(t)) > 300) return false;
  const expected = createHmac('sha256', secret).update(`${t}.${rawBody}`).digest('hex');
  // After a rotation there is one v1 per secret for 24 h: any match is enough.
  return parts.some(([k, v]) => k === 'v1' && v.length === expected.length
    && timingSafeEqual(Buffer.from(v), Buffer.from(expected)));
}

Retries

Answer with a 2xx within 10 seconds. Timeouts, network errors, 408, 429 and 5xx are retried after 1 min, 5 min, 30 min, 2 h, 6 h, 12 h and 24 h; other 4xx and redirects are not retried, and 410 Gone disables the webhook. Endpoints must use HTTPS.

Events

job.startedJob's container started running
job.completedJob finished successfully
job.failedJob failed (non-zero exit, error, or could not start)
job.timed_outJob exceeded its timeout
job.cancelledJob cancelled
batch.completedAll jobs of a batch finished
pipeline.publishedPipeline published
pipeline.updatedNew pipeline version saved
member.addedUser joined the organization
member.removedMember removed from the organization

Something missing?

Open a support request from the Help page.

Contact Support