Documentation
Complete reference for building and running pipelines on Marathoon.
Overview
Marathoon is a multi-tenant pipeline orchestration platform. You define pipelines as Docker-based workflows, submit jobs via the UI or API, and get structured results including logs, artifacts, and visual reports.
Projects & Pipelines
Organize work into projects; define reusable pipelines.
Jobs & Batches
Run pipelines on demand or submit hundreds at once.
Visual Reports
Render charts, tables, and metrics directly on job pages.
Base API URL
https://<your-domain>/api/v1/tenants/<tenantId>Getting Started
1. Create a Project
A Project is the top-level container for pipelines. Go to Projects → New Project in the dashboard.
Projects control access: team members are granted access per-project with roles (Owner, Maintainer, Analyst, Viewer).
2. Define a Pipeline
A Pipeline defines a Docker image, its inputs/outputs, and the command to run. Use the visual pipeline editor or POST via API.
{
"name": "My Pipeline",
"description": "Processes input files and produces a report",
"definition": {
"imageUri": "myregistry.io/my-pipeline",
"imageTag": "latest",
"command": ["python", "main.py"],
"inputs": [
{ "name": "input_file", "type": "file", "required": true },
{ "name": "threshold", "type": "string", "required": false, "defaultValue": "0.5" }
],
"outputs": [
{ "name": "results", "type": "file", "path": "/outputs/results.csv" },
{ "name": "report", "type": "file", "path": "/outputs/report.json" }
]
},
"defaultParameters": {
"cpuUnits": 2,
"memoryMb": 2048,
"timeoutSeconds": 3600,
"environmentVariables": {},
"secretRefs": []
}
}After creation, Publish the pipeline to make it runnable.
3. Submit a Job
Submit a job via the UI (Run button on the pipeline page) or via API:
curl -X POST \
https://<domain>/api/v1/tenants/<tenantId>/jobs/run/<pipelineId> \
-H "X-API-Key: <api-key>" \
-H "Content-Type: application/json" \
-d '{
"inputs": {
"input_file": "s3://my-bucket/data.csv",
"threshold": "0.8"
}
}'4. Monitor & Retrieve Results
Track job status in real-time via the Jobs page (SignalR live updates). Each job shows:
- Status timeline (Queued → Running → Succeeded / Failed)
- Structured logs (searchable, paginated)
- Container stdout/stderr
- Output artifacts (downloadable)
- Visual report (if report.json produced)
Core Concepts
Projects
Logical grouping for pipelines. Members access projects through roles. A project can be archived, which blocks new job submissions while preserving history.
Pipelines
Versioned Docker-based workflow definitions. Each publish increments the version number. Old versions are preserved and jobs reference the version they ran against. Pipelines must be published before users can submit jobs.
Jobs & Batches
A Job is a single pipeline execution with specific inputs. A Batch groups multiple jobs submitted together (e.g., processing 1000 files). Jobs transition through: Queued → Running → Succeeded | Failed | Cancelled.
Credits
Credits are the billing unit (100 credits = $1). Each job submission deducts credits based on plan limits. Credit balance is visible in the sidebar. Plans: Free (200), Research (1,500), Pro (7,500), Enterprise (custom). Insufficient credits returns 402 Payment Required.
Permissions
Two-layer authorization: JWT (user session) and API keys (per-key permission set). Permissions are scoped to the tenant. Key presets: ReadOnly, Operator, Admin, or custom granular permissions.
Pipeline Definition
The pipeline definition is the core schema stored with each pipeline. It is versioned on every update.
Top-Level Fields
| Field | Type | Description |
|---|---|---|
| imageUri | string (required) | Docker image URI (without tag) |
| imageTag | string | Docker image tag, defaults to latest |
| command | string[] | Override the container entrypoint command |
| entryPoint | string | Alternative entrypoint override |
| inputs | PipelineInput[] | Declared input slots |
| outputs | PipelineOutput[] | Declared output slots |
| tasks | PipelineTask[] | Graph nodes (visual editor) |
| connections | PipelineConnection[] | Graph edges (visual editor) |
Inputs & Outputs
{
"inputs": [
{
"name": "input_file",
"type": "file",
"required": true,
"description": "Input CSV to process"
},
{
"name": "mode",
"type": "string",
"required": false,
"defaultValue": "fast"
}
],
"outputs": [
{
"name": "results",
"type": "file",
"path": "/outputs/results.csv",
"description": "Processed output"
}
]
}Input values are passed as environment variables to the container: INPUT_FILE, MODE, etc. (uppercased).
Parameters (Resources)
{
"cpuUnits": 2,
"memoryMb": 4096,
"timeoutSeconds": 7200,
"environmentVariables": {
"LOG_LEVEL": "info"
},
"secretRefs": []
}cpuUnits maps to CPU cores. Parameters can be overridden per job submission. secretRefs is reserved for managed pipeline secrets, which are not available yet: never put secret values in parameters or environment variables.
Visual Reports
Any pipeline can produce a rich visual report. Write a file named report.json to the output directory. Marathoon auto-detects it and renders it on the job page.
Top-Level Schema
{
"id": "report-001",
"jobId": "<jobId>",
"title": "Analysis Results",
"description": "Optional subtitle shown below the title",
"generatedAtUtc": "2026-06-17T12:00:00Z",
"blocks": [ ... ],
"tabs": [
{ "label": "Overview", "blocks": [ ... ] },
{ "label": "Details", "blocks": [ ... ] }
]
}Use tabs for multi-section reports. When present, the renderer shows a tab bar instead of a flat block list. blocks at root level is used when no tabs.
Block Types
{ "id":"b1","order":1,"type":"text","data":{ "content":"Hello world","style":"heading1" } }
{ "id":"b2","order":2,"type":"markdown","data":{ "markdown":"**Bold** and _italic_" } }text styles: paragraph, heading1, heading2, heading3, quote
{
"id": "b3", "order": 3, "type": "stats",
"data": {
"stats": [
{ "label": "Accuracy", "value": "94.2", "unit": "%", "change": "+2.1%", "changeDirection": "up" },
{ "label": "Samples", "value": "10000" }
]
}
}{
"id": "b4", "order": 4, "type": "table",
"data": {
"columns": [
{ "key": "name", "header": "Name", "type": "text", "sortable": true },
{ "key": "score", "header": "Score", "type": "number", "sortable": true },
{ "key": "pass", "header": "Pass", "type": "boolean" }
],
"rows": [
{ "name": "Sample A", "score": 0.92, "pass": true },
{ "name": "Sample B", "score": 0.51, "pass": false }
],
"sortable": true,
"filterable": true,
"paginated": true,
"pageSize": 20
}
}Column types: text number date boolean badge link progress
{
"id": "b5", "order": 5, "type": "chart", "title": "Loss Curve",
"data": {
"chartType": "line",
"series": [
{
"name": "Train",
"data": [{"x":1,"y":0.9},{"x":2,"y":0.7},{"x":3,"y":0.5}]
},
{
"name": "Val",
"data": [{"x":1,"y":0.95},{"x":2,"y":0.75},{"x":3,"y":0.6}]
}
],
"options": { "xAxisLabel": "Epoch", "yAxisLabel": "Loss", "showLegend": true }
}
}chartType: line bar area pie donut scatter radar
{
"id": "b6", "order": 6, "type": "vega",
"data": {
"spec": {
"$schema": "https://vega.github.io/schema/vega-lite/v5.json",
"mark": "bar",
"data": { "values": [{"x":"A","y":3},{"x":"B","y":7}] },
"encoding": {
"x": { "field": "x", "type": "ordinal" },
"y": { "field": "y", "type": "quantitative" }
}
}
}
}Full Vega-Lite (or Vega) spec. Use for advanced / custom visualizations.
{
"id": "b7", "order": 7, "type": "image",
"data": { "url": "/output/confusion_matrix.png", "altText": "Confusion matrix" }
}URL references another output artifact by its path. Marathoon resolves it from blob storage.
{ "id":"b8","order":8,"type":"alert", "data":{"message":"Dataset has nulls","severity":"warning"} }
{ "id":"b9","order":9,"type":"progress","data":{"value":750,"max":1000,"unit":"samples"} }alert severities: info success warning error
Example Pipelines
Four pipelines ready to import from Pipelines → View examples, or via Import JSON. Each pairs with the file agent: drop a sub-folder → a job starts automatically.
Comment Marathoon échange des fichiers avec votre conteneur
Chaque conteneur reçoit ses données via deux répertoires montés :
| Répertoire | Rôle |
|---|---|
/inputs | Fichiers uploadés par l'agent ou soumis manuellement |
/outputs | Tout fichier déposé ici devient un artefact téléchargeable |
Variables d'environnement. Les inputs scalaires (type string) sont injectés en majuscules :
mode: "fast" → MODE=fast dans l'environnement du conteneur.
Rapport visuel. Si votre script écrit /outputs/report.json au format Marathoon, la page du job l'affiche automatiquement sous forme de graphiques, tableaux et métriques.
Sorties déclarées. Le champ path d'une sortie (ex. /outputs/report.json) indique à Marathoon l'artefact principal à mettre en avant. Les autres fichiers dans /outputs sont quand même capturés.
1. Quick Check
Pipeline minimal — aucune entrée, un job rapide. Commencez ici.
Pipeline le plus simple possible : il démarre, simule une courte analyse et s'arrête avec succès.
Objectif pédagogique : vérifier que votre agent est bien appairé et que déposer un dossier déclenche bien un job de bout en bout, avant de passer aux exemples avec fichiers.
Concept clé : même sans entrées déclarées, le conteneur peut écrire dans /outputs — tout fichier y déposé devient un artefact téléchargeable sur la page du job.
Avec l'agent fichier : déposez n'importe quel sous-dossier dans le répertoire surveillé. Chaque sous-dossier = un job. Le contenu est ignoré.
ADockerfile — how the image is built
# Image de base légère — Alpine Linux avec bash, curl, jq
FROM alpine:3.19
RUN apk add --no-cache bash curl jq
WORKDIR /app
COPY analyze.sh /app/analyze.sh
RUN chmod +x /app/analyze.sh
# Paramètres tunable via environmentVariables à la soumission du job
ENV ANALYSIS_DURATION=30 \
OUTPUT_DIR=/outputs
ENTRYPOINT ["/app/analyze.sh"]BEntrypoint script: analyze.sh
#!/bin/bash
set -e
# ── Configuration ──────────────────────────────────────────────────────────
# Marathoon injecte les paramètres de la pipeline en variables d'environnement
# (uppercased). Ici on peut les surcharger via parameterOverrides à la soumission.
DURATION=${ANALYSIS_DURATION:-30}
OUTPUT_DIR=${OUTPUT_DIR:-/outputs}
mkdir -p "$OUTPUT_DIR"
echo "[INFO] Démarrage — durée simulée : ${DURATION}s"
# ── Travail simulé ──────────────────────────────────────────────────────────
for step in "Initialisation" "Chargement" "Traitement" "Agrégation"; do
echo "[INFO] $step..."
sleep $((DURATION / 4))
done
# ── Sorties ─────────────────────────────────────────────────────────────────
# Tout fichier écrit dans /outputs devient un artefact téléchargeable.
cat > "$OUTPUT_DIR/summary.json" <<JSON
{
"status": "completed",
"duration_seconds": $DURATION,
"records_processed": 985432,
"records_skipped": 14568
}
JSON
cat > "$OUTPUT_DIR/metrics.csv" <<CSV
metric,value,unit
records_processed,985432,count
records_skipped,14568,count
execution_time,$DURATION,seconds
CSV
echo "[INFO] Terminé. Résultats dans $OUTPUT_DIR/"
CInputs / Outputs
Entrées : aucune. Le conteneur ignore le contenu du dossier déposé — il a juste besoin d'un déclencheur.
Sorties déclarées :
result→/outputs/result.json(déclaré dans la pipeline)
Sorties produites réellement :
/outputs/summary.json— résumé JSON de l'exécution/outputs/metrics.csv— métriques au format CSV
Ces deux fichiers apparaissent comme artefacts même s'ils ne sont pas déclarés explicitement dans la définition de la pipeline.
Folder to drop into the watched directory
watch/
└── run-001/
└── anything.txtPipeline definition (importable JSON)
{
"name": "Quick Check",
"description": "Minimal sample pipeline — fast run, no inputs.",
"definition": {
"imageUri": "marathoon-analyzer-simple",
"imageTag": "latest",
"inputs": [],
"outputs": [
{
"name": "result",
"type": "file",
"path": "/outputs/result.json",
"description": "Run result"
}
],
"entryPoint": null,
"command": [],
"tasks": [],
"connections": []
},
"defaultParameters": {
"cpuUnits": 1,
"memoryMb": 256,
"timeoutSeconds": 120,
"environmentVariables": {},
"secretRefs": []
}
}2. File Analyzer
Lit tous les fichiers d'un dossier et produit un rapport d'analyse.
Traite l'ensemble des fichiers déposés dans /inputs et génère un rapport d'analyse par fichier, plus un rapport agrégé.
Objectif pédagogique : illustrer le binding mount-all — l'agent uploade tous les fichiers d'un sous-dossier dans /inputs, sans nommage requis. Le script parcourt le répertoire entier.
Concepts clés :
/inputsest peuplé automatiquement par Marathoon avec les fichiers uploadés- Le script peut traiter n'importe quel type de fichier (texte, JSON, CSV, binaire)
- Chaque fichier produit son propre rapport JSON dans
/outputs/reports/
Avec l'agent fichier : déposez un dossier contenant des fichiers à analyser. Tous les fichiers (n'importe quel type) sont envoyés ensemble dans un seul job.
ADockerfile — how the image is built
FROM alpine:3.19
# python3 pour le script de traitement par type de fichier
RUN apk add --no-cache bash curl jq python3 py3-pip
WORKDIR /app
# Marathoon monte /inputs et /outputs — on les crée pour la clarté
RUN mkdir -p /inputs /outputs
COPY analyze.sh process_file.py /app/
RUN chmod +x /app/analyze.sh
ENV INPUT_DIR=/inputs \
OUTPUT_DIR=/outputs
# VOLUME déclare les points de montage (optionnel mais explicite)
VOLUME ["/inputs", "/outputs"]
ENTRYPOINT ["/app/analyze.sh"]BEntrypoint script: analyze.sh
#!/bin/bash
set -e
# ── Configuration ──────────────────────────────────────────────────────────
INPUT_DIR=${INPUT_DIR:-/inputs}
OUTPUT_DIR=${OUTPUT_DIR:-/outputs}
mkdir -p "$OUTPUT_DIR/reports"
# ── Inventaire des entrées ──────────────────────────────────────────────────
echo "[INFO] Scan de $INPUT_DIR..."
FILE_COUNT=$(find "$INPUT_DIR" -type f | wc -l | tr -d ' ')
if [ "$FILE_COUNT" -eq 0 ]; then
echo "[WARN] Aucun fichier dans $INPUT_DIR"
exit 0
fi
echo "[INFO] $FILE_COUNT fichier(s) trouvé(s)"
# ── Traitement fichier par fichier ──────────────────────────────────────────
# process_file.py détecte le type (txt, json, csv, binaire) et
# génère un rapport JSON d'analyse pour chaque fichier.
find "$INPUT_DIR" -type f | while read -r file; do
name=$(basename "$file")
echo "[INFO] Traitement : $name"
out="$OUTPUT_DIR/reports/${name%.*}_analysis.json"
python3 /app/process_file.py "$file" "$out"
done
# ── Rapport agrégé ──────────────────────────────────────────────────────────
DONE=$(find "$OUTPUT_DIR/reports" -name "*.json" | wc -l | tr -d ' ')
cat > "$OUTPUT_DIR/aggregate_report.json" <<JSON
{
"total_files_processed": $DONE,
"input_directory": "$INPUT_DIR",
"output_directory": "$OUTPUT_DIR",
"status": "completed"
}
JSON
echo "[INFO] Rapport agrégé écrit. $DONE fichier(s) traité(s)."
CEntrypoint script: process_file.py
#!/usr/bin/env python3
"""
Analyse un fichier selon son type (texte, JSON, CSV, binaire)
et écrit un rapport JSON dans le répertoire de sortie.
Usage: python3 process_file.py <input_path> <output_path>
"""
import json
import sys
import os
import hashlib
from datetime import datetime
from pathlib import Path
def get_file_info(filepath: str) -> dict:
path = Path(filepath)
stat = path.stat()
with open(filepath, 'rb') as f:
file_hash = hashlib.sha256(f.read()).hexdigest()
return {
"filename": path.name,
"size_bytes": stat.st_size,
"extension": path.suffix.lower(),
"sha256": file_hash[:16] + "...",
}
def analyze_text(filepath: str) -> dict:
with open(filepath, 'r', encoding='utf-8', errors='ignore') as f:
content = f.read()
lines = content.split('\n')
return {
"type": "text",
"lines": len(lines),
"words": len(content.split()),
"characters": len(content),
"non_empty_lines": len([l for l in lines if l.strip()]),
}
def analyze_json(filepath: str) -> dict:
with open(filepath, 'r', encoding='utf-8') as f:
data = json.load(f)
def depth(obj, d=0):
if isinstance(obj, dict):
return max((depth(v, d + 1) for v in obj.values()), default=d)
if isinstance(obj, list):
return max((depth(v, d + 1) for v in obj), default=d)
return d
return {
"type": "json",
"root_type": type(data).__name__,
"keys": len(data) if isinstance(data, dict) else None,
"items": len(data) if isinstance(data, list) else None,
"max_depth": depth(data),
}
def analyze_csv(filepath: str) -> dict:
with open(filepath, 'r', encoding='utf-8', errors='ignore') as f:
lines = f.readlines()
if not lines:
return {"type": "csv", "rows": 0, "columns": 0}
first = lines[0]
delimiter = ',' if ',' in first else ('\t' if '\t' in first else ';')
return {
"type": "csv",
"rows": len(lines),
"columns": len(first.split(delimiter)),
"delimiter": repr(delimiter),
}
def analyze_binary(filepath: str) -> dict:
with open(filepath, 'rb') as f:
header = f.read(4)
return {
"type": "binary",
"magic": header.hex(),
}
def process(filepath: str) -> dict:
info = get_file_info(filepath)
ext = info["extension"]
if ext in ('.txt', '.log', '.md', '.py', '.js', '.sh'):
analysis = analyze_text(filepath)
elif ext == '.json':
try:
analysis = analyze_json(filepath)
except json.JSONDecodeError:
analysis = {"type": "json", "error": "Invalid JSON"}
elif ext in ('.csv', '.tsv'):
analysis = analyze_csv(filepath)
else:
analysis = analyze_binary(filepath)
return {
"file_info": info,
"analysis": analysis,
"processed_at": datetime.utcnow().isoformat() + "Z",
}
def main():
if len(sys.argv) < 2:
print(json.dumps({"error": "Usage: process_file.py <input> [output]"}))
sys.exit(1)
filepath = sys.argv[1]
output_path = sys.argv[2] if len(sys.argv) > 2 else None
if not os.path.exists(filepath):
print(json.dumps({"error": f"File not found: {filepath}"}))
sys.exit(1)
result = process(filepath)
if output_path:
os.makedirs(os.path.dirname(output_path), exist_ok=True)
with open(output_path, 'w') as f:
json.dump(result, f, indent=2)
print(f"[INFO] Rapport écrit : {output_path}")
else:
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
DInputs / Outputs
Entrées : tous les fichiers présents dans /inputs (binding mount-all).
Aucun nommage requis — le script parcourt le répertoire entier.
Types supportés : .txt, .log, .md, .py, .js, .sh, .json, .csv, .tsv, binaires.
Sorties produites :
/outputs/reports/<nom>_analysis.json— un rapport par fichier/outputs/aggregate_report.json— résumé global
Sortie déclarée : report → /outputs/analysis-report.json
Folder to drop into the watched directory
watch/
└── batch-2024-06/
├── document-a.txt
├── document-b.txt
└── data.jsonPipeline definition (importable JSON)
{
"name": "File Analyzer",
"description": "Reads all files in /inputs and emits an analysis report.",
"definition": {
"imageUri": "marathoon-analyzer-file",
"imageTag": "1.0",
"inputs": [],
"outputs": [
{
"name": "report",
"type": "file",
"path": "/outputs/analysis-report.json",
"description": "Analysis report"
}
],
"entryPoint": null,
"command": [],
"tasks": [],
"connections": []
},
"defaultParameters": {
"cpuUnits": 2,
"memoryMb": 512,
"timeoutSeconds": 300,
"environmentVariables": {
"ANALYSIS_DURATION": "5",
"LOG_INTERVAL": "1"
},
"secretRefs": []
}
}3. CSV → JSON
Convertit chaque fichier CSV en JSON et affiche un rapport de conversion.
Pour chaque *.csv dans /inputs, écrit un <nom>.json correspondant (tableau d'objets ligne) dans /outputs, plus un report.json rendu automatiquement sur la page du job.
Objectif pédagogique : montrer deux mécanismes de binding en parallèle :
- Name-match binding : un input nommé
datalie le fichierdata.csvautomatiquement - Mount-all : tout autre
*.csvest quand même traité via le scan du répertoire
Concept clé : le report.json au format Marathoon est détecté automatiquement par la plateforme et rendu sur la page du job avec stats + tableau.
Avec l'agent fichier : déposez un dossier contenant des fichiers CSV.
ADockerfile — how the image is built
# Python 3.12 alpine — stdlib csv/json suffisent, aucune dépendance externe
FROM python:3.12-alpine
WORKDIR /app
RUN mkdir -p /inputs /outputs
COPY convert.py /app/convert.py
ENV INPUT_DIR=/inputs \
OUTPUT_DIR=/outputs
VOLUME ["/inputs", "/outputs"]
ENTRYPOINT ["python", "/app/convert.py"]BEntrypoint script: convert.py
#!/usr/bin/env python3
"""
CSV → JSON converter sample.
Lit tous les *.csv sous INPUT_DIR (récursivement), écrit un <nom>.json
(tableau d'objets ligne) pour chacun dans OUTPUT_DIR, puis génère un
report.json au format Marathoon — rendu automatiquement sur la page du job.
"""
import csv
import json
import os
import sys
from datetime import datetime, timezone
INPUT_DIR = os.environ.get("INPUT_DIR", "/inputs")
OUTPUT_DIR = os.environ.get("OUTPUT_DIR", "/outputs")
def log(msg: str) -> None:
ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%S.%fZ")
print(f"[INFO] {ts} {msg}", flush=True)
def find_csv_files(root: str):
"""Parcourt récursivement root et yield chaque *.csv trouvé."""
for dirpath, _dirs, files in os.walk(root):
for name in files:
if name.lower().endswith(".csv"):
yield os.path.join(dirpath, name)
def convert_file(path: str) -> dict:
"""Convertit un CSV en liste de dicts et l'écrit dans OUTPUT_DIR."""
rel = os.path.relpath(path, INPUT_DIR)
try:
with open(path, newline="", encoding="utf-8-sig") as fh:
rows = list(csv.DictReader(fh))
except Exception as exc:
log(f"ERREUR lecture {rel}: {exc}")
return {"file": rel, "rows": 0, "status": f"error: {exc}"}
out_name = os.path.splitext(os.path.basename(path))[0] + ".json"
out_path = os.path.join(OUTPUT_DIR, out_name)
with open(out_path, "w", encoding="utf-8") as out:
json.dump(rows, out, indent=2, ensure_ascii=False)
log(f"Converti {rel} → {out_name} ({len(rows)} lignes)")
return {"file": rel, "rows": len(rows), "status": "ok"}
def write_report(summary: list) -> None:
"""
Écrit /outputs/report.json au format Marathoon.
La plateforme détecte ce fichier et le rend sur la page du job
sous forme de stats + tableau.
"""
total_rows = sum(s["rows"] for s in summary)
report = {
"title": "CSV → JSON conversion",
"description": "Un fichier JSON écrit par CSV converti.",
"generatedAtUtc": datetime.now(timezone.utc).isoformat(),
"blocks": [
{
"type": "stats",
"stats": [
{"label": "Fichiers convertis", "value": str(len(summary))},
{"label": "Lignes totales", "value": str(total_rows)},
],
},
{
"type": "table",
"title": "Fichiers convertis",
"columns": [
{"key": "file", "header": "Fichier"},
{"key": "rows", "header": "Lignes"},
{"key": "status", "header": "Statut"},
],
"rows": summary,
},
],
}
out = os.path.join(OUTPUT_DIR, "report.json")
with open(out, "w", encoding="utf-8") as fh:
json.dump(report, fh, indent=2)
log(f"report.json écrit ({os.path.getsize(out)} octets)")
def main() -> int:
os.makedirs(OUTPUT_DIR, exist_ok=True)
log(f"Scan de {INPUT_DIR} pour les fichiers CSV")
csv_files = sorted(find_csv_files(INPUT_DIR))
if not csv_files:
log("Aucun CSV trouvé — rien à convertir")
write_report([])
return 0
summary = [convert_file(path) for path in csv_files]
write_report(summary)
log(f"Terminé : {len(summary)} fichier(s) traité(s)")
return 0
if __name__ == "__main__":
sys.exit(main())
CInputs / Outputs
Entrées : tous les *.csv dans /inputs (récursif).
L'input nommé data permet de binder data.csv explicitement, mais tous les autres CSV sont quand même traités.
Sorties produites :
/outputs/<nom>.json— un tableau JSON de lignes par CSV/outputs/report.json— rapport de conversion (stats + tableau), rendu automatiquement sur la page du job
Sortie déclarée : report → /outputs/report.json
Folder to drop into the watched directory
watch/
└── export-001/
├── data.csv
└── more.csvPipeline definition (importable JSON)
{
"name": "CSV to JSON",
"description": "Converts CSV files to JSON with a conversion report.",
"definition": {
"imageUri": "marathoon-csv-json",
"imageTag": "1.0",
"inputs": [
{
"name": "data",
"type": "file",
"required": false,
"defaultValue": null,
"description": "A CSV file to convert"
}
],
"outputs": [
{
"name": "report",
"type": "file",
"path": "/outputs/report.json",
"description": "Conversion report"
}
],
"entryPoint": null,
"command": [],
"tasks": [],
"connections": []
},
"defaultParameters": {
"cpuUnits": 1,
"memoryMb": 256,
"timeoutSeconds": 120,
"environmentVariables": {},
"secretRefs": []
}
}4. Visual Report
Émet un report.json qui se rend en graphiques et tableaux sur la page du job.
Génère un report.json couvrant tous les types de blocs Marathoon : stats, alerte, markdown, graphiques (bar, donut), Vega-Lite, tableau.
Objectif pédagogique : servir de référence complète pour le schéma de rapport visuel. Ouvrez le job terminé pour voir le rendu, puis copiez les blocs qui vous intéressent.
Types de blocs disponibles :
| Type | Description |
|---|---|
stats | Métriques avec valeur, unité, variation |
alert | Bandeau info/success/warning/error |
markdown | Texte formaté, titres, listes |
chart | Graphiques (line, bar, area, pie, donut, scatter, radar) |
vega | Spec Vega-Lite complète pour visualisations avancées |
table | Tableau trié/filtré/paginé |
image | Image depuis un artefact de sortie |
progress | Barre de progression |
Avec l'agent fichier : déposez n'importe quel dossier pour déclencher un run — le rapport est généré par la pipeline elle-même.
ADockerfile — how the image is built
# BusyBox suffit — le script est en shell pur FROM busybox:latest COPY report.sh /report.sh RUN chmod +x /report.sh ENTRYPOINT ["/report.sh"]
BEntrypoint script: report.sh
#!/bin/sh
set -e
# Aucune entrée requise — ce pipeline génère ses propres données.
mkdir -p /outputs
# ── Génération du rapport visuel ────────────────────────────────────────────
# Marathoon détecte report.json automatiquement et le rend sur la page du job.
# Utilisez des tabs pour organiser plusieurs sections.
cat > /outputs/report.json <<'JSON'
{
"title": "Sample Analysis Report",
"description": "Généré par le pipeline Visual Report pour illustrer le schéma de rapport.",
"tabs": [
{
"label": "Overview",
"blocks": [
{
"type": "stats",
"title": "Métriques clés",
"stats": [
{ "label": "Lignes traitées", "value": "12,840" },
{ "label": "Erreurs", "value": "3", "changeDirection": "down", "change": "-40%" },
{ "label": "Durée", "value": "8.2", "unit": "s" },
{ "label": "Précision", "value": "98.6", "unit": "%", "changeDirection": "up", "change": "+1.2%" }
]
},
{
"type": "alert",
"severity": "success",
"message": "Toutes les validations sont passées."
},
{
"type": "markdown",
"markdown": "## Résumé\n\nLa pipeline a traité le jeu de données avec succès.\n\n- Deux valeurs aberrantes supprimées\n- Tendance **positive** trimestre sur trimestre"
}
]
},
{
"label": "Graphiques",
"blocks": [
{
"type": "chart",
"title": "Débit mensuel",
"chartType": "bar",
"series": [
{ "name": "Jobs", "data": [{"x":0,"y":120},{"x":1,"y":180},{"x":2,"y":150},{"x":3,"y":240}] }
],
"options": { "xAxisLabel": "Mois", "yAxisLabel": "Jobs", "showGrid": true, "showLegend": true }
},
{
"type": "chart",
"title": "Succès vs Échec",
"chartType": "donut",
"series": [
{ "name": "Résultat", "data": [{"label":"Succès","x":0,"y":92},{"label":"Échec","x":1,"y":8}] }
],
"options": { "showLegend": true }
},
{
"type": "vega",
"title": "Heatmap de corrélation (Vega-Lite)",
"spec": {
"$schema": "https://vega.github.io/schema/vega-lite/v5.json",
"data": { "values": [
{"x":"A","y":"A","v":1.0},{"x":"A","y":"B","v":0.3},{"x":"A","y":"C","v":0.6},
{"x":"B","y":"A","v":0.3},{"x":"B","y":"B","v":1.0},{"x":"B","y":"C","v":0.1},
{"x":"C","y":"A","v":0.6},{"x":"C","y":"B","v":0.1},{"x":"C","y":"C","v":1.0}
]},
"mark": "rect",
"height": 200,
"encoding": {
"x": {"field":"x","type":"nominal"},
"y": {"field":"y","type":"nominal"},
"color": {"field":"v","type":"quantitative","scale":{"scheme":"viridis"}}
}
}
}
]
},
{
"label": "Données",
"blocks": [
{
"type": "table",
"title": "Top résultats",
"sortable": true,
"filterable": true,
"columns": [
{ "key": "name", "header": "Nom", "type": "text" },
{ "key": "score", "header": "Score", "type": "number" },
{ "key": "passed", "header": "Validé", "type": "boolean" }
],
"rows": [
{ "name": "alpha", "score": 0.98, "passed": true },
{ "name": "beta", "score": 0.71, "passed": true },
{ "name": "gamma", "score": 0.42, "passed": false }
]
}
]
}
]
}
JSON
echo "report.json écrit ($(wc -c < /outputs/report.json) octets)"
CInputs / Outputs
Entrées : aucune — le dossier déposé sert uniquement de déclencheur.
Sorties produites :
/outputs/report.json— rapport multi-tabs au format Marathoon, rendu automatiquement sur la page du job
Sortie déclarée : report → /outputs/report.json
Pour votre propre pipeline : générez vos graphiques avec matplotlib/seaborn/Vega-Lite, sérialisez le résultat dans le schéma Marathoon, et écrivez-le dans
/outputs/report.json.
Folder to drop into the watched directory
watch/
└── report-run-1/
└── trigger.txtPipeline definition (importable JSON)
{
"name": "Visual Report",
"description": "Emits a report.json rendered as charts/tables on the job page.",
"definition": {
"imageUri": "marathoon-sample-report",
"imageTag": "latest",
"inputs": [],
"outputs": [
{
"name": "report",
"type": "file",
"path": "/outputs/report.json",
"description": "Visual report"
}
],
"entryPoint": null,
"command": [],
"tasks": [],
"connections": []
},
"defaultParameters": {
"cpuUnits": 1,
"memoryMb": 256,
"timeoutSeconds": 120,
"environmentVariables": {},
"secretRefs": []
}
}Bring Your Own Infrastructure
By default, jobs run on Marathoon's hosted compute and outputs are stored in hosted blob storage. You can instead route a project's jobs to your own Docker host and keep files in your own object storage. When compute is yours, there is no per-run credit charge — only a flat platform fee for orchestration, queueing and the UI. Marketplace pipelines are the exception: they always run on Marathoon's hosted compute.
Connect a runner
Open Settings → Infrastructure → Connect infrastructure → Runner. The guided setup creates a Connected Runner provider, mints a one-time enrollment token (15 minutes, single use) and shows the exact command to run on your machine: a Docker one-liner, a Linux package, or the Windows installer. The machine enrols itself, sends a heartbeat, and the wizard ends with a test job executed end to end on your hardware.
Each project resolves a compute target and a storage target. Override either per project; unset falls back to the tenant default, which falls back to hosted.
How the Connected Runner works
The runner is a small agent that runs next to a Docker engine you own — or inside a Kubernetes cluster. It is pull-only: it polls Marathoon for jobs assigned to it, runs each job in a container (on your Docker, or as a Kubernetes Job), stages inputs and uploads the outputs. Marathoon never connects to your machine.
- Enrolment exchanges the one-time token for an API key limited to the runner endpoints (it cannot submit jobs or read pipelines) and bound to the machine's agent id.
- Every job is leased to one runner at a time; a runner that stops heartbeating loses the lease and the job is re-queued.
- Inputs and outputs are streamed through the Marathoon API to and from your storage provider. Container logs are relayed as they are produced.
- Revoke the runner at any time by deleting its API key or the agent: it stops receiving work on its next poll.
Network prerequisites
The runner needs outbound HTTPS only. Nothing has to be opened on your side.
| Direction | Target | Purpose |
|---|---|---|
| Outbound | Marathoon API (443) | Enrolment, heartbeats, job claims, input download, output upload |
| Outbound | Your image registries (443) | Pulling the job images your pipelines reference |
| Inbound | none | No port, no public IP, no VPN required |
Use your own storage
Point a project at your own S3-compatible bucket or Azure Blob container so job inputs and outputs are kept in storage you control. Enter the credentials once in Settings → Infrastructure (Providers tab); they are stored encrypted on Marathoon's side (or in your Key Vault when configured) and are never displayed again. The guided setup also generates a least-privilege IAM policy scoped to the bucket.
Data flow: files are streamed through the Marathoon API when a runner downloads inputs or uploads outputs; Marathoon does not keep a copy of file contents. Metadata (job records, log lines, artifact names) stays in Marathoon.
Managing machines
Settings → Connected machines lists every machine running the agent — file watchers and enrolled runners — with its live status (online, Docker reachable), the providers whose jobs route to it, the API key it authenticates with and the enrollment tokens still waiting for a machine. Revoking a runner deactivates its keys, removes its provider bindings and deletes its record in one action; the machine is refused on its next poll. Revocation applies to file agents too, and is the only way to invalidate a bound key.
Private image registries
A Connected Runner provider can carry credentials for one private registry (host, username, password or access token). They are stored in the platform secret store like any other provider credential and delivered to the enrolled runner only when it claims a job whose image is hosted on that registry: public images pull anonymously, and the credential never appears in any API response or log. Prefer a read-only token scoped to the images you run.
Cloud VM templates
The connect wizard also produces a cloud-init user-data document and Terraform snippets (AWS EC2, Azure VM) that boot a fresh instance straight into an enrolled runner with no inbound rule. The enrollment token inside them is single-use and expires 15 minutes after generation: apply immediately and never bake the document into a machine image.
Running on Kubernetes
The connect wizard's Kubernetes tab gives you a helm install for the marathoon-runner chart. One runner pod per release; every job becomes its own Kubernetes Job in the namespace, with no Docker socket and no privileged container.
- Each job pod has an emptyDir workspace: a staging init container waits while the runner streams the inputs in through the exec API, the job container sees /inputs and /outputs, and a small sidecar keeps the pod alive so the runner can pull the outputs out the same way. Job pods mount no service-account token and hold no Marathoon credential.
- The runner's Role is limited to jobs, pods, pods/exec, pods/log and secrets in its namespace; install it in a dedicated namespace. Private registries use a per-job kubernetes.io/dockerconfigjson pull Secret owned by the Job, so it disappears with it.
- Enrolled credentials (API key, agent id) live in the Secret <release>-credentials, so a restarted pod keeps its identity without a persistent volume. Set marathoon.resetCredentials=true with a fresh token to re-enrol.
- jobs.resources sets requests/limits on every job pod (required with a ResourceQuota); jobs.nodeSelector, jobs.workspaceSizeLimit and jobs.helperImage cover node placement, workspace size and air-gapped clusters. Cancellation and timeouts delete the Job; ttlSecondsAfterFinished is the safety net.
Direct runner ↔ storage transfers
When a storage provider has "Runners talk to this storage directly" enabled, a Connected Runner asks the API for a short-lived pre-authorized URL (15 minutes, one blob) and downloads inputs or uploads outputs straight from and to your bucket; the API only records the resulting artifact. Job content then never transits through Marathoon. If the runner cannot reach the bucket, it falls back to the API for that file and says so in the job log. Never available for the hosted default storage.
Zero shared secret
Instead of storing a key, you can grant Marathoon’s own cloud identity access: on AWS, a cross-account role Marathoon assumes with a per-provider External ID you put in the trust policy; on Azure, a Storage Blob Data Contributor assignment for Marathoon’s managed identity on your container. Revoking is a change on your side, with nothing to rotate. The identities to trust are shown in the connect form and come from the operator’s configuration.
How billing changes
| Mode | Compute | Billing |
|---|---|---|
| Hosted | Marathoon runs the container | Metered in credits per run |
| Your infra | Your Docker host runs it | Flat platform fee, no per-run compute charge |
Marketplace
The marketplace lets creators sell pipelines and lets any tenant install them. A listing wraps a global pipeline with pricing and goes through admin review before it becomes purchasable.
As a buyer
Browse the catalog at /marketplace, open a listing, and install it. Installed pipelines become runnable in your projects; they always run on Marathoon's hosted compute, never on your own runners.
As a creator
Become a creator in Dashboard → Marketplace → Creator, list a published pipeline, submit it for review, and earn on every purchase and every run.
Pricing models
| Model | Charge |
|---|---|
| OneTimePurchase | A one-time unlock price in USD |
| PerExecutionMarkup | Extra credits added on top of compute, on every run |
| Both | One-time unlock plus a per-run credit markup |
Revenue split: the platform keeps 20% and the creator 80% of both one-time sales and per-run markup — 15% / 85% once a creator's net earnings pass $10,000 over 12 months. Stripe fees are on the platform. Creators withdraw via Stripe once their balance reaches $50.
Listing lifecycle
Draft → PendingReview → Approved (or Rejected with notes). Approved listings can be Delisted at any time. Only Approved listings are purchasable and appear in the public catalog.
Agent & Batches
Two ways to run at scale without the UI: submit a batch of jobs at once, or install the desktop agent to turn a watched folder into automatic runs.
Batches
A batch groups many jobs submitted together — ideal for processing hundreds or thousands of files against one pipeline. Track them on the Batches page; each child job has its own status, logs and artifacts.
curl -X POST \
https://<domain>/api/v1/tenants/<tenantId>/jobs/batch/<pipelineId> \
-H "X-API-Key: <api-key>" \
-H "Content-Type: application/json" \
-d '{
"inputsList": [
{ "input_file": "s3://bucket/a.csv" },
{ "input_file": "s3://bucket/b.csv" }
]
}'Desktop agent
Pair the desktop agent from Dashboard → Agent → Pair. One click logs you in, mints a pipeline-scoped API key, and writes the local config. The agent then watches a directory:
- Drop a single file → a one-shot run starts
- Drop a sub-folder → one run with all files in it as inputs (batch mode)
- Live log tail and run details are shown in the agent window
The paired key is scoped to a single pipeline, so the agent can only submit jobs for that pipeline.
Authentication
All API requests require authentication. Two methods are supported.
JWT (User Session)
Obtained via POST /api/v1/auth/login. Pass the token as a Bearer header:
Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...
API Keys
Create API keys from Settings → API Keys or via API. Keys are shown once on creation — copy immediately. Send the key in the X-API-Key header (the Authorization header with the ApiKey scheme also works).
X-API-Key: mk_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Creating a key via API:
curl -X POST \
https://<domain>/api/v1/tenants/<tenantId>/api-keys \
-H "Authorization: Bearer <jwt>" \
-H "Content-Type: application/json" \
-d '{
"name": "CI Pipeline Key",
"preset": "Operator",
"expiresAtUtc": "2027-01-01T00:00:00Z"
}'Permission presets:
| Preset | Capabilities |
|---|---|
| ReadOnly | Read jobs, pipelines, projects. No write access. |
| Operator | Submit/cancel/retry jobs, download artifacts, read everything. |
| Admin | Full tenant access: manage API keys, settings, team, billing. |
| Custom | Granular permissions specified in customPermissions array. |
API Reference
All endpoints are prefixed with /api/v1/tenants/{tenantId}
Jobs
/jobsList jobs (filterable by pipeline, project, status, date)/jobs/{jobId}Get job detail/jobs/run/{pipelineId}Submit a single job/jobs/batch/{pipelineId}Submit a batch of jobs/jobs/{jobId}/cancelCancel a running or queued job/jobs/{jobId}/retryRetry a failed job/jobs/{jobId}Delete a completed/failed job/jobs/{jobId}/logsGet structured logs (paginated)/jobs/{jobId}/container-logsGet raw container stdout/stderr/jobs/{jobId}/artifacts/{artifactId}/downloadDownload an output artifactSubmit job body:
{
"inputs": { "key": "value" },
"parameterOverrides": {
"cpuUnits": 4,
"memoryMb": 8192,
"timeoutSeconds": 1800,
"environmentVariables": { "DEBUG": "true" }
}
}Pipelines
Prefix: /projects/{projectId}/pipelines
/projects/{projectId}/pipelinesList pipelines in a project/projects/{projectId}/pipelines/{pipelineId}Get pipeline detail/projects/{projectId}/pipelinesCreate a pipeline/projects/{projectId}/pipelines/{pipelineId}Update definition (creates new version)/projects/{projectId}/pipelines/{pipelineId}Delete a pipeline/projects/{projectId}/pipelines/{pipelineId}/publishPublish (make runnable)/projects/{projectId}/pipelines/{pipelineId}/unpublishUnpublish/projects/{projectId}/pipelines/{pipelineId}/versionsList all versions/projects/{projectId}/pipelines/{pipelineId}/versions/{n}Get a specific version/projects/{projectId}/pipelines/{pipelineId}/versions/{n}/revertRevert to a previous version/projects/{projectId}/pipelines/{pipelineId}/versions/{a}/diff/{b}Diff two versions/projects/{projectId}/pipelines/{pipelineId}/validateValidate a definition without saving/projects/{projectId}/pipelines/{pipelineId}/test-runsSubmit a test run (unpublished ok)/projects/{projectId}/pipelines/{pipelineId}/parametersUpdate default parametersAPI Keys
/api-keysList all API keys/api-keys/presetsList available permission presets/api-keys/{apiKeyId}Get a specific key/api-keysCreate an API key (key shown once)/api-keys/{apiKeyId}Update name/description/permissions/api-keys/{apiKeyId}/revokeDeactivate without deleting/api-keys/{apiKeyId}Permanently deleteError Format
All error responses follow this shape:
{
"error": "Human-readable error message",
"code": "Domain.ErrorCode"
}| HTTP | Meaning |
|---|---|
| 400 | Validation error or bad request |
| 401 | Missing or invalid authentication |
| 402 | Insufficient credits |
| 403 | Insufficient permissions |
| 404 | Resource not found |
| 409 | Conflict (e.g., job already cancelled) |
Webhooks
Configure webhooks from Settings → Webhooks to receive a signed HTTP POST when a job or batch finishes. Deduplicate on the payload id: retries and replays reuse it.
Payload
{
"id": "6f1c2a4e-8a43-4a0e-9a8e-2d5c3c6b1f00",
"type": "job.completed",
"createdAt": "2026-06-17T12:34:56Z",
"organizationId": "...",
"data": {
"jobId": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"sequence": 1,
"executionId": "...",
"pipelineId": "...",
"status": "Succeeded",
"exitCode": 0,
"startedAtUtc": "2026-06-17T12:30:01Z",
"endedAtUtc": "2026-06-17T12:34:56Z",
"durationSeconds": 295
}
}Signature
Every request carries X-Marathoon-Signature: t=<unix seconds>,v1=<hex>, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your signing secret (shown once, at creation or rotation). Recompute it over the raw body, accept the request if any v1 matches (for 24 hours after a rotation there is one per secret), and reject requests whose t is more than 5 minutes old. X-Marathoon-Event carries the event type and X-Marathoon-Delivery the delivery id. Deliveries can arrive out of order: reorder a job's events by data.sequence. Use the Test button to send a signed webhook.ping.
import { createHmac, timingSafeEqual } from 'node:crypto';
function verify(rawBody, header, secret) {
const parts = header.split(',').map((p) => p.split('='));
const t = parts.find(([k]) => k === 't')?.[1];
if (!t || Math.abs(Date.now() / 1000 - Number(t)) > 300) return false;
const expected = createHmac('sha256', secret).update(`${t}.${rawBody}`).digest('hex');
// After a rotation there is one v1 per secret for 24 h: any match is enough.
return parts.some(([k, v]) => k === 'v1' && v.length === expected.length
&& timingSafeEqual(Buffer.from(v), Buffer.from(expected)));
}Retries
Answer with a 2xx within 10 seconds. Timeouts, network errors, 408, 429 and 5xx are retried after 1 min, 5 min, 30 min, 2 h, 6 h, 12 h and 24 h; other 4xx and redirects are not retried, and 410 Gone disables the webhook. Endpoints must use HTTPS.
Events
job.startedJob's container started runningjob.completedJob finished successfullyjob.failedJob failed (non-zero exit, error, or could not start)job.timed_outJob exceeded its timeoutjob.cancelledJob cancelledbatch.completedAll jobs of a batch finishedpipeline.publishedPipeline publishedpipeline.updatedNew pipeline version savedmember.addedUser joined the organizationmember.removedMember removed from the organization