小王算钱

Interview Prep · Data Infra

Job Scheduling
From cron to Airflow and Temporal

cron runs a command when the clock says so, and nothing more. A data pipeline needs four other things: dependencies between tasks, a record of each run, retries and backfills after a failure, and no reliance on one machine. The seven systems below are in the order they appeared. Each one adds something the previous one lacked, and brings a new problem with it.

Timeline

From cron to Temporal: what each step adds that the last one lacked

  1. Origin

    Started by Maxime Beauchemin at Airbnb in October 2014, officially announced in June 2015, entered the Apache Incubator in March 2016 and became a top-level project in January 2019

    The problem it set out to solve

    Oozie's XML was hard to write and it could not leave Hadoop. Luigi had no triggering and no history of each run. Airflow wanted Python definitions, its own scheduling, the state of every run visible in a UI, and failed runs that can be rerun from that UI.
    Airflow 3: its parts
    • DAG filesPython code, kept in a DAG bundle
    • DAG processorParses DAG files and writes the serialized DAGs to the metadata database
    • Scheduler (contains the executor)Creates DAG runs, picks the tasks that can run and hands them to the executor
    • Metadata databasePostgreSQL or MySQL; holds the state of DAGs, runs and tasks
    • WorkerExecutes tasks; in Airflow 3 it reports state through the API server and does not connect to the database
    • API serverThe REST API and the UI; also receives the state tasks report
    • Triggerer (optional)Waits on deferred tasks in an asyncio event loop

    How it works

    • DAGs are written in Python files. The DAG processor parses those files and stores the serialized structure in the metadata database.
    • The scheduler loops over three jobs: create DAG runs for DAGs that are due; find tasks whose upstream tasks have all succeeded; hand tasks to the executor within the limits of pools and concurrency settings.
    • The executor decides where a task runs. LocalExecutor runs it on the scheduler's machine, CeleryExecutor hands it to long-running workers, and KubernetesExecutor starts a pod per task.
    • When scheduling by data interval (the default in Airflow 2), each DAG run has a period of time it is meant to process, the logical date is the start of that period, and the run is triggered only after the period has ended. In Airflow 3 a plain cron expression no longer defines a period by default, and the logical date is simply the moment the run is scheduled to trigger.
    • All state lives in the metadata database, usually PostgreSQL or MySQL. Clearing a task instance makes the scheduler run it again.

    Defining and submitting a workflow

    1. Write a Python file that defines a DAG: its schedule, its start date and its tasks.
    2. Declare dependencies between tasks with >> or with TaskFlow function calls.
    3. Put the file in the DAG bundle.
    4. The DAG processor parses it and writes it to the metadata database, after which it appears in the UI.

    One run, step by step

    1. When the trigger time arrives, the scheduler creates the DAG run and its task instances.
    2. The scheduler marks tasks whose upstream tasks have all succeeded as scheduled.
    3. Within pool and concurrency limits it marks them queued and hands them to the executor.
    4. A worker executes the task and reports its state.
    5. When a task succeeds, the scheduler releases its downstream tasks on the next loop.
    6. When a task fails with retries left, its state becomes up_for_retry and the scheduler schedules it again after the retry delay.
    7. When every task has ended, the DAG run is marked success or failed.

    Strengths

    • A DAG is Python code and can be generated dynamically.
    • It schedules on its own, and each run has its own logical date, so reruns and backfills are built in.
    • The UI shows the state and logs of every task in every run.
    • There are many ready-made operators and providers, and the executor can be swapped.
    • Several schedulers can run at once and cover for one another.

    Weaknesses

    • A task does not start the moment its upstream finishes. It waits for the scheduler's next loop, so Airflow is a poor fit for large numbers of very short tasks.
    • The metadata database sits at the centre of every component. When it is slow, the whole system is slow.
    • Tasks pass data to each other through XCom, which is meant for small values only. Large data has to go to external storage.
    • A DAG is code. Before Airflow 3, changing the DAG file made past runs show up in the UI with the new structure.

    Where it is used, and where it still is

    Batch pipelines that run on a timetable and have a clear start and end: ETL, reporting, periodic model training.

    What replaced it, and why

    It has not been replaced. Systems like Dagster go further in scheduling by data instead of by task, and Airflow 3 is adding features in the same direction.

    In an interview, one sentence

    DAGs defined in Python, a scheduler that creates one run per time interval, state in a database and pluggable execution.

    In an interview, two minutes

    At its core Airflow has a DAG processor, a scheduler, a metadata database and an API server, plus the workers that execute tasks. The DAG processor parses Python files and stores each DAG's structure in the metadata database. The scheduler loops over three jobs: creating runs for DAGs that are due, finding tasks whose upstream tasks are complete, and handing them to the executor within concurrency limits. The executor is a setting inside the scheduler that decides where a task runs, which can be locally, on Celery workers or in a pod per task. The metadata database holds all state. The key concept is the logical date: each run has a fixed logical time that does not change on a rerun, which is what makes reruns and backfills well defined. When scheduling by data interval, a run processes one definite period and is triggered only after that period ends, so for a daily DAG the run on October 2 has a logical date of October 1. Airflow 3 changed the default: with a plain cron expression there is no period, and the logical date equals the trigger time. Its weak points are scheduling delay, which makes it a poor fit for many short tasks, and heavy reliance on the metadata database. Version 2.0 added multiple schedulers. Version 3.0 made workers report through an API instead of connecting to the database, and added DAG versioning.

    Worth adding

    The follow-up you are most likely to get is the logical date. For a daily DAG scheduled by data interval, the run triggered at midnight on October 2 has a logical date of October 1, because it processes the data for October 1. Airflow 3 changed the default for cron expressions, and the same run has a logical date of October 2. Before answering, say which version and which timetable you mean.

Seven concepts first

Words that keep coming back

DAG
A directed acyclic graph. Nodes are tasks and arrows mean "waits for". With no cycles there is always an order to run things in.
Nearly every batch scheduler uses one. A cycle would mean two tasks waiting on each other, so neither could ever start.
logical date / nominal time
Which period of data a run is meant to process, as opposed to the moment it actually starts. Oozie calls it nominal time, Airflow calls it the logical date.
It is what gives reruns and backfills a meaning: rerunning the October 1 run still processes the data for October 1.
Idempotent
Running the same task for the same period once or several times gives the same result.
No scheduler can promise that a task runs exactly once. Retries, backfills and failures of the scheduler itself can each make the same task run again.
backfill / catchup
Running past periods. Catchup is the scheduler filling in periods it missed on its own; a backfill is a person asking for a stretch of history to be run. In Airflow 3, catchup is off by default.
A new pipeline needs its history filled in, and changed logic needs recomputing. Neither is possible if tasks are not idempotent.
Executor
The layer that decides where and in what form a task executes: a local process, a long-running worker, or a container per task.
It sets the isolation and the startup delay. Long-running workers start quickly but share resources; a pod per task isolates well but starts slowly.
Sensor / data trigger
Running when a condition becomes true instead of at a time: a file has arrived, an upstream table has been updated.
With time-only triggers you guess when upstream will finish. Guess early and you read incomplete data; guess late and you wait for nothing.
Durable execution
Persisting how far a piece of code has got, so that after a crash it continues from there instead of starting over.
It is the line between systems like Temporal and DAG schedulers.

In depth · 1

One Airflow run, from start to finish

Take a DAG that runs once a day and is scheduled by data interval, and follow how the data for October 1 gets processed.

  1. 1

    The DAG file is parsed

    The DAG processor reads the DAG file, executes it to obtain the DAG's structure, serializes that and writes it to the metadata database. From then on the scheduler reads only that stored structure and never touches the DAG code.

  2. 2

    A time interval ends

    For a daily DAG, the data interval for October 1 ends at midnight at the start of October 2. Before that moment the run is not created. This is how scheduling by data interval behaves, and it is the default in Airflow 2. In Airflow 3 a plain cron expression triggers the run at the same moment, midnight on October 2, but gives it a logical date of October 2 and no period; to get the earlier behaviour you use CronDataIntervalTimetable.

  3. 3

    The scheduler creates the DAG run

    The scheduler sees that this DAG is due a new run. It inserts a DAG run into the database with a logical date of October 1, and one task instance for each task in the DAG.

  4. 4

    It picks the tasks that can run

    The scheduler examines the task instances. Those whose upstream tasks have all succeeded (the default trigger rule) are set to scheduled. Then, within pool and concurrency limits, it sets them to queued and hands them to the executor. This step locks rows of the pool table, so several schedulers running at once do not exceed a pool's limit.

  5. 5

    The task executes and reports

    A worker receives the task and executes it in a subprocess. In Airflow 3 the task reports its state through the API server. In Airflow 2, task code connected to the metadata database directly.

  6. 6

    Failure and retry

    If the task fails and has retries left, its state becomes up_for_retry and it returns to the scheduler once the retry delay has passed. When retries run out the state is failed, and downstream tasks become upstream_failed.

  7. 7

    A person reruns it

    Once the problem is fixed, you clear the failed task instance. Its state is reset and the scheduler picks it up on its next loop. The logical date does not change, so it still processes the data for October 1.

In depth · 2

A scheduler cannot give you exactly-once, so tasks have to be idempotent

This is the follow-up interviewers ask most often about scheduling. It does not depend on which system you use.

  1. 1

    Why exactly-once is out of reach

    A worker finishes a task and loses its network connection before reporting success. The scheduler never hears the result and cannot tell "did not run" from "ran but did not report". It has two choices. Run the task again, and it may execute twice. Do not run it again, and it may never have completed. With retries configured you get the first, called at-least-once. Without retries you get the second, at-most-once. Neither is exactly-once.

  2. 2

    Overwrite by period, do not append

    The task writes "the partition for October 1" and replaces it whole each time. However many times it runs, the result is the same. If it INSERTs into a table, running twice gives two copies of the data.

  3. 3

    Use the logical date, not the current time

    If a task says now() or "yesterday", then backfilling last month's run processes today's data. The period has to be passed in by the scheduler.

  4. 4

    Write to a temporary place, then rename atomically

    A failure halfway through leaves half a file, and downstream tasks treat it as complete. Luigi decides completion by whether the output exists, which is why it lists atomic file operations among its design goals.

  5. 5

    Give calls to outside systems an idempotency key

    Sending an email or charging a card cannot be overwritten. Give each call a key derived from the process ID and the step, and let the other side deduplicate. Temporal retries an Activity without limit by default, and its documentation advises writing Activities to be idempotent.

Side by side

All seven systems in one table

SystemHow a workflow is definedWhat triggers a runWhere state livesWhere tasks runBest for
cronA line in a crontabTimeNone keptA process on the same machineIndependent timed tasks on one machine
OozieXML (hPDL)Time + data availabilityRelational databaseJobs on the Hadoop clusterHadoop clusters that are still running
LuigiPython classesNone; started by cronWhether the output file existsThe worker process that was startedSmaller pipelines whose outputs are files
AirflowPython (DAGs)Time, asset updates, external eventsMetadata databaseDecided by the executor: local, Celery workers or podsBatch pipelines that run on a timetable
Argo WorkflowsYAML (a Kubernetes CRD)Runs when submitted; CronWorkflow for schedulesThe Workflow resource's status (etcd)One pod per stepContainerized work already on Kubernetes, such as ML training and CI
DagsterPython (assets)Time, sensors, upstream asset changesA database recording runs and asset eventsDecided by the run launcher and executorData platforms that care about lineage and freshness
TemporalOrdinary code (a Workflow function)API calls, signals, schedulesEvent HistoryWorkers the user deploysLong-running business processes

Where the industry is heading · checked October 2026

Where things have been going

  • The Hadoop-era schedulers are leaving

    Apache Oozie was retired in February 2025 and moved to the Apache Attic. Apache's page gives no reason. A reasonable explanation is that once computation left Hadoop, a scheduler designed only for Hadoop had no place.

  • Airflow 3 separates tasks from the scheduler's database

    Airflow 3.0 was released on April 22, 2025. Tasks no longer connect to the metadata database and report through the API server instead. It also added DAG versioning, backfills managed by the scheduler and scheduling on assets. Version 3.0 also changed two defaults: catchup is off, and a plain cron expression no longer defines a data interval, so the logical date equals the trigger time. After that, 3.1 (September 2025) added human-in-the-loop and deadline alerts, and 3.2 (April 2026) added asset partitioning. The latest version found was 3.3.2, released on September 17, 2026.

  • From scheduling tasks to scheduling data

    Dagster makes the asset its basic unit. Airflow has been able to trigger DAGs on data updates since 2.4, and in Airflow 3 the concept is called an asset. The direction is the same: let the scheduler know what a task produced.

  • The companies behind orchestrators are merging

    On July 13, 2026, Prefect announced it was acquiring Dagster Labs. The announcement says Dagster keeps its name and its open-source licence, and that the combined company operates under the Prefect name. For anyone choosing a tool, the company behind the project is part of the choice.

  • Time is no longer the only trigger

    Oozie's coordinator could wait for data long ago. Today's systems make this a first-class feature: an upstream asset update, an external event or a sensor can each trigger a run.

  • Execution is handed to Kubernetes

    Argo graduated from the CNCF in December 2022. Airflow and Dagster can both run each task in a pod. With Kubernetes available, a scheduler no longer has to maintain its own fleet of long-running workers.

  • Durable execution is a category of its own

    Systems like Temporal do not compete with DAG schedulers for data pipelines; they are aimed at business processes. Asked in an interview to design a reliable multi-step process, first work out which of the two kinds the question is about.

Common misconceptions

Seven things that are easy to get wrong in an interview

  • Misconception: Airflow's logical date is the time the task starts running.

    The logical date is a logical time the scheduler assigns to the run. It does not change on a rerun, and it is not the moment the task actually starts. When scheduling by data interval (the default in Airflow 2), it is the start of the period being processed: for a daily DAG, the run triggered at midnight on October 2 has a logical date of October 1. In Airflow 3 with a plain cron expression, the logical date defaults to the scheduled trigger time, midnight on October 2. A task should decide which day's data to read from the time the scheduler passes in, never from the current time.

  • Misconception: A scheduler can guarantee each task executes exactly once.

    No scheduler can guarantee that. With retries configured you get at-least-once: if a worker finishes a task but does not get to report it, the task runs again. Without retries you get at-most-once. The Kubernetes documentation is blunter about CronJob: it may create two Jobs, or none. The effect of exactly-once has to come from the task being idempotent.

  • Misconception: Airflow is an engine that processes data.

    Airflow is an orchestrator: it decides what runs when. The computation usually happens elsewhere, in Spark, a data warehouse or a container. Nor is it designed for stream processing; it manages batch work that has a start and an end.

  • Misconception: Temporal is a newer, better Airflow.

    They do not solve the same problem. Airflow processes data one period at a time on a timetable. Temporal keeps individual long-running business processes from losing progress. Temporal has no backfills or data intervals, and Airflow has no step-by-step persistence.

  • Misconception: Luigi's central scheduler executes the tasks.

    It executes nothing. It only makes sure two instances of the same task do not run at once, and serves a UI. Tasks are run by worker processes, and a worker has to be started by cron or a person.

  • Misconception: Airflow with the KubernetesExecutor is the same as Argo Workflows.

    Only the execution layer is the same: a pod per task. Airflow still has its own scheduler and metadata database, and keeps state in that database. Argo keeps state in the Workflow resource in Kubernetes, and a controller does the scheduling.

  • Misconception: After a reboot, cron runs the tasks it missed.

    It does not. cron looks only at which lines match the current minute and ignores the past. Making up missed runs is something workflow schedulers can do, and only because each run corresponds to a definite period.

Staff-level interview questions

Answer first, then open

▶The company runs 200 scripts from cron, and downstream scripts often start before upstream ones have finished. What would you move to?
  • Start by naming the problem. What is missing here is dependencies and state, not execution capacity.
  • If the scripts are batch jobs that process data by day or by hour, a DAG scheduler such as Airflow fits: dependencies are written down, failures retry, and history can be backfilled.
  • The migration need not happen in one go. The scheduler can call the existing scripts unchanged at first, so the dependencies exist, and the scripts can be made idempotent one by one afterwards.
  • State the cost: there is now a scheduler, a database and workers to operate. For a dozen scripts with no dependencies, cron plus alerting may be enough.
  • In the end it depends on how dense the dependencies are, how much a failure costs in manual work, and whether the team has someone to run a new system.
▶Design a scheduler in which the scheduler process dying does not stop tasks from triggering on time.
  • Begin with where state lives. DAG runs and task instances are in a database and the scheduler process holds no state of its own, so a replacement can carry on.
  • Then how several schedulers avoid triggering the same thing twice. There are two routes: leader election (one works at a time, usually with the help of ZooKeeper, etcd or a lock in the database) or active-active (several work at once, and database row locks make sure each task is taken by only one).
  • Airflow 2.0 chose active-active, locking with SELECT ... FOR UPDATE and adding no extra coordination service. Oozie's high availability instead has several servers sharing a database and using ZooKeeper for locks.
  • Missed triggers also need handling. If every scheduler is down for ten minutes, on recovery it must work out which runs were due in that time. That requires "is a run due" to be computed from the timetable and the records in the database, not from timers held in memory.
  • Which route to take depends on what you already have. With a dependable database and no wish to operate another component, use row locks. Only when triggering is so frequent that database locks become the bottleneck is sharding or a dedicated coordination service worth it.
▶Airflow and Temporal are both called workflow engines. When do you use which?
  • Look at what triggers the process. A timetable that processes one period of data is Airflow's model. One instance per incoming request that waits on outside events is Temporal's.
  • Look at its shape. An Airflow DAG is largely fixed before it runs. A Temporal process is code: it can loop, branch on intermediate results and sleep for thirty days.
  • Look at the number of instances. How often an Airflow DAG runs is set by its timetable. Temporal is built for very many Workflow instances existing at once.
  • Look at what else you need. Backfills, data intervals and lineage come ready-made in Airflow and Dagster, and Temporal has none of them. In return, Temporal's persistence of every step of the code is absent from Airflow.
  • The two often sit side by side: Airflow for the data platform, Temporal for the product backend. In the end it comes down to which your process resembles: a timetable redrawn every day, or many separate business processes that each live a long time.
▶A daily task failed and nobody noticed for three days. How do you backfill it, and what do you watch for?
  • First check whether the task is idempotent. If it is not, rerunning creates duplicate data, so fix the task or clear the half-written results for those three days first.
  • Backfill by logical date, one run per day, not one run covering three days. Each run then has the same inputs and outputs as usual, and a problem is easy to locate.
  • Mind the downstream tasks. For three days they either did not run or read incomplete results, and they need rerunning too. When the scheduler knows the dependencies (in Airflow, clearing with downstream included) this can be automatic.
  • Mind concurrency. Will three days at once overwhelm the database or the cluster? Limit how many runs execute together.
  • Mind whether the upstream data still exists. Some sources keep only a few days, and once it is gone it cannot be recovered.
  • What needs adding afterwards is alerting: three days of failure unnoticed means SLA or deadline alerts are missing. How exactly to backfill depends on how idempotent the task is, how far downstream the damage reaches, and how long the source data is kept.
▶A scheduler has to support a million tasks a day. Where will the bottlenecks be?
  • Look at how long tasks run first. A million tasks of an hour each and a million tasks of two seconds each are two different problems.
  • For short tasks the bottleneck is the fixed cost per task: the scheduler's loop interval, the time to start a pod, a database write for every state change. Consider merging small tasks into one, or moving to long-running workers.
  • For the scheduler itself the bottleneck is usually the database. Every loop asks which tasks can run, and that gets slower as the task-instance table grows. The remedies are indexes, pruning history and sharing the load across several schedulers.
  • For systems like Argo that keep state in Kubernetes, the bottleneck is the API server and etcd, along with the 1 MB limit on a single Workflow resource.
  • Parsing is a bottleneck too. DAGs are code, and thousands of DAG files each have to be executed before their structure is known.
  • Measure before changing anything. The answer depends on task duration, whether strong isolation is needed, and how much scheduling delay is acceptable.
▶Why are newer schedulers moving from tasks to data (assets)? Is it worth migrating?
  • Scheduling by task, the scheduler knows only that a task succeeded. It does not know which table the task wrote, so it cannot say whether a table is current or what breaks if it does.
  • Scheduling by asset, dependencies are declared between pieces of data. The scheduler can then do three things: trigger downstream only when upstream has been updated, recompute only what is affected, and give you lineage directly.
  • The cost is remodelling. Tasks that produce no data do not fit, and the team has to write things differently.
  • It does not have to mean a new system. Airflow has had datasets since 2.4, and Airflow 3 renamed them assets and schedules on them, so they can be added to existing DAGs gradually.
  • Whether it is worth it depends on whether your pain is not knowing the state of your data. If the pain is only that tasks fail now and then, scheduling by task with good alerting is enough.

Hands-on exercise · about 5 minutes

A mini scheduler in forty-odd lines: dependencies, state, reruns

It uses only the Python standard library (3.9 or later), so there is nothing to install. Save the code below as mini_scheduler.py. It runs a four-step DAG once per date and records each task's state in state.json. The transform task for October 2 fails the first time.

import json
import os
from graphlib import TopologicalSorter

# task -> the tasks it depends on
DAG = {"extract": [], "transform": ["extract"], "load": ["transform"], "report": ["load"]}
DATES = ["2026-10-01", "2026-10-02", "2026-10-03"]
STATE_FILE = "state.json"  # plays the part of the metadata database


def run_task(task, date, attempt):
    # transform fails the first time it runs for 10-02, like an upstream outage.
    if task == "transform" and date == "2026-10-02" and attempt == 1:
        raise RuntimeError("upstream timeout")
    if task == "load":
        # Idempotent write: one file per date, overwritten, never appended.
        with open(f"out_{date}.txt", "w") as f:
            f.write(f"rows for {date}\n")


def schedule():
    state = json.load(open(STATE_FILE)) if os.path.exists(STATE_FILE) else {}
    for date in DATES:
        for task in TopologicalSorter(DAG).static_order():
            key = f"{date}/{task}"
            record = state.setdefault(key, {"status": "none", "tries": 0})
            if record["status"] == "success":
                print(f"{key:24} skip (already success)")
                continue
            if any(state[f"{date}/{up}"]["status"] != "success" for up in DAG[task]):
                record["status"] = "upstream_failed"
                print(f"{key:24} upstream_failed")
                continue
            record["tries"] += 1
            try:
                run_task(task, date, record["tries"])
                record["status"] = "success"
            except RuntimeError:
                record["status"] = "failed"
            print(f"{key:24} {record['status']} (try {record['tries']})")
    json.dump(state, open(STATE_FILE, "w"), indent=1)


schedule()
print("output files:", sorted(f for f in os.listdir(".") if f.startswith("out_")))

Run it twice in a row:

echo "=== run 1"; python3 mini_scheduler.py
echo "=== run 2"; python3 mini_scheduler.py

The real output of the two runs:

=== run 1
2026-10-01/extract       success (try 1)
2026-10-01/transform     success (try 1)
2026-10-01/load          success (try 1)
2026-10-01/report        success (try 1)
2026-10-02/extract       success (try 1)
2026-10-02/transform     failed (try 1)
2026-10-02/load          upstream_failed
2026-10-02/report        upstream_failed
2026-10-03/extract       success (try 1)
2026-10-03/transform     success (try 1)
2026-10-03/load          success (try 1)
2026-10-03/report        success (try 1)
output files: ['out_2026-10-01.txt', 'out_2026-10-03.txt']
=== run 2
2026-10-01/extract       skip (already success)
2026-10-01/transform     skip (already success)
2026-10-01/load          skip (already success)
2026-10-01/report        skip (already success)
2026-10-02/extract       skip (already success)
2026-10-02/transform     success (try 2)
2026-10-02/load          success (try 1)
2026-10-02/report        success (try 1)
2026-10-03/extract       skip (already success)
2026-10-03/transform     skip (already success)
2026-10-03/load          skip (already success)
2026-10-03/report        skip (already success)
output files: ['out_2026-10-01.txt', 'out_2026-10-02.txt', 'out_2026-10-03.txt']
  • In the first run, transform fails for October 2. Its downstream tasks, load and report, are not executed and end as upstream_failed. October 1 and October 3 are unaffected, because each date is a separate run.
  • The second run is in effect a rerun. Every task that already succeeded is skipped, and only the failed transform for October 2 and its downstream tasks actually execute. That is what keeping state buys you.
  • transform is on try 2 the second time, while load is on try 1. load was never executed the first time, so it did not count as a try.
  • load writes its file by overwriting. Delete state.json and run twice more: every task executes again, but the contents of the output files do not change. Change "w" to "a" and try once more, and you will see a task that is not idempotent leave duplicate lines after a rerun.
  • What these forty-odd lines leave out is most of what a real scheduler does: triggering by time, running tasks in parallel, locking when several schedulers run at once, and collecting state from tasks that run on other machines.

Cheat sheet

Ten minutes before the interview

cron
Runs a command on time. No dependencies, state, retries or backfills, and only on one machine.
Oozie
The scheduler for Hadoop. DAGs in XML; coordinators trigger on time plus data availability. Retired February 2025.
Luigi
Make, in Python. A task is complete when its output exists. No built-in triggering.
Airflow
DAGs in Python, a scheduler that creates runs per data interval, state in the metadata database, swappable executors.
Argo Workflows
A workflow is a Kubernetes CRD, each step is a pod, and state is in the resource's status.
Dagster
Orchestration by asset. Declare the data and how to compute it; the system derives dependencies and gives you lineage.
Temporal
Durable execution. Records every step and recovers by replaying the Event History. Workflows must be deterministic, Activities idempotent.
logical date
A run's logical time, unchanged on a rerun, and not the moment the task actually starts. The start of the period when scheduling by data interval; equal to the trigger time for a cron expression in Airflow 3 by default.
Idempotency
No scheduler guarantees exactly-once; retries and backfills make tasks run again. Overwrite by period, and use the time the scheduler gives you, not the current time.
Scheduler high availability
State in a database, stateless schedulers. Several schedulers avoid double triggering by leader election or row locks.
How to choose
A DAG scheduler for processing data on a timetable; durable execution for one long process per request; a Kubernetes-native engine when the work is already containers.

Sources

Version numbers, release dates and project status were checked in October 2026. They may have changed since.

DISCUSSION

Join the discussion

Share a thought or question. Comments publish immediately; add your email only if you would like reply notifications.

Be the first to start the conversation.