Interview Prep · Data Infra
Job Scheduling
From cron to Airflow and Temporal
cron runs a command when the clock says so, and nothing more. A data pipeline needs four other things: dependencies between tasks, a record of each run, retries and backfills after a failure, and no reliance on one machine. The seven systems below are in the order they appeared. Each one adds something the previous one lacked, and brings a new problem with it.
Timeline
From cron to Temporal: what each step adds that the last one lacked
Origin
Started by Maxime Beauchemin at Airbnb in October 2014, officially announced in June 2015, entered the Apache Incubator in March 2016 and became a top-level project in January 2019The problem it set out to solve
Oozie's XML was hard to write and it could not leave Hadoop. Luigi had no triggering and no history of each run. Airflow wanted Python definitions, its own scheduling, the state of every run visible in a UI, and failed runs that can be rerun from that UI.Airflow 3: its parts - DAG filesPython code, kept in a DAG bundle
- DAG processorParses DAG files and writes the serialized DAGs to the metadata database
- Scheduler (contains the executor)Creates DAG runs, picks the tasks that can run and hands them to the executor
- Metadata databasePostgreSQL or MySQL; holds the state of DAGs, runs and tasks
- WorkerExecutes tasks; in Airflow 3 it reports state through the API server and does not connect to the database
- API serverThe REST API and the UI; also receives the state tasks report
- Triggerer (optional)Waits on deferred tasks in an asyncio event loop
How it works
- DAGs are written in Python files. The DAG processor parses those files and stores the serialized structure in the metadata database.
- The scheduler loops over three jobs: create DAG runs for DAGs that are due; find tasks whose upstream tasks have all succeeded; hand tasks to the executor within the limits of pools and concurrency settings.
- The executor decides where a task runs. LocalExecutor runs it on the scheduler's machine, CeleryExecutor hands it to long-running workers, and KubernetesExecutor starts a pod per task.
- When scheduling by data interval (the default in Airflow 2), each DAG run has a period of time it is meant to process, the logical date is the start of that period, and the run is triggered only after the period has ended. In Airflow 3 a plain cron expression no longer defines a period by default, and the logical date is simply the moment the run is scheduled to trigger.
- All state lives in the metadata database, usually PostgreSQL or MySQL. Clearing a task instance makes the scheduler run it again.
Defining and submitting a workflow
- Write a Python file that defines a DAG: its schedule, its start date and its tasks.
- Declare dependencies between tasks with >> or with TaskFlow function calls.
- Put the file in the DAG bundle.
- The DAG processor parses it and writes it to the metadata database, after which it appears in the UI.
One run, step by step
- When the trigger time arrives, the scheduler creates the DAG run and its task instances.
- The scheduler marks tasks whose upstream tasks have all succeeded as scheduled.
- Within pool and concurrency limits it marks them queued and hands them to the executor.
- A worker executes the task and reports its state.
- When a task succeeds, the scheduler releases its downstream tasks on the next loop.
- When a task fails with retries left, its state becomes up_for_retry and the scheduler schedules it again after the retry delay.
- When every task has ended, the DAG run is marked success or failed.
Strengths
- A DAG is Python code and can be generated dynamically.
- It schedules on its own, and each run has its own logical date, so reruns and backfills are built in.
- The UI shows the state and logs of every task in every run.
- There are many ready-made operators and providers, and the executor can be swapped.
- Several schedulers can run at once and cover for one another.
Weaknesses
- A task does not start the moment its upstream finishes. It waits for the scheduler's next loop, so Airflow is a poor fit for large numbers of very short tasks.
- The metadata database sits at the centre of every component. When it is slow, the whole system is slow.
- Tasks pass data to each other through XCom, which is meant for small values only. Large data has to go to external storage.
- A DAG is code. Before Airflow 3, changing the DAG file made past runs show up in the UI with the new structure.
Where it is used, and where it still is
Batch pipelines that run on a timetable and have a clear start and end: ETL, reporting, periodic model training.What replaced it, and why
It has not been replaced. Systems like Dagster go further in scheduling by data instead of by task, and Airflow 3 is adding features in the same direction.In an interview, one sentence
DAGs defined in Python, a scheduler that creates one run per time interval, state in a database and pluggable execution.
In an interview, two minutes
At its core Airflow has a DAG processor, a scheduler, a metadata database and an API server, plus the workers that execute tasks. The DAG processor parses Python files and stores each DAG's structure in the metadata database. The scheduler loops over three jobs: creating runs for DAGs that are due, finding tasks whose upstream tasks are complete, and handing them to the executor within concurrency limits. The executor is a setting inside the scheduler that decides where a task runs, which can be locally, on Celery workers or in a pod per task. The metadata database holds all state. The key concept is the logical date: each run has a fixed logical time that does not change on a rerun, which is what makes reruns and backfills well defined. When scheduling by data interval, a run processes one definite period and is triggered only after that period ends, so for a daily DAG the run on October 2 has a logical date of October 1. Airflow 3 changed the default: with a plain cron expression there is no period, and the logical date equals the trigger time. Its weak points are scheduling delay, which makes it a poor fit for many short tasks, and heavy reliance on the metadata database. Version 2.0 added multiple schedulers. Version 3.0 made workers report through an API instead of connecting to the database, and added DAG versioning.
Worth adding
The follow-up you are most likely to get is the logical date. For a daily DAG scheduled by data interval, the run triggered at midnight on October 2 has a logical date of October 1, because it processes the data for October 1. Airflow 3 changed the default for cron expressions, and the same run has a logical date of October 2. Before answering, say which version and which timetable you mean.
Seven concepts first
Words that keep coming back
- DAG
- A directed acyclic graph. Nodes are tasks and arrows mean "waits for". With no cycles there is always an order to run things in.
- Nearly every batch scheduler uses one. A cycle would mean two tasks waiting on each other, so neither could ever start.
- logical date / nominal time
- Which period of data a run is meant to process, as opposed to the moment it actually starts. Oozie calls it nominal time, Airflow calls it the logical date.
- It is what gives reruns and backfills a meaning: rerunning the October 1 run still processes the data for October 1.
- Idempotent
- Running the same task for the same period once or several times gives the same result.
- No scheduler can promise that a task runs exactly once. Retries, backfills and failures of the scheduler itself can each make the same task run again.
- backfill / catchup
- Running past periods. Catchup is the scheduler filling in periods it missed on its own; a backfill is a person asking for a stretch of history to be run. In Airflow 3, catchup is off by default.
- A new pipeline needs its history filled in, and changed logic needs recomputing. Neither is possible if tasks are not idempotent.
- Executor
- The layer that decides where and in what form a task executes: a local process, a long-running worker, or a container per task.
- It sets the isolation and the startup delay. Long-running workers start quickly but share resources; a pod per task isolates well but starts slowly.
- Sensor / data trigger
- Running when a condition becomes true instead of at a time: a file has arrived, an upstream table has been updated.
- With time-only triggers you guess when upstream will finish. Guess early and you read incomplete data; guess late and you wait for nothing.
- Durable execution
- Persisting how far a piece of code has got, so that after a crash it continues from there instead of starting over.
- It is the line between systems like Temporal and DAG schedulers.
In depth · 1
One Airflow run, from start to finish
Take a DAG that runs once a day and is scheduled by data interval, and follow how the data for October 1 gets processed.
- 1
The DAG file is parsed
The DAG processor reads the DAG file, executes it to obtain the DAG's structure, serializes that and writes it to the metadata database. From then on the scheduler reads only that stored structure and never touches the DAG code.
- 2
A time interval ends
For a daily DAG, the data interval for October 1 ends at midnight at the start of October 2. Before that moment the run is not created. This is how scheduling by data interval behaves, and it is the default in Airflow 2. In Airflow 3 a plain cron expression triggers the run at the same moment, midnight on October 2, but gives it a logical date of October 2 and no period; to get the earlier behaviour you use CronDataIntervalTimetable.
- 3
The scheduler creates the DAG run
The scheduler sees that this DAG is due a new run. It inserts a DAG run into the database with a logical date of October 1, and one task instance for each task in the DAG.
- 4
It picks the tasks that can run
The scheduler examines the task instances. Those whose upstream tasks have all succeeded (the default trigger rule) are set to scheduled. Then, within pool and concurrency limits, it sets them to queued and hands them to the executor. This step locks rows of the pool table, so several schedulers running at once do not exceed a pool's limit.
- 5
The task executes and reports
A worker receives the task and executes it in a subprocess. In Airflow 3 the task reports its state through the API server. In Airflow 2, task code connected to the metadata database directly.
- 6
Failure and retry
If the task fails and has retries left, its state becomes up_for_retry and it returns to the scheduler once the retry delay has passed. When retries run out the state is failed, and downstream tasks become upstream_failed.
- 7
A person reruns it
Once the problem is fixed, you clear the failed task instance. Its state is reset and the scheduler picks it up on its next loop. The logical date does not change, so it still processes the data for October 1.
In depth · 2
A scheduler cannot give you exactly-once, so tasks have to be idempotent
This is the follow-up interviewers ask most often about scheduling. It does not depend on which system you use.
- 1
Why exactly-once is out of reach
A worker finishes a task and loses its network connection before reporting success. The scheduler never hears the result and cannot tell "did not run" from "ran but did not report". It has two choices. Run the task again, and it may execute twice. Do not run it again, and it may never have completed. With retries configured you get the first, called at-least-once. Without retries you get the second, at-most-once. Neither is exactly-once.
- 2
Overwrite by period, do not append
The task writes "the partition for October 1" and replaces it whole each time. However many times it runs, the result is the same. If it INSERTs into a table, running twice gives two copies of the data.
- 3
Use the logical date, not the current time
If a task says now() or "yesterday", then backfilling last month's run processes today's data. The period has to be passed in by the scheduler.
- 4
Write to a temporary place, then rename atomically
A failure halfway through leaves half a file, and downstream tasks treat it as complete. Luigi decides completion by whether the output exists, which is why it lists atomic file operations among its design goals.
- 5
Give calls to outside systems an idempotency key
Sending an email or charging a card cannot be overwritten. Give each call a key derived from the process ID and the step, and let the other side deduplicate. Temporal retries an Activity without limit by default, and its documentation advises writing Activities to be idempotent.
Side by side
All seven systems in one table
| System | How a workflow is defined | What triggers a run | Where state lives | Where tasks run | Best for |
|---|---|---|---|---|---|
| cron | A line in a crontab | Time | None kept | A process on the same machine | Independent timed tasks on one machine |
| Oozie | XML (hPDL) | Time + data availability | Relational database | Jobs on the Hadoop cluster | Hadoop clusters that are still running |
| Luigi | Python classes | None; started by cron | Whether the output file exists | The worker process that was started | Smaller pipelines whose outputs are files |
| Airflow | Python (DAGs) | Time, asset updates, external events | Metadata database | Decided by the executor: local, Celery workers or pods | Batch pipelines that run on a timetable |
| Argo Workflows | YAML (a Kubernetes CRD) | Runs when submitted; CronWorkflow for schedules | The Workflow resource's status (etcd) | One pod per step | Containerized work already on Kubernetes, such as ML training and CI |
| Dagster | Python (assets) | Time, sensors, upstream asset changes | A database recording runs and asset events | Decided by the run launcher and executor | Data platforms that care about lineage and freshness |
| Temporal | Ordinary code (a Workflow function) | API calls, signals, schedules | Event History | Workers the user deploys | Long-running business processes |
Where the industry is heading · checked October 2026
Where things have been going
The Hadoop-era schedulers are leaving
Apache Oozie was retired in February 2025 and moved to the Apache Attic. Apache's page gives no reason. A reasonable explanation is that once computation left Hadoop, a scheduler designed only for Hadoop had no place.
Airflow 3 separates tasks from the scheduler's database
Airflow 3.0 was released on April 22, 2025. Tasks no longer connect to the metadata database and report through the API server instead. It also added DAG versioning, backfills managed by the scheduler and scheduling on assets. Version 3.0 also changed two defaults: catchup is off, and a plain cron expression no longer defines a data interval, so the logical date equals the trigger time. After that, 3.1 (September 2025) added human-in-the-loop and deadline alerts, and 3.2 (April 2026) added asset partitioning. The latest version found was 3.3.2, released on September 17, 2026.
From scheduling tasks to scheduling data
Dagster makes the asset its basic unit. Airflow has been able to trigger DAGs on data updates since 2.4, and in Airflow 3 the concept is called an asset. The direction is the same: let the scheduler know what a task produced.
The companies behind orchestrators are merging
On July 13, 2026, Prefect announced it was acquiring Dagster Labs. The announcement says Dagster keeps its name and its open-source licence, and that the combined company operates under the Prefect name. For anyone choosing a tool, the company behind the project is part of the choice.
Time is no longer the only trigger
Oozie's coordinator could wait for data long ago. Today's systems make this a first-class feature: an upstream asset update, an external event or a sensor can each trigger a run.
Execution is handed to Kubernetes
Argo graduated from the CNCF in December 2022. Airflow and Dagster can both run each task in a pod. With Kubernetes available, a scheduler no longer has to maintain its own fleet of long-running workers.
Durable execution is a category of its own
Systems like Temporal do not compete with DAG schedulers for data pipelines; they are aimed at business processes. Asked in an interview to design a reliable multi-step process, first work out which of the two kinds the question is about.
Common misconceptions
Seven things that are easy to get wrong in an interview
Misconception: Airflow's logical date is the time the task starts running.
The logical date is a logical time the scheduler assigns to the run. It does not change on a rerun, and it is not the moment the task actually starts. When scheduling by data interval (the default in Airflow 2), it is the start of the period being processed: for a daily DAG, the run triggered at midnight on October 2 has a logical date of October 1. In Airflow 3 with a plain cron expression, the logical date defaults to the scheduled trigger time, midnight on October 2. A task should decide which day's data to read from the time the scheduler passes in, never from the current time.
Misconception: A scheduler can guarantee each task executes exactly once.
No scheduler can guarantee that. With retries configured you get at-least-once: if a worker finishes a task but does not get to report it, the task runs again. Without retries you get at-most-once. The Kubernetes documentation is blunter about CronJob: it may create two Jobs, or none. The effect of exactly-once has to come from the task being idempotent.
Misconception: Airflow is an engine that processes data.
Airflow is an orchestrator: it decides what runs when. The computation usually happens elsewhere, in Spark, a data warehouse or a container. Nor is it designed for stream processing; it manages batch work that has a start and an end.
Misconception: Temporal is a newer, better Airflow.
They do not solve the same problem. Airflow processes data one period at a time on a timetable. Temporal keeps individual long-running business processes from losing progress. Temporal has no backfills or data intervals, and Airflow has no step-by-step persistence.
Misconception: Luigi's central scheduler executes the tasks.
It executes nothing. It only makes sure two instances of the same task do not run at once, and serves a UI. Tasks are run by worker processes, and a worker has to be started by cron or a person.
Misconception: Airflow with the KubernetesExecutor is the same as Argo Workflows.
Only the execution layer is the same: a pod per task. Airflow still has its own scheduler and metadata database, and keeps state in that database. Argo keeps state in the Workflow resource in Kubernetes, and a controller does the scheduling.
Misconception: After a reboot, cron runs the tasks it missed.
It does not. cron looks only at which lines match the current minute and ignores the past. Making up missed runs is something workflow schedulers can do, and only because each run corresponds to a definite period.
Staff-level interview questions
Answer first, then open
▶The company runs 200 scripts from cron, and downstream scripts often start before upstream ones have finished. What would you move to?
- Start by naming the problem. What is missing here is dependencies and state, not execution capacity.
- If the scripts are batch jobs that process data by day or by hour, a DAG scheduler such as Airflow fits: dependencies are written down, failures retry, and history can be backfilled.
- The migration need not happen in one go. The scheduler can call the existing scripts unchanged at first, so the dependencies exist, and the scripts can be made idempotent one by one afterwards.
- State the cost: there is now a scheduler, a database and workers to operate. For a dozen scripts with no dependencies, cron plus alerting may be enough.
- In the end it depends on how dense the dependencies are, how much a failure costs in manual work, and whether the team has someone to run a new system.
▶Design a scheduler in which the scheduler process dying does not stop tasks from triggering on time.
- Begin with where state lives. DAG runs and task instances are in a database and the scheduler process holds no state of its own, so a replacement can carry on.
- Then how several schedulers avoid triggering the same thing twice. There are two routes: leader election (one works at a time, usually with the help of ZooKeeper, etcd or a lock in the database) or active-active (several work at once, and database row locks make sure each task is taken by only one).
- Airflow 2.0 chose active-active, locking with SELECT ... FOR UPDATE and adding no extra coordination service. Oozie's high availability instead has several servers sharing a database and using ZooKeeper for locks.
- Missed triggers also need handling. If every scheduler is down for ten minutes, on recovery it must work out which runs were due in that time. That requires "is a run due" to be computed from the timetable and the records in the database, not from timers held in memory.
- Which route to take depends on what you already have. With a dependable database and no wish to operate another component, use row locks. Only when triggering is so frequent that database locks become the bottleneck is sharding or a dedicated coordination service worth it.
▶Airflow and Temporal are both called workflow engines. When do you use which?
- Look at what triggers the process. A timetable that processes one period of data is Airflow's model. One instance per incoming request that waits on outside events is Temporal's.
- Look at its shape. An Airflow DAG is largely fixed before it runs. A Temporal process is code: it can loop, branch on intermediate results and sleep for thirty days.
- Look at the number of instances. How often an Airflow DAG runs is set by its timetable. Temporal is built for very many Workflow instances existing at once.
- Look at what else you need. Backfills, data intervals and lineage come ready-made in Airflow and Dagster, and Temporal has none of them. In return, Temporal's persistence of every step of the code is absent from Airflow.
- The two often sit side by side: Airflow for the data platform, Temporal for the product backend. In the end it comes down to which your process resembles: a timetable redrawn every day, or many separate business processes that each live a long time.
▶A daily task failed and nobody noticed for three days. How do you backfill it, and what do you watch for?
- First check whether the task is idempotent. If it is not, rerunning creates duplicate data, so fix the task or clear the half-written results for those three days first.
- Backfill by logical date, one run per day, not one run covering three days. Each run then has the same inputs and outputs as usual, and a problem is easy to locate.
- Mind the downstream tasks. For three days they either did not run or read incomplete results, and they need rerunning too. When the scheduler knows the dependencies (in Airflow, clearing with downstream included) this can be automatic.
- Mind concurrency. Will three days at once overwhelm the database or the cluster? Limit how many runs execute together.
- Mind whether the upstream data still exists. Some sources keep only a few days, and once it is gone it cannot be recovered.
- What needs adding afterwards is alerting: three days of failure unnoticed means SLA or deadline alerts are missing. How exactly to backfill depends on how idempotent the task is, how far downstream the damage reaches, and how long the source data is kept.
▶A scheduler has to support a million tasks a day. Where will the bottlenecks be?
- Look at how long tasks run first. A million tasks of an hour each and a million tasks of two seconds each are two different problems.
- For short tasks the bottleneck is the fixed cost per task: the scheduler's loop interval, the time to start a pod, a database write for every state change. Consider merging small tasks into one, or moving to long-running workers.
- For the scheduler itself the bottleneck is usually the database. Every loop asks which tasks can run, and that gets slower as the task-instance table grows. The remedies are indexes, pruning history and sharing the load across several schedulers.
- For systems like Argo that keep state in Kubernetes, the bottleneck is the API server and etcd, along with the 1 MB limit on a single Workflow resource.
- Parsing is a bottleneck too. DAGs are code, and thousands of DAG files each have to be executed before their structure is known.
- Measure before changing anything. The answer depends on task duration, whether strong isolation is needed, and how much scheduling delay is acceptable.
▶Why are newer schedulers moving from tasks to data (assets)? Is it worth migrating?
- Scheduling by task, the scheduler knows only that a task succeeded. It does not know which table the task wrote, so it cannot say whether a table is current or what breaks if it does.
- Scheduling by asset, dependencies are declared between pieces of data. The scheduler can then do three things: trigger downstream only when upstream has been updated, recompute only what is affected, and give you lineage directly.
- The cost is remodelling. Tasks that produce no data do not fit, and the team has to write things differently.
- It does not have to mean a new system. Airflow has had datasets since 2.4, and Airflow 3 renamed them assets and schedules on them, so they can be added to existing DAGs gradually.
- Whether it is worth it depends on whether your pain is not knowing the state of your data. If the pain is only that tasks fail now and then, scheduling by task with good alerting is enough.
Hands-on exercise · about 5 minutes
A mini scheduler in forty-odd lines: dependencies, state, reruns
It uses only the Python standard library (3.9 or later), so there is nothing to install. Save the code below as mini_scheduler.py. It runs a four-step DAG once per date and records each task's state in state.json. The transform task for October 2 fails the first time.
import json
import os
from graphlib import TopologicalSorter
# task -> the tasks it depends on
DAG = {"extract": [], "transform": ["extract"], "load": ["transform"], "report": ["load"]}
DATES = ["2026-10-01", "2026-10-02", "2026-10-03"]
STATE_FILE = "state.json" # plays the part of the metadata database
def run_task(task, date, attempt):
# transform fails the first time it runs for 10-02, like an upstream outage.
if task == "transform" and date == "2026-10-02" and attempt == 1:
raise RuntimeError("upstream timeout")
if task == "load":
# Idempotent write: one file per date, overwritten, never appended.
with open(f"out_{date}.txt", "w") as f:
f.write(f"rows for {date}\n")
def schedule():
state = json.load(open(STATE_FILE)) if os.path.exists(STATE_FILE) else {}
for date in DATES:
for task in TopologicalSorter(DAG).static_order():
key = f"{date}/{task}"
record = state.setdefault(key, {"status": "none", "tries": 0})
if record["status"] == "success":
print(f"{key:24} skip (already success)")
continue
if any(state[f"{date}/{up}"]["status"] != "success" for up in DAG[task]):
record["status"] = "upstream_failed"
print(f"{key:24} upstream_failed")
continue
record["tries"] += 1
try:
run_task(task, date, record["tries"])
record["status"] = "success"
except RuntimeError:
record["status"] = "failed"
print(f"{key:24} {record['status']} (try {record['tries']})")
json.dump(state, open(STATE_FILE, "w"), indent=1)
schedule()
print("output files:", sorted(f for f in os.listdir(".") if f.startswith("out_")))
Run it twice in a row:
echo "=== run 1"; python3 mini_scheduler.py
echo "=== run 2"; python3 mini_scheduler.pyThe real output of the two runs:
=== run 1
2026-10-01/extract success (try 1)
2026-10-01/transform success (try 1)
2026-10-01/load success (try 1)
2026-10-01/report success (try 1)
2026-10-02/extract success (try 1)
2026-10-02/transform failed (try 1)
2026-10-02/load upstream_failed
2026-10-02/report upstream_failed
2026-10-03/extract success (try 1)
2026-10-03/transform success (try 1)
2026-10-03/load success (try 1)
2026-10-03/report success (try 1)
output files: ['out_2026-10-01.txt', 'out_2026-10-03.txt']
=== run 2
2026-10-01/extract skip (already success)
2026-10-01/transform skip (already success)
2026-10-01/load skip (already success)
2026-10-01/report skip (already success)
2026-10-02/extract skip (already success)
2026-10-02/transform success (try 2)
2026-10-02/load success (try 1)
2026-10-02/report success (try 1)
2026-10-03/extract skip (already success)
2026-10-03/transform skip (already success)
2026-10-03/load skip (already success)
2026-10-03/report skip (already success)
output files: ['out_2026-10-01.txt', 'out_2026-10-02.txt', 'out_2026-10-03.txt']- In the first run, transform fails for October 2. Its downstream tasks, load and report, are not executed and end as upstream_failed. October 1 and October 3 are unaffected, because each date is a separate run.
- The second run is in effect a rerun. Every task that already succeeded is skipped, and only the failed transform for October 2 and its downstream tasks actually execute. That is what keeping state buys you.
- transform is on try 2 the second time, while load is on try 1. load was never executed the first time, so it did not count as a try.
- load writes its file by overwriting. Delete state.json and run twice more: every task executes again, but the contents of the output files do not change. Change "w" to "a" and try once more, and you will see a task that is not idempotent leave duplicate lines after a rerun.
- What these forty-odd lines leave out is most of what a real scheduler does: triggering by time, running tasks in parallel, locking when several schedulers run at once, and collecting state from tasks that run on other machines.
Cheat sheet
Ten minutes before the interview
- cron
- Runs a command on time. No dependencies, state, retries or backfills, and only on one machine.
- Oozie
- The scheduler for Hadoop. DAGs in XML; coordinators trigger on time plus data availability. Retired February 2025.
- Luigi
- Make, in Python. A task is complete when its output exists. No built-in triggering.
- Airflow
- DAGs in Python, a scheduler that creates runs per data interval, state in the metadata database, swappable executors.
- Argo Workflows
- A workflow is a Kubernetes CRD, each step is a pod, and state is in the resource's status.
- Dagster
- Orchestration by asset. Declare the data and how to compute it; the system derives dependencies and gives you lineage.
- Temporal
- Durable execution. Records every step and recovers by replaying the Event History. Workflows must be deterministic, Activities idempotent.
- logical date
- A run's logical time, unchanged on a rerun, and not the moment the task actually starts. The start of the period when scheduling by data interval; equal to the trigger time for a cron expression in Airflow 3 by default.
- Idempotency
- No scheduler guarantees exactly-once; retries and backfills make tasks run again. Overwrite by period, and use the time the scheduler gives you, not the current time.
- Scheduler high availability
- State in a database, stateless schedulers. Several schedulers avoid double triggering by leader election or row locks.
- How to choose
- A DAG scheduler for processing data on a timetable; durable execution for one long process per request; a Kubernetes-native engine when the work is already containers.
Sources
Version numbers, release dates and project status were checked in October 2026. They may have changed since.
- Kubernetes documentation: CronJob (concurrency policy, missed schedules, idempotency)
- Wikipedia: cron (Version 7 Unix, Vixie cron 1987)
- anacron manual (cron assumes the machine runs continuously and does not make up missed runs)
- Apache Attic: Oozie (retired February 2025)
- Oozie documentation: overview (workflow, coordinator, bundle)
- Oozie Workflow specification (hPDL, control nodes, callback and polling)
- Oozie Coordinator specification (nominal time, datasets, done-flag, concurrency controls)
- Oozie installation guide (databases, ZooKeeper-based high availability)
- Apache announcement: Oozie 5.0.0 (launcher moved from a MapReduce mapper to a YARN ApplicationMaster)
- Apache Incubator: Oozie incubation record (entered 2011-07-11, graduated 2012-08-28)
- Yahoo's Oozie repository (created June 2010)
- Luigi repository (authors, open-sourced late 2012, the comparison with GNU Make)
- Luigi documentation: design and limitations (scale, no built-in triggering, no distributed execution)
- Luigi documentation: the central scheduler
- Luigi documentation: Tasks (requires, output, run)
- Airflow documentation: home page (what Airflow is for: batch workflows with a start and an end)
- Airflow documentation: project history
- Airflow documentation: architecture overview
- Airflow documentation: DAG runs (data interval, logical date, catchup, backfill)
- Airflow documentation: timetables (CronTriggerTimetable versus CronDataIntervalTimetable, the Airflow 3 default)
- Airflow documentation: asset scheduling (added in 2.4)
- Airflow documentation: scheduler (multiple schedulers, row locks)
- Airflow documentation: executors
- Airflow documentation: Tasks (task instance states)
- Airflow documentation: XComs
- Airflow blog: Airflow 2.0 (December 17, 2020, scheduler high availability)
- Airflow release notes (3.0.0, 3.1.0, 3.2.0, 3.3.2)
- Argo Workflows documentation: home page (use cases)
- Argo Workflows documentation: architecture
- Argo Workflows documentation: compressing and offloading large workflows
- Argo Workflows documentation: CronWorkflow
- Argo Workflows documentation: the backfill pattern (a workflow you write that loops over dates)
- CNCF announcement: Argo graduates (December 6, 2022)
- Akuity blog: why the Argo project was created (Applatix, 2017)
- Nick Schrock: Introducing Dagster (July 2019)
- Dagster 0.14.0 release notes (February 2022, software-defined assets)
- Dagster: announcement of Prefect acquiring Dagster Labs (July 13, 2026)
- Dagster blog: Dagster 1.0 (August 2022)
- Dagster documentation: assets
- Dagster documentation: open-source deployment architecture
- Temporal: about the company (founders, the origin in Cadence)
- Amplify Partners: on investing in Temporal (from Cadence to Temporal in 2019)
- Temporal documentation: Workflow Execution (durability, replay, determinism)
- Temporal documentation: Event History limits
- Temporal documentation: retry policies (Activity defaults)
- Temporal documentation: the four services of the Temporal Server
DISCUSSION
Join the discussion
Share a thought or question. Comments publish immediately; add your email only if you would like reply notifications.Be the first to start the conversation.