[Proposal] Self-Healing Retries #73638
Unanswered
daniyarndh
asked this question in
Ideas
Replies: 1 comment
|
That is - again - discussion that should happen on devlist (and Ideally in human words). Most likely you will also need to come and explain your ideas, consequences, and answer questions of people (after initial discussion) on a dev call - see our wiki for schedule. One of such callsis tomorrow. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Currently, Airflow task retries are passive and fixed. When a task fails (e.g.,
retries=3, retry_delay=5m), Airflow re-executes the exact same code with the exact same executor configuration and resource requests.For transient code bugs, this is fine. However, in modern data engineering, many failures are caused by transient infrastructure and operational bottlenecks:
OOMKilledsignal. Retrying with the same memory allocation fails 3 times in a row, paging on-call engineers at 2 AM.Engineers routinely over-provision resources, for instance, requesting 32 GB RAM for a task that needs 4 GB on 29 out of 30 days, solely to prevent edge-case retry failures - wasting substantial compute budget.
Proposed Solution
We propose introducing Self-Healing Retries capability. Instead of treating all retries identically, Airflow should allow tasks to dynamically mutate their executor specs, execution queues, or retry strategies based on the caught exception or exit code. As a result, Airflow can automatically double the memory allocation after an Out-of-Memory crash, back off during API rate limits, or move jobs to a low-concurrency queue during database locks.
Airflow here acts as the decision brain (intercepting error classes and altering retry specs), while the underlying executor (Kubernetes, Celery, ECS) acts as the execution muscle (provisioning the updated container).
All reactions