# Who watches the watcher: finding a scheduled job that vanished

> A deleted scheduled job leaves no error. Compare the expected roles with what the scheduler really holds, and alert on the time of the last success.

- Canonical: https://jaemyeong.com/en/blog/watch-the-watcher-retry-cron-reconciliation/
- Published: 2026.08.06
- Updated: 2026.10.04
- Category: IT/기술
- Tags: #CI-CD, #Automation, #Observability, #Reliability, #Cron

While checking the automation of an app, I once found the records and reality out of step. The job ledger said a retry job had been registered, but querying the scheduler returned only one gate poller. The gate kept doing its work, so the missing retry did not show. There was not enough evidence to pin down when or why it had been deleted.

A scheduled job that has been deleted cannot report its own failure. This post goes through a state comparison for finding that kind of gap, and a setup in which the monitoring does not stop together with what it monitors. Cron here is not only the Unix crontab. It also covers the GitHub Actions schedule and the Kubernetes CronJob. The original query output and logs, and measurements after a repair, are not included. The versions of the scheduler and tools used at the time, and the environment they ran in, were not recorded either.

## Zero errors is not zero runs

Different kinds of failure leave different traces.

- An error during a run leaves a start record, the error, and a failure count, and the job can retry from the inside.
- A run that stalls partway shows up as a start record with a stale heartbeat. Whether it can recover from the inside depends on conditions.
- An object that was made inactive is still in the scheduler but cannot recover from the inside.
- A deleted object, or a stopped scheduler, sends no report from the inside at all, and cannot recover from the inside either.

So zero errors alone cannot be read as healthy. The possibility of zero runs has to be ruled out first.

## Roles in the ledger, real objects in the query result

An ID written in the ledger is only a reference for looking things up. It is not evidence that the object is registered now. The desired state holds the role, the schedule, and the command, and the actual state holds the objects that were returned and whether they are active. Recreating an object can change its ID, so the comparison uses role keys such as `retry` and `gate`. Putting a version on the schedule and on the run template also helps detect outdated settings. The pseudocode for comparing the sets of roles is this.

```text
expected = desired_jobs_by_role
actual = scheduler.list()

missing    = expected.roles - actual.roles
duplicates = actual.group_by(role).where(count > 1)
mismatched = actual.where(spec_version != expected.spec_version)
```

It states what the comparison means, and a result of running it is not in this post. The pseudocode alone misses two cases. First, comparing only the version does not catch someone editing the schedule or the command without bumping the version. The actual schedule and command, or a hash of the whole desired specification, have to be compared for each role. Second, an object whose role is not in the expected list is neither missing nor a duplicate, so it would stay. `actual.roles - expected.roles` needs its own category, with an alert to investigate ownership and no automatic deletion.

One desired state is matched against four observed stages.

- desired: the expected roles in a versioned specification
- registered: whether the real object exists, and its enabled or suspend state
- executed: the last start time and the run history
- succeeded: a last success within the time limit
- progressed: a checkpoint, the queue depth, the completion time

Each stage tells you something different. According to [the Kubernetes CronJob documentation](https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/), when `.spec.suspend` is true, later runs stop even though the object exists. A run can also stall or fail after it starts. Even when the success time is recent, only a checkpoint tells whether the input was empty or a condition was wrong. I suggest obtaining the list of real objects and the last success time first, and widening to business metrics when the need comes up.

## What monitoring inside the same scheduler cannot see

The structure in the case I first saw had a flaw. The relation ran one way, with retry repairing the other poller, while that poller only did its own work. If retry is deleted, nothing is left to do the repair. Making the two check each other does not help against a failure that hits both in the same scheduler.

Comparing when a session restarts or when settings change provides a basic safeguard. It leaves a gap between one check and the next, though. When detection has to be continuous, an outside dead-man monitor checks the time of the last success. The outside monitor and the delivery of its alerts are things to check as well. The aim is to run the monitoring and what it monitors on separate foundations, so that one stopped scheduler does not stop both.

[The monitoring chapter of the Google SRE book](https://sre.google/sre-book/monitoring-distributed-systems/) calls monitoring based on metrics exposed by the internals of a system white-box, and testing externally visible behavior as a user would see it black-box. Seen this way, the internal list, the run history, and the heartbeat are white-box. My reading is that it is black-box only when a separate system directly checks the end result from the user's point of view, and that installing the monitor outside does not by itself make it black-box.

## The order for comparing and repairing

The repair is split into five steps.

```text
observe → classify → re-read → repair → verify
```

Query the list and the active state, then classify the result as missing, duplicate, inactive, drift, or unexpected, which is a role that is not in the expected list. Query once more to filter out races and temporary delays. After that, create only what is missing from a version-controlled literal template, check the role, the schedule, the command, and the spec version again, and update the ID in the ledger. The goal is one object per expected role. Several objects with the same role raise a duplicate alert, and a role that is not in the expected list raises an unexpected alert. Either way, an object whose ownership is unknown is not deleted.

Automatic repair is limited to missing objects whose ownership and specification are certain. The other states are handled like this.

- inactive: check the change history, then activate it explicitly.
- duplicate: alert on the risk of concurrent runs, and have a person look into the ownership.
- drift: check the expected version and who made the change.
- unexpected: have a person find out who created it and why. It is not deleted automatically.

An inactive state or a drift that someone made on purpose must not be overwritten. This comparison runs when a process starts or resumes, and right after settings change. When detection has to be faster, add a periodic check from outside.

## Alerting on the time of the last success

[The Prometheus instrumentation guide](https://prometheus.io/docs/practices/instrumentation/) says the key metric of a batch job is the last time it succeeded. The same document says that to track the time since something happened, you export the Unix timestamp at which it happened, not the time since. The elapsed time is calculated at query time. The condition for whether the time since the last success of retry has passed the threshold is this.

```text
time() - job_last_success_timestamp_seconds{job="retry"} > stale_after_seconds
```

As a starting threshold I suggest twice the run interval plus the longest expected run time. [The Prometheus alerting guide](https://prometheus.io/docs/practices/alerting/) also says to page when a batch job has not succeeded recently enough, and that this should generally be at least enough time for 2 full runs. The value is not a number settled by measurement. It is adjusted by how much the run time varies and how much detection delay is acceptable. A normal success that comes one interval late has its alert held back within a grace window, and a run that passes its timeout after starting is classified as a stale execution or a hang.

The case where the series does not exist at all is looked at separately. [The PromQL functions reference](https://prometheus.io/docs/prometheus/latest/querying/functions/) has `absent()` and `absent_over_time()`. According to [the Pushgateway guide](https://prometheus.io/docs/practices/pushing/), the Pushgateway does not remove the series it received on its own. It keeps exposing them until they are deleted by hand, so a stale success value can remain. That is why a missing series and a stale success time are told apart.

- When the object is missing, look at registration, the comparison process, and the real list.
- When the success metric is missing, look at instrumentation, collection, the first run, and scrape or push.
- When the success time is stale, look for delay, failure, or a stall in the history and logs.

Where to look also depends on the observed stage.

- With no recent start, look at the trigger and the scheduler.
- With a start and no success, look at the logs and the timeout.
- With a recent success and a checkpoint that has stopped moving, look at the input and the business conditions.
- With no heartbeat from the watcher, look at the outside probe and the alerts.

What a scheduler guarantees also differs by product. According to the Kubernetes documentation, a CronJob creates a Job about once per scheduled time. In some circumstances two Jobs are created, or none, so Jobs should be idempotent. `startingDeadlineSeconds` and `concurrencyPolicy` are settings that control delay and concurrent runs, and they do not guarantee exactly-once execution. According to [the GitHub Actions schedule documentation](https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows#schedule), the schedule event can be delayed during periods of high load and queued jobs may be dropped, and it runs only when the workflow file is on the default branch, and only on that branch. Whatever the product, the things to look at in common are the object, the run history, and the success time.

The part that compares and classifies sets can be checked with unit tests. Real delays, duplicates, and gaps are things to check in staging or in an isolated production check. Results of running them are not in this post. Why the first gap happened was not settled either. Comparing intermittently leaves a gap in detection, and the outside monitor and its alerts can fail too. There can be a delay before a newly created object shows up in a query. Querying again right after creation can still show it as missing, and a second object would then be created. To avoid a duplicate, wait with a bounded retry until the object is visible, or make creation idempotent by using the role key as a unique name.
