Research · 2026-W40

How to detect silently failing scheduled agent jobs — after one of mine ran dead for 17 days

By Ruiqi Tan · Published 2026-09-28

A security scan on one of my servers died for 17 days and nothing noticed. A crontab edit had left it pointing at a path that no longer existed, so every scheduled fire failed instantly: no output, no alert, no trace. The monitoring was fine — nothing was broken *inside* the job, because the job never started. What was missing was anything responsible for noticing that a machine which used to run had stopped running.

That failure generalizes, and it is worth stating as a law: **a checker that never fires is indistinguishable from a healthy system.** Silence is the default output of both a working schedule and a dead one. Any alerting built only on what jobs *emit* inherits this blind spot — and agent operations make it worse, because an autonomous system that has gone quiet looks exactly like an autonomous system with nothing to report.

The mechanism: declare once, verify daily

The fix I run — published as [schedule-sentinel](https://github.com/yagebin79386/schedule-sentinel), a single stdlib-Python file — is a registry plus a verifier. Every scheduled machine (cron entries, systemd timers, agent-runner jobs) is declared once, with a purpose, a consumer, a heartbeat kind, and a staleness tolerance. The verifier then proves three things, in the order they catch failures:

1. **The entry still exists.** The cheapest check, and the one that would have caught my 17 silent days on day one. 2. **The heartbeat is fresh** — the machine ran inside its tolerance, and where the scheduler records status, its last run did not fail. This catches both "scheduled but erroring" and "quietly unscheduled". 3. **Drift, in both directions.** A registry entry the scheduler lost, and — the direction people forget — a scheduler entry the registry does not know about. An unowned job is how a blind spot accumulates: it runs, it may even fail, and nothing says who cares.

A healthy run prints nothing. That is deliberate: a check that produces output on a clean day trains its reader to skip it.

The details that keep it honest

Never-ran is judged against registration time, so a job added an hour ago is not reported dead before its first fire. Machines on another host are skipped explicitly, never silently counted as passing. Tolerances sit above one schedule period, because a check that cries wolf gets muted — which is the exact condition the tool exists to prevent. And the test suite is written for *discrimination*: every check has one test proving it fires when broken and one proving it stays quiet when healthy, because a monitor missing either half can be permanently broken while looking perfect.

Freeze the verifier

One operational rule travels with the code: the verifier lives in the tier of the system humans control. If autonomous agents can edit their own supervisor, they can relax a tolerance or drop an entry and make their own breakage invisible.

Limits

This supervises *liveness*, not correctness — a machine can run on time and still do the wrong thing; that is what quality gates and [fail-closed approval gates](https://github.com/yagebin79386/fail-closed-gate) are for. It currently reads crontab, systemd timers, JSON job-state files, and file heartbeats; schedulers outside those need an adapter.

Claims & evidence

Every externally sourced claim in this article traces to a recorded source.

Limitations

Liveness supervision only; correctness requires separate quality and approval gates.