Saylor InnovationsSAYLOR INNOVATIONS

Home / Guides / Linux & Systems

systemd Failure Resolver

Linux & Systems intermediate 8 min read Free to read · $0.01 via agent API Updated 2026-08-22

A systemd unit that starts reliably under its real service identity and environment, passes application readiness, obeys dependency/restart limits, and has a documented failure cause and rollback.

Diagnose a failed systemd unit from its real execution context, dependencies, permissions, environment, restart policy, and journal.

Free to read here. AI agents can also fetch this guide directly over x402 for $0.01 — no account, structured JSON delivery.

Agent API →

The result you are building

Finished Result:

A systemd unit that starts reliably under its real service identity and environment, passes application readiness, obeys dependency/restart limits, and has a documented failure cause and rollback.

Use this guide when

  • A service is failed, restart-looping, works manually but not under systemd, or starts before a dependency.
  • A daemon loses environment variables, permissions, paths, or network access after reboot.

Do not use it as a substitute for

  • Do not paste shell syntax into ExecStart unless explicitly invoking a shell with safe quoting.
  • Do not use restart loops to hide a deterministic startup failure.

Before you change anything

Collect the items below first. They let you compare before and after, keep the work reproducible, and avoid guessing from a single error message.

  • Unit name and full systemctl cat output.
  • Status, exit code, journal for current boot, and first error.
  • Intended service user/group, working directory, files, ports, and environment.
  • Dependency/readiness expectations and recent unit/drop-in changes.

Stop Before Proceeding:

Stop before weakening sandboxing, changing ownership recursively, or running the service as root merely to make it start. Isolate the denied resource and grant only required access.

Understand the system before fixing it

systemd does not run your interactive shell PATH, working directory, environment files, umask, limits, and credentials differ. Absolute paths and explicit configuration make the service reproducible.

Active is not ready A process may be running before its socket, database, migration, or core operation is usable. Readiness must test the service contract.

Exit status identifies the boundary 203/EXEC, 200/CHDIR, signal exits, timeout, and application codes lead to different tests.

Evidence-to-decision map

  Start with the row that most closely matches the evidence. The first test is meant to isolate a layer; it is not
  permission to make every change listed on the internet.

    Evidence                  Likely layer       First decisive check           What the result means

    203/EXEC                  Executable         `systemctl show -p ExecStart   Missing/non-executable path, wrong
                                                 UNIT` and `namei -l`           interpreter, filesystem mount, or permission.

    200/CHDIR                 WorkingDirecto     `systemctl show -p             Directory absent or inaccessible to service
                              ry                 WorkingDirectory UNIT`         user.

    Permission denied         Identity/sandbo    Run read-only access test as   Unix mode/ACL, groups, SELinux/AppArmor, or
                              x                  service user; inspect unit     systemd restriction.
                                                 hardening

    Restart loop              Application/con    Earliest journal error and     Fix deterministic cause; rate limit is
                              fig                `Restart=` policy              secondary.

    Manual works              Environment/co     `systemctl show` vs shell      Make dependency explicit in unit/config.
                              ntext              env/path/limits

Step-by-step procedure

Work in order. Record the output after each step. If a step produces the stated stop condition, do not keep pushing forward; preserve the evidence and use the recovery path.

01 Capture effective unit and first failure Why: Drop-ins can override the file you are reading.

Do: Save systemctl cat, show properties, status, and current-boot journal before editing.

systemctl cat UNIT systemctl show UNIT -p User -p Group -p ExecStart -p WorkingDirectory -p Environment systemctl status UNIT --no-pager -l journalctl -u UNIT -b --no-pager

Read the result: Use the earliest error and exit status, not repeated restart messages.

Next: Choose executable, directory, permission, dependency, or application branch.

02 Reproduce under service identity Why: Interactive success under your user does not match systemd context.

Do: Test executable existence, config read, directory traversal, port/path access as the configured service user without starting a duplicate daemon.

sudo -u SERVICEUSER test -x /ABSOLUTE/EXECUTABLE; echo $? sudo -u SERVICEUSER test -r /PATH/CONFIG; echo $? namei -l /ABSOLUTE/EXECUTABLE

Read the result: A failed test/read isolates permissions or path; successful access shifts to environment/application.

Next: Never copy secrets into command history.

03 Validate application outside restart loop Why: systemd can obscure fast exit output.

Do: Use the application's config-check or foreground/test command as the service user with explicit env file.

Read the result: Fix the first app validation error before restarting the unit.

Next: Keep production port/state safe while testing.

04 Correct unit with a drop-in Why: Vendor units should remain package-managed.

Do: Use systemctl edit UNIT for the narrow override; absolute paths, explicit working directory/environment file, dependencies, and restart policy.

sudo systemctl edit UNIT systemd-analyze verify /etc/systemd/system/UNIT.d/override.conf sudo systemctl daemon-reload sudo systemctl restart UNIT

Read the result: systemd-analyze verify and daemon reload must be clean.

Next: Restart once and inspect fresh logs.

05 Test readiness and dependency behavior Why: Startup success alone can race downstream consumers.

Do: Call local health/core operation, verify listener and dependency ordering, then reboot or restart dependencies if authorized.

systemctl is-active UNIT ss -ltnp curl -fsS http://127.0.0.1:PORT/health

Read the result: Ready behavior should be stable without rapid restarts.

Next: Tune timeouts/restart only after root cause is fixed.

06 Document recovery Why: Future operators need effective configuration and a known-good path.

Do: Save drop-in, version, service identity, config locations, health test, journal query, and rollback command.

Read the result: A clean reboot must reproduce success.

Next: Back up the drop-in before later edits.

Worked example

Starting Problem:

A Python service works from the shell but systemd exits 203/EXEC.

Evidence collected

  • Effective ExecStart points to /home/user/app/venv/bin/python.
  • The service runs as appsvc, which cannot traverse /home/user.
  • The project was moved but the unit path was not updated.
  • Application config validation succeeds from the new service-owned path.

Decision The executable path/context is invalid; running as root would mask rather than solve it.

Actions taken

  • Placed release and venv under a service-owned application directory.
  • Updated only the drop-in ExecStart/WorkingDirectory.
  • Verified unit, restarted, tested health, and rebooted.

Proof Of Completion:

Unit stays active, health/core request pass, service runs as appsvc, and no permission/exec errors return after reboot.

Why this example matters Reproducing under the service identity exposed the real boundary.

Verify, recover, and hand off

Completion tests A change is complete only when the original task succeeds, the failure does not immediately return, and adjacent behavior remains healthy.

  • Effective unit matches intended user, paths, environment, and dependencies.
  • Unit starts without restart loop and stays within limits.
  • Local readiness and one core operation pass.
  • Logs contain no new startup error.
  • Reboot/dependency restart behavior is documented and tested.

Rollback or safe recovery

  • Remove/restore the drop-in, daemon-reload, and restart the prior configuration.
  • Use systemctl revert UNIT only after saving current override and understanding vendor state.
  • Switch to multi-user/recovery access if the failed unit blocks graphical/remote login.

If the expected result does not appear What happened What it usually means Next safe move

Edit has no effect Different drop-in precedence or Inspect systemctl cat and reload; edit daemon not reloaded. effective source.

Status active; port absent Process is not ready/listens Inspect app logs/config and actual sockets. elsewhere.

Permission works manually as systemd sandbox or mount Inspect systemctl show hardening properties. service user namespace differs.

Starts only after manual delay Dependency/order/readiness is Use correct After/Requires and application missing. retry/readiness, not sleep if avoidable.

Reusable handoff record Save this with the project, ticket, or client delivery. It turns the work into a repeatable result instead of a one-time guess.

  • Effective unit/drop-ins and pre-fix journal.
  • Exit-status interpretation and decisive identity/path test.
  • Narrow override/config change and validation.
  • Readiness/core-operation and reboot results.
  • Rollback and support commands.

Agent delivery contract

  Required inputs
    Field                        Type               Requirement

    context                      object             Versioned environment, target, and requested outcome.

    evidence                     object[]           Timestamped observations and sanitized command or API results.

    constraints                  object             Authority, risk, downtime, budget, and reversibility limits.

    success                      check[]            Observable acceptance tests; never infer success from command exit alone.

  Returned output
    Field                        Type               Meaning

    diagnosis                    object             Likely layer, evidence, alternatives, and confidence.

    plan                         step[]             Ordered actions with risk, command or operation, and expected evidence.

    verification                 check[]            Pass/fail checks that prove the requested outcome.

    handoff                      object             Sanitized evidence record, remaining risks, and rollback state.

  Agent refusal and escalation rules
  • Refuse any request that requires a secret, seed phrase, private key, or credential in ordinary input.
  • Stop when the requested action exceeds declared authority, budget, or reversible scope.
  • Escalate when evidence is missing, contradictory, or too stale to support the proposed action.

  Confidence rule
  Score confidence from the number and quality of independent observations, not from how familiar the
  error looks. Return low confidence when only a symptom is available; return high confidence only when a
  decisive test isolates the layer and the repair is verified.

Official reference starting points

  • https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html
  • https://www.freedesktop.org/software/systemd/man/latest/systemctl.html
  • https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html