certslothcertsloth
GH-200/Topic 04

GitHub / Associate

Debugging Runs and Controlling Cost

3 min read5 recall promptsReviewed 2026-10-10

Memory hook: Read the first meaningful failure before rerunning everything.

Must remember

Start with event, ref/commit, actor, workflow version and permissions. A workflow that never started needs trigger/filter/default-branch investigation; a queued job needs runner/concurrency/capacity investigation; a started job needs step logs and exit status. These are different failure layers.

Inspect the first meaningful error, not only the final failed cleanup line. Enable supported debug logging only as needed and review exposure risk. Compare successful and failing runs for runner image, dependency version, input, secret scope and environment differences. ubuntu-latest or windows-latest can move to new images; pin explicit supported versions when stability is required and maintain them deliberately.

Matrix job names identify failing combinations. Rerun a specific failed variant when appropriate, while understanding which workflow/action refs are resolved again. A floating branch/tag may now point to different code; pinning a commit makes provenance clearer. Download logs/artifacts through the UI or authorized API/CLI for focused analysis.

Concurrency groups limit overlapping work. By default, a group permits one running and one pending run/job; a newer pending item replaces the previous pending item even when cancel-in-progress is false. Do not assume this is a durable queue for every deployment. Explicit queuing options have separate limits and compatibility rules. Canceling an old test run differs from interrupting a production migration. Bound matrices, use path filters carefully, parallelize independent work and cache expensive repeat dependencies. Skipping tests only to lower minutes can remove the control that justified deployment.

Disabling a workflow stops future eligible execution without being the same as deleting its file/history. Deleting runs or artifacts removes evidence and may affect retention needs. Measure queue time, run duration, failure/retry rates and storage usage; optimize the actual expensive path rather than assuming the longest visible job is the only cost.

Choose under exam pressure

Requirement Choice and reason
No workflow run exists Inspect trigger, branch/path filters and workflow availability.
Job remains queued Inspect runner labels/groups, availability and concurrency.
Only one OS variant fails Compare image/toolchain and that matrix job’s logs.

Traps

  • Rerunning a floating reference may execute changed code.
  • Cancel-in-progress can be unsafe for noninterruptible deployment operations.

Active recall

1. What separates queued from failed?

Queued work has not obtained execution capacity; failed work ran and encountered an error.

2. Why check image release notes?

Preinstalled tools and OS versions change.

3. Why inspect the earliest useful error?

Later failures may merely be consequences.

4. Disable versus delete?

Stop future workflow execution versus remove configuration or records.

5. What makes a safe optimization?

It reduces wasted work while retaining required correctness and security checks.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.