From adcdd2edcd54b80a44bfaa657b403f8c80e3b8cf Mon Sep 17 00:00:00 2001 From: zhangwenjian Date: Mon, 7 Sep 2026 13:32:43 +0800 Subject: [PATCH 1/2] =?UTF-8?q?docs=F0=9F=93=9D:=20name=20the=20job=20sche?= =?UTF-8?q?duler=20as=20the=20other=20reason=20for=20one=20replica?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The comment on replicas gave one obstacle to raising it, the shared log volume, which reads as the only one. Someone who moves the log path off that volume would conclude the way is clear. The scheduler in app/jobs is the second, and it is the one that does not announce itself. Its handle on a job lives in sys_job.entry_id, one column shared by every process, and startup zeroes the whole column before writing its own ids. A second pod therefore erases the first pod's, and both pods run the full enabled list. Stopping a job from the UI then removes an entry from whichever process is asked, by an id that belongs to another one, and answers 200. See #915. --- scripts/k8s/deploy.yml | 18 ++++++++++++++---- 1 file changed, 14 insertions(+), 4 deletions(-) diff --git a/scripts/k8s/deploy.yml b/scripts/k8s/deploy.yml index cac85e8e..e9ad5e24 100644 --- a/scripts/k8s/deploy.yml +++ b/scripts/k8s/deploy.yml @@ -23,10 +23,20 @@ metadata: version: v1 spec: # One replica, and the drain window below buys nothing at one replica: there - # is nowhere to send the traffic this pod stops taking. Raising it needs one - # more change than the number - the volume below is shared by every replica, - # and the log path in settings.yml lives on it, so a second pod would append - # to the same rotating file. + # is nowhere to send the traffic this pod stops taking. Raising it needs two + # changes that are not this number: + # + # The volume below is shared by every replica, and the log path in + # settings.yml lives on it, so a second pod would append to the same + # rotating file. + # + # The job scheduler is per process while its handle on a job is one shared + # column. Startup runs `UPDATE sys_job SET entry_id = 0 WHERE entry_id > 0` + # across the whole table (app/jobs/jobbase.go), so a second pod erases the + # first pod's ids and writes its own, and every pod registers the whole + # enabled list in its own scheduler. Neither symptom logs anything: an + # enabled job fires once per pod, and stopping one from the UI removes an + # entry from the wrong process and still answers 200. See #915. replicas: 1 selector: matchLabels: From 63bcc912ef22c18be5aeccbf2b43dd2e82d54825 Mon Sep 17 00:00:00 2001 From: zhangwenjian Date: Mon, 7 Sep 2026 14:14:39 +0800 Subject: [PATCH 2/2] =?UTF-8?q?ci=F0=9F=91=B7:=20stop=20the=20k8s=20manife?= =?UTF-8?q?sts=20from=20redeploying=20the=20demo?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The comment at the top of this workflow says documentation-only changes skip it, because a push to master pushes an image, runs the migrations and restarts the demo container. The ignore list did not cover scripts/k8s, so editing a manifest that the deploy never reads bought the site an outage. The pattern is scripts/k8s/** rather than scripts/** because scripts/Dockerfile is a build input - go.yml builds the release image from it on a tag. This workflow file stays outside the list on purpose. paths-ignore skips only when every changed path matches, so a change that edits the deploy still runs it, which is the point. --- .github/workflows/build.yml | 17 ++++++++++++++++- 1 file changed, 16 insertions(+), 1 deletion(-) diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index d3e3578e..94cd4cea 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -1,12 +1,25 @@ name: Build -# Documentation-only changes skip this workflow entirely. +# Documentation-only changes, and changes confined to the Kubernetes +# manifests, skip this workflow entirely. # # A push to master here does not just build - it pushes an image, runs the # migrations and restarts the demo container, so the site takes a short outage. # Paying that for a README edit is waste at best; at worst a deploy fails for a # reason unrelated to anything in the change. Code coverage is unaffected, # because go.yml still builds every push and pull request. +# +# scripts/k8s holds deploy.yml, storage.yml and prerun.sh, and the deploy below +# reads none of them - it is an ssh into one host that runs docker, building the +# Dockerfile at the repository root. Those manifests are for people deploying to +# a cluster of their own. The pattern is scripts/k8s/** rather than scripts/** +# because scripts/Dockerfile is a build input: go.yml builds the release image +# from it on a tag. +# +# A file outside these patterns still runs the workflow even when the rest of +# the change is ignorable: paths-ignore skips only when every changed path +# matches. Editing this file is one such case, on purpose - a deploy script +# that is never exercised by the change that broke it is worse than an outage. on: push: branches: [ master ] @@ -15,6 +28,7 @@ on: - 'docs/**' - 'LICENSE*' - '.github/ISSUE_TEMPLATE/**' + - 'scripts/k8s/**' pull_request: branches: [ master ] paths-ignore: @@ -22,6 +36,7 @@ on: - 'docs/**' - 'LICENSE*' - '.github/ISSUE_TEMPLATE/**' + - 'scripts/k8s/**' # One deploy at a time. Two merges seconds apart raced here: both runs did # docker rm -f then docker run, the second removed the container the first had