Posts tagged with “infra”

How a Baton test killed every tmux session with `kill(-1)`

A desktop restart can look like an XRDP or systemd problem. In this incident, the real culprit was a Rust test for Baton: it accidentally broadcast SIGTERM to every process owned by the test user.

The audit record showed:

proctitle=kill -TERM -1072950
syscall=kill ... a0=0xffffffff a1=SIGTERM
ppid=1072584 ... comm=kill exe=/usr/bin/kill

The parent process was named baton-5b9d31243. That is a cargo test binary under target/debug/deps, not a resident Baton service. The test was exercising max-duration handling with a sleep 5 child, so Baton really was only active because an agent was running tests on the Baton project.

The code intended to signal a process group:

kill -TERM -1234

On Linux, procps parses the negative group ID as another option unless the option list is terminated. The resulting syscall is kill(-1, SIGTERM), which signals every process the caller is allowed to signal. That killed the desktop session, kitty, browsers, the systemd user manager, and tmux sessions. mat-keeper could not protect itself because it ran as the same user.

The safe form is:

kill -TERM -- -1234

The fix was applied to both the Rust Baton implementation and my-ai-team's shell Baton transport, including test cleanup paths. The RDP connection switch was a coincidence in the timeline; it did not directly invoke Baton. Full Baton tests and the relevant my-ai-team regressions pass with the corrected argument handling.

XRDP session teardown can kill your default tmux server

An XRDP/Xfce connection is tied to a desktop and user-manager lifecycle, not only to the RDP window. If that session tears down, the default tmux server, kitty processes, and user services can disappear together without anyone running tmux kill-server.

On XPS, xrdp-sesman logged the window manager exiting with signal 15, then cleaned up the X server. XFCE reported broken PipeWire, D-Bus, and ICE connections; systemd subsequently stopped the user manager and kitty scopes. A simultaneous NVIDIA “fallen off the bus” failure was a strong suspect, but the old logs could not identify the original signal sender.

Keep long-lived agents on an explicit tmux socket supervised outside the desktop session:

sudo systemctl enable --now [email protected]
tmux -L mat list-sessions
tmux -L mat attach -t <session>

The mat-keeper session is only the service sentinel. MAT agent sessions share the mat server and should survive an XRDP switch. The default tmux a command still refers to the desktop-bound default socket, so use -L mat when recovering the fleet.

To make the next teardown attributable, enable auditd with a persistent signal rule and inspect the caller:

sudo ausearch -k process-signal -ts recent -i
sudo ausearch -k service-control -ts recent -i

Look for comm, exe, pid, ppid, and the target in OBJ_PID; these distinguish xrdp-sesman, systemd, and other senders. Auditd cannot reconstruct an earlier incident, and a desktop process can also exit because of D-Bus or GPU failure without a signal, so correlate its records with /var/log/xrdp-sesman.log, .xsession-errors, and journalctl.

`printf: write error: Broken pipe` on a self-hosted GitHub runner? Check SIGPIPE

A bash pipeline that is quiet locally but fails only on one self-hosted runner may be a signal-disposition problem, not a data or code problem. Our required test failed in unrelated files with printf: write error: Broken pipe, and every failure landed on the same Linux runner.

The useful check is the effective job-shell behavior, not only the runner listener:

systemctl show actions.runner.SHUKE-LABS.<runner>.service -p IgnoreSIGPIPE

PATH=/usr/local/libexec/actions-runner:$PATH bash -c '
  set -o pipefail
  printf "%s\n" {1..10000} | head -n1 >/dev/null
  printf "pipeline rc=%s\n" "$?"
'

The normal result is a SIGPIPE exit (usually 141). With SIGPIPE inherited as SIG_IGN, bash's builtin printf receives EPIPE instead, prints the broken-pipe diagnostic, and pipefail turns it into a failure.

The fix has two parts. Add a systemd drop-in:

[Service]
IgnoreSIGPIPE=no

Then put a root-owned launcher first in the runner's .path file:

#!/usr/bin/perl
$SIG{PIPE} = "DEFAULT";
exec "/usr/bin/bash", @ARGV or die "exec bash: $!";

The launcher matters because the Actions Runner's Node/Listener/Worker chain can set SIGPIPE to ignored again after systemd starts the service. Verify with a fresh job shell and the pipeline test; the listener process's SigIgn alone is not a reliable success criterion.

Check the whole Linux fleet, not only the runner named in the first failure. The same defect was present on all three sydney2 runners and on the mbp15-vm Lima guest. After fixing the signal boundary, two independent failures surfaced: a jq 1.6-incompatible test mock and newer-main tmux/test timing regressions. Keeping those layers separate prevented a correct runner diagnosis from being mistaken for a complete CI diagnosis.

The final CI-shaped local run passed 166/166, and the GitHub required jobs passed after the fleet fix.

Routing Claude Code through a gateway that's behind HTTP Basic Auth

Say you expose an Anthropic-compatible gateway on the public internet and put nginx Basic Auth in front of it. Point Claude Code at it and every request 401s — or worse, the gateway rejects it with a cryptic error. The reason is a header collision, and it's easy to untangle once you see it.

Basic Auth lives in the Authorization header (Authorization: Basic <base64>). Claude Code's own auth also wants a header — but which one depends on how you authenticate:

  • ANTHROPIC_API_KEY → sent as X-Api-Key
  • ANTHROPIC_AUTH_TOKEN (and subscription/OAuth login) → sent as Authorization: Bearer <token>

So the trick is to stay in API-key mode. Then Authorization is free for Basic Auth and your key rides in the separate X-Api-Key header — no collision:

export ANTHROPIC_BASE_URL="https://gateway.example.com"
export ANTHROPIC_CUSTOM_HEADERS="Authorization: Basic $(echo -n 'user:pass' | base64)"
export ANTHROPIC_API_KEY="<your-key>"
unset ANTHROPIC_AUTH_TOKEN   # Bearer would fight Basic for the Authorization header

On the nginx side, strip the client's Authorization after Basic Auth passes, so the Basic credential never leaks upstream to the gateway:

location / {
    auth_basic           "gateway";
    auth_basic_user_file /etc/nginx/.htpasswd;

    proxy_pass http://backend:9949;
    proxy_set_header Authorization "";   # consumed by auth_basic; don't forward
}

Two gotchas worth knowing. If you're logged into Claude Code with a subscription, it sends Authorization: Bearer <oauth> and ignores ANTHROPIC_API_KEY — so it collides with Basic and never sends X-Api-Key. Log out (or use a config dir with no login) to force API-key mode. And if your gateway reads a pool/backend selector from X-Api-Key (some proxies do), the value must be the selector name, not a real sk-ant-... key — a real key or a dummy matches no pool and fails closed.

The failure mode that eats the most time isn't any of the above, though: a typo in the variable name. ANTROPIC_API_KEY (missing the H) sets nothing, Claude Code silently falls back to whatever else it can find, and you'll chase the server for an hour before spotting the letter. Echo env | grep -i anthropic before blaming the proxy.

Moving a Guacamole (or any stateful Docker stack) to a new host? Copy the volume, don't re-init

When you move a Guacamole stack between machines, the temptation is to spin up a fresh stack on the new host and let initdb.sql build the database. Don't — that gives you an empty install. Every saved connection, every user, and (critically) every TOTP/MFA enrollment lives inside the MariaDB data volume. Re-running the init script wipes all of it, and your users have to re-scan their authenticator QR codes.

The fix is a cold, byte-exact copy of the DB volume. Stop the stack first so the copy is consistent:

# on the OLD host
cd ~/path/to/guacamole && docker compose down
docker run --rm -v guacamole_db-data:/v -v /tmp:/out alpine \
  tar czf /out/guacdb.tgz --numeric-owner -C /v .

--numeric-owner matters: MariaDB's files are owned by uid 999 inside the container, and you want that uid preserved, not remapped to whatever user happens to exist on the new box.

Ship the tarball over, then restore it into a fresh named volume before the first up:

# on the NEW host
docker volume create guacamole_db-data
docker run --rm -v guacamole_db-data:/v -v /tmp:/in alpine \
  tar xzf /in/guacdb.tgz --numeric-owner -C /v
cd ~/path/to/guacamole && docker compose up -d

MariaDB's entrypoint checks whether the data directory is empty. Since you just populated it, it skips initialization entirely and comes up with all your data intact — confirm with docker logs guacamoledb | grep "ready for connections" and no Initializing database line.

One more thing worth doing while you're at it: if a reverse proxy fronts the app (e.g. nginx terminating TLS on a separate edge box), point its proxy_pass at a DNS name that tracks the new host's IP rather than the raw IP. Then a future move is a one-line edge change — or zero, if the name already follows the host. A quick sanity check that the whole path works, without needing to log in:

curl -s -o /dev/null -w "%{http_code}\n" -X POST https://your.guac.example/api/tokens \
  -H "Content-Type: application/x-www-form-urlencoded" --data "username=x&password=y"

A 403 here is success — it means the webapp reached the database and rejected bad credentials. (A 500 usually just means you forgot the application/x-www-form-urlencoded content type, not that anything's broken.)