Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
130 changes: 130 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,130 @@
# Changelog

## Unreleased

### Fixed

- **A seed read that fails now says why.** The backend's message used to go to
`/dev/null`, so `open`, `doctor` and `status` said "no seed — run store-seed" even when
the seed was stored and the keychain had only refused one request. They now repeat the
backend's own reason. Every failed read is appended to `HS_TOTP_ERROR_LOG` (default
`$HS_CONFIG_DIR/<profile>.totp-errors.log`) with its exit status and, for the keychain,
the macOS session it ran in. Only the error is written; the seed never is.
- **A failed read is tried once more** after `HS_TOTP_READ_RETRY_DELAY` seconds (default
2), which absorbs a refusal that clears by itself. Not after a cancelled prompt, which
would ask again someone who just said no, and not for a missing seed file.

## 0.1.0 — 2026-08-21

First tagged release. The repository has been public since late July 2026 and there was
no tag before this one, so anyone who cloned in between holds a copy with the defects
listed under **Fixed** — several of which are silent. Upgrading is worth it.

Two audits, in July and August, produced the work below. Every fix carries a test, and
each test was checked against the specific defect it guards rather than only against the
previous state of `main`. The suite went from 37 assertions to 194.

### Changed — read before upgrading

- **`submit`'s extra arguments are quoted** and no longer re-parsed by the remote shell.
`README` always promised they reached `sbatch` unchanged; they now do. An argument that
relied on remote expansion (a `$VAR`, a backtick) will stop expanding.
- **A remote script path given to `submit` is literal.** `HS_REMOTE_WORKDIR` is still
expanded by the cluster, deliberately — it is where a profile's `$USER` has to resolve.
- **`watch` clamps its interval** to `HS_WATCH_MIN_INTERVAL` (30 s) and stops at
`HS_WATCH_MAX_SECONDS` (one hour), saying so in both cases.
- **`watch` has a third exit status.** `0` finished, `1` unknown, **`3` stopped at the cap
with the job still queued**. Only `0` means the job is done.
- **`queue` prints the whole job id.** The column is ragged now; it used to be truncated,
which turned the array task `12345678_10` into `12345678_1` — a valid id for a
different task.
- **`watch`, `fetch`, `cancel` and `submit` reject anything that is not a job id**, before
it reaches a remote command line.
- **The no-argument subcommands reject trailing arguments.** `open -p bigiron` used to
open the *default* cluster while appearing to select a profile, and spend a code on it.
- **`local` no longer claims to help `sshfs`**, which reads neither `RSYNC_RSH` nor
`GIT_SSH_COMMAND`. Use `-o ssh_command="ssh $(hpc-session ssh-opts)"`.
- **A VPN configured with `HS_VPN_STATUS_CMD` alone is recognised**, which is the shape
`docs/vpn-hooks.md` recommends for a tunnel you raise yourself. It previously reported
as "not configured".
- **The tool's own remote commands run under `sh`**, whatever login shell the account
uses. `hpc-session run` is untouched: that is your command line.

### Added

- **CI.** The suite runs on Linux and on macOS's `/bin/bash` 3.2 — the real floor this
tool is written to — plus `shellcheck`.
- **`hpc-session templates`**, and `render <name>` resolved against `HS_TEMPLATE_DIR`, so
a skill that owns a job type can own its job shape without hardcoding a path into this
repository. `SKILL.md` states the contract across that seam.
- `HS_WATCH_INTERVAL`, `HS_WATCH_MIN_INTERVAL`, `HS_WATCH_MAX_SECONDS`,
`HS_WATCH_MAX_MISSES`, `HS_TEMPLATE_DIR`.
- `SLURM_NTASKS` as a template placeholder. It was hard-coded to 1 while `--nodes` was a
placeholder, so any `HS_SLURM_NODES` above 1 allocated nodes that then sat idle.
- **`hpc-session version`**, also reported by `doctor`. It answers before the profile
loads, so a broken setup can still say which copy it is; a test holds it equal to the
newest heading in this file.
- A second worked example profile, for the ordinary key-only cluster.

### Fixed

- **`watch` reported a running job as finished** whenever a remote query failed with
anything other than ssh's own 255 — an invalid job id, a `squeue` behind a module, a
restarting controller, a `csh` login shell. The output was identical to a real
completion. It now requires positive evidence that the controller was reached.
- **Configuration precedence was inverted.** The profile was sourced last, so every
documented one-off override (`HS_HOST=other hpc-session status`) was silently discarded.
- **The shipped `HS_REMOTE_WORKDIR` broke `submit` on every OpenSSH ≥ 9.0 client**, which
speaks SFTP and runs no remote shell to expand `$USER`.
- **`fetch` could deliver an unrelated file** from the remote home: `ls -1` on a matching
*directory* prints its contents as bare relative names. It also skipped a workdir whose
last component was a symlink, reporting "nothing matching" — which reads as "the job
wrote nothing".
- **`fetch` returned 0 on a partial retrieval.** Its status was the last copy's.
- **`init` could not create a profile named with `-p`**, the only selector `README`
documents, and the error's own suggested remedy failed when `HS_PROFILE` was exported.
- **The `command` TOTP backend could not open a session at all** — it was asked for a
stored seed it has none of by design.
- **A bad key under `AuthenticationMethods publickey,keyboard-interactive`** was retried
three times over ~90 s instead of failing at once.
- **A local code-generation failure waited out three full time steps** for something that
could never change, judging a stale error string.
- **A failed control-directory `mkdir` was swallowed**, so the session opened and simply
never multiplexed — which looks like a slow cluster, not a broken setup.
- **A mistyped `HS_TOTP_ALGO` or `HS_TOTP_PERIOD` was reported as an invalid seed**,
sending the user to re-enrol over a typo.
- **`--help` printed executable code**, having sliced a line range that had drifted.
- `sbatch`'s federated message form made `submit` return the **cluster name** as the job
id; `--parsable` made it fail *after* the job was queued.
- `config.example` omitted `HS_CONTROL_DIR`, `HS_CONFIG_DIR` and `HS_OTP`. A test now
holds every default to appearing there.
- Four documentation claims the code did not honour, including two in `SECURITY.md`.

### Security

- **The TOTP seed no longer passes through `argv`** on the keychain path, where `ps` and
any exec-logging agent could see it. It travels on stdin.
- **A live code no longer survives a signal.** A trap removes the temporary file on
`EXIT`, `INT`, `TERM` and `HUP`; previously a Ctrl-C between writing it and `ssh`
reading it left a valid second factor on disk with nothing left to delete it.
- **The keychain identifiers are refused if they contain a quote, a backslash or a
newline.** `security -i` re-tokenises the storing command and runs one command per line,
so either could inject options — or a whole second command — into something running
against the user's keychain.
- **The `file` backend enforces mode 0600** on an existing file. `umask` governs creation
only, so a seed written into a file that was already 0644 stayed world-readable.
- `SECURITY.md` now records what `-T /usr/bin/security` costs: any process running as you
can read the seed back without a prompt, which is what makes unattended `open` work.

### Known limitations

- `README` states OpenSSH 6.7+, which is correct for multiplexing. The unattended TOTP
path additionally needs `SSH_ASKPASS_REQUIRE`, which is newer
([#13](https://github.com/HolobiomicsLab/hpc-session/issues/13)). A caller with a
terminal is unaffected.
- Multi-cluster (federated) submission is out of scope. `submit` says so when `sbatch`
reports another cluster.
- The test suite is entirely offline, against stubs. Separately, one end-to-end run was
made against a live SLURM controller (a TOTP-gated cluster behind a VPN): `open`,
`render`, `submit`, `queue`, `watch` to completion, `fetch`, `cancel`, `close`, all as
documented. That is one site and one SLURM version, not a compatibility claim.
11 changes: 9 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -285,15 +285,22 @@ in either skill.
## Contributing

Issues and pull requests are welcome, particularly worked profiles for other clusters
(as `examples/<site>.conf`, with placeholders instead of real logins) and VPN hooks for
clients not yet covered. Please keep `tests/run_tests.sh` green and add a case for any
(as `examples/<site>.conf`, with placeholders instead of real logins — see
[`examples/`](examples/) for the two shapes already there) and VPN hooks for clients not
yet covered. Please keep `tests/run_tests.sh` green and add a case for any
behaviour you change — CI runs it on Linux and on macOS's bash 3.2, plus `shellcheck`.

## Authors

Developed at the [HolobiomicsLab](https://github.com/HolobiomicsLab) — CNRS and
Université Côte d'Azur.

## Changelog

[CHANGELOG.md](CHANGELOG.md). Behaviour changes worth knowing before upgrading are listed
first in each release. `hpc-session version` says which copy you have, and `doctor` opens
with the same line — quote it when reporting anything.

## License

MIT — see [LICENSE](LICENSE). Copyright CNRS and Université Côte d'Azur.
4 changes: 4 additions & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,10 @@ TOTP seed if you configure a backend that stores one.
code (in which case nothing is stored here).
- **The seed never appears in `argv` or the environment.** It is piped into the code
generator on stdin, so `ps` cannot see it.
- **A failed seed read is logged; the seed never is.** The backend's error message, its
exit status and the time go to `HS_TOTP_ERROR_LOG`, created mode 0600 beside your
profiles. A successful read hands the seed on with a shell builtin, so it still reaches
no `argv`, no environment and no file.
- **The generated code** is written to a `mktemp` file (mode 0600) and read exactly once:
the askpass helper prints it and deletes it. That single-shot behaviour is deliberate —
a helper that kept answering would let `ssh` retry a stale code in a loop.
Expand Down
3 changes: 3 additions & 0 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,10 +20,13 @@ of `hpc-session`; see [README.md](README.md) for the full reference and
## Before anything else

```bash
hpc-session version # which release this is — behaviour differs between them
hpc-session doctor # no network — reports what is configured and what is missing
hpc-session status # is the master up? the VPN? is a TOTP seed present?
```

`version` answers even when the profile is missing or broken, so it is safe to ask first.

If no profile exists, run `hpc-session init` — or `hpc-session -p <name> init` for a second
cluster — and tell the user which keys they must fill in (`HS_HOST` at minimum). Do not
invent a hostname, account or partition — ask.
Expand Down
11 changes: 11 additions & 0 deletions bin/hpc-session
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@
# hpc-session watch 12345 # poll until it leaves the queue
# hpc-session fetch 12345 ./results # bring the job's files back
# hpc-session close # drop master + VPN, freeing the link
# hpc-session version # which release this copy is
#
# Setup: hpc-session init [profile] -> edit the profile -> hpc-session doctor
# Docs: README.md, docs/2fa-enrollment.md, docs/vpn-hooks.md
Expand Down Expand Up @@ -82,6 +83,7 @@ hs_doctor_vpn() {

# Everything checkable without touching the network — safe to run before a first login.
hs_doctor() {
echo "hpc-session $HS_VERSION"
echo "profile $HS_PROFILE ($(hs_profile_path "$HS_PROFILE"))"
[ -n "$HS_HOST" ] && hs_report host ok "$HS_HOST" || hs_report host fail "HS_HOST is unset"
hs_doctor_tool ssh "the whole tool depends on it"
Expand Down Expand Up @@ -122,6 +124,15 @@ if [ "${1:-}" = init ]; then
exit
fi

# Dispatched before the loader, like `init`: a missing or broken profile is exactly when
# someone needs to say which copy of the tool they are running, and the answer does not
# depend on a profile. CHANGELOG.md is the other half of this string; a test holds them
# together so a tag cannot drift from what the tool reports.
HS_VERSION="0.1.0"
case "${1:-}" in
version|--version) printf 'hpc-session %s\n' "$HS_VERSION"; exit 0 ;;
esac

hs_load_profile

case "${1:-}" in
Expand Down
6 changes: 6 additions & 0 deletions config.example
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,12 @@ HS_TOTP_DIGITS="6"
HS_TOTP_PERIOD="30"
HS_TOTP_ALGO="sha1"

# A seed read that fails is tried once more after this many seconds, and the backend's
# own message — never the seed — is appended to the log. Empty means the default,
# $HS_CONFIG_DIR/<profile>.totp-errors.log. open, doctor and status repeat the latest one.
HS_TOTP_READ_RETRY_DELAY="2"
HS_TOTP_ERROR_LOG=""

# --- connection tuning --------------------------------------------------------------
HS_CONTROL_PERSIST="8h" # how long an idle master survives
HS_CONNECT_TIMEOUT="25" # seconds
Expand Down
1 change: 1 addition & 0 deletions docs/2fa-enrollment.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,6 +106,7 @@ stall for up to 30 seconds waiting for a fresh step.
|---|---|---|
| `auth failed … waiting Ns for a fresh code` | the previous code was already consumed | it retries by itself; wait one step |
| `could not open the master`, repeatedly | VPN down, or wrong seed | `hpc-session status`; compare `hpc-session code` with the phone |
| `the seed in 'keychain' could not be read: …` | the keychain refused this read, often because it wanted to show a prompt and the command ran outside your logged-in session | the message names the reason; `HS_TOTP_ERROR_LOG` keeps every failed read with the session it ran in. Run `doctor` from your own terminal to compare |
| Commands hang instead of failing | stale socket after the tunnel dropped | `hpc-session close` then `open` (it also cleans up on its own) |
| Every code rejected | clock skew — TOTP is time-based | check the machine's clock is NTP-synced |
| Codes rejected only sometimes | your site may not use the default profile | set `HS_TOTP_DIGITS` / `HS_TOTP_PERIOD` / `HS_TOTP_ALGO` |
Expand Down
37 changes: 37 additions & 0 deletions examples/keyonly-direct.conf
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# Example profile: the ordinary case — a cluster you reach directly, with an SSH key and
# no second factor. Start here; the other example adds a VPN and TOTP on top.
#
# Copy to ~/.config/hpc-session/<name>.conf, replace every PLACEHOLDER, then:
# hpc-session -p <name> doctor
#
# Contains no credentials, and should never contain any.

# Requires a matching block in ~/.ssh/config, which is where your key, user and any jump
# host belong — every other tool can see them there too:
# Host mycluster
# HostName login.cluster.example.edu
# User YOUR_CLUSTER_LOGIN
HS_HOST="mycluster"

# Job scripts are copied here and submitted from here. Escape $USER so the REMOTE shell
# expands it — since OpenSSH 9.0 scp speaks SFTP and runs no remote shell, so an
# unexpanded one would reach the copy verbatim. Prefer scratch over your home directory.
HS_REMOTE_WORKDIR="/scratch/\$USER/jobs"

# Defaults for `hpc-session render`. Leave a value empty and its #SBATCH line is dropped,
# so an unset account does not become an `--account=` that SLURM rejects.
HS_SLURM_ACCOUNT=""
HS_SLURM_PARTITION="YOUR_PARTITION"
HS_SLURM_TIME="02:00:00"
HS_SLURM_CPUS="4"
HS_SLURM_MEM="16G"

# No VPN and no second factor: leave every hook empty and the tool behaves as if neither
# existed. `open` is then a plain key login, and `doctor` needs no python3.
HS_VPN_UP_CMD=""
HS_VPN_STATUS_CMD=""
HS_TOTP_BACKEND="none"

# Nothing here monopolises your link, so a long-lived master is pure win: one login covers
# a working day. Shorten it if your site drops idle connections.
HS_CONTROL_PERSIST="8h"
4 changes: 4 additions & 0 deletions lib/config.sh
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,10 @@ hs_apply_defaults() {
: "${HS_TOTP_DIGITS:=6}"
: "${HS_TOTP_PERIOD:=30}"
: "${HS_TOTP_ALGO:=sha1}"
# A failed seed read is tried once more after this many seconds, and the backend's own
# message — never the seed — goes to this log. See hs_seed_read.
: "${HS_TOTP_READ_RETRY_DELAY:=2}"
: "${HS_TOTP_ERROR_LOG:=$HS_CONFIG_DIR/$HS_PROFILE.totp-errors.log}"
: "${HS_REMOTE_WORKDIR:=.}"
# Where `render <name>` looks. Point it at another skill's assets and that skill owns its
# own job shapes without owning any paths — see "Composing with other skills" in SKILL.md.
Expand Down
Loading
Loading