A Telegram notifier for Mimir, the AI-settled prediction market on Stellar. It polls Mimir's two Soroban contracts for new on-chain events and posts them, human-readable, into one or more named chats or channels:
🆕 New claim #7
Category: crypto
Creator: GBMGZ…IR2Y
ledger 4226691 · tx
⚔️ Claim #7 challenged
Stake: 2.0000000 USDC
Challenger: GDZCB…X4UH
ledger 4226692 · tx
⚖️ Claim #7 resolved — winner: challengers
Confidence: 100%
Onchain smoke — challengers awarded so the payout pull can be exercised
ledger 4226728 · tx
Built with grammy and
@stellar/stellar-sdk. Reads only —
it holds no keys and signs nothing.
| Contract | Events it notifies on |
|---|---|
mimir-market |
claim_created, claim_challenged, claim_resolved, claim_cancelled, market_settled, challenger_paid, fee_claimed, withdrawal, withdrawal_pending |
mimir-squad |
market_created, deposited, withdrawn, resolved, claimed, fees_claimed |
Admin events (oracle_changed, ownership_transferred, fee_policy_*,
fee_accrued, agent_attributed) are decoded far enough to be recognised and
then skipped — they are retained as bounded unknown events and logged with
their event name, ledger, and bounded decode reason, not posted. New event names
and malformed XDR follow the same path, so they cannot stop polling or advance a
log with an unbounded remote payload.
The notifier can be restricted to a subset of claim categories by setting an
optional allowlist in .env:
NOTIFY_CATEGORIES=crypto,politicsWhen set, only claim_created events whose category matches an allowlisted name
are sent to Telegram; all other notification types continue to flow normally.
The default (unset or empty) preserves the current behaviour and notifies every
claim category.
Message @BotFather on Telegram, send /newbot, follow
the prompts, and copy the token it gives you (123456789:AA…).
- Private chat: message @userinfobot; it replies with your numeric id.
- Group: add your bot to the group, send any message, then open
https://api.telegram.org/bot<YOUR_TOKEN>/getUpdatesand readresult[].message.chat.id. Group and supergroup ids are negative (-1001234567890). - Channel: add the bot as an administrator with "Post messages" permission.
Either use the numeric id from
getUpdatesor, for a public channel, the@channelusername.
If your group has privacy mode
on (the default), the bot only sees messages that are commands or replies to it —
which covers /status and the operator controls below.
Set OPERATOR_TELEGRAM_USER_ID to the numeric user id returned by
@userinfobot to enable /audit, /pause, and /resume. The notification
TELEGRAM_CHAT_ID is intentionally not accepted as authorization: in a group,
everyone can send messages from that chat. If this variable is omitted, existing
deployments continue unchanged and the operator commands are ignored.
cp .env.example .env # then fill in BOT_TOKEN and TELEGRAM_CHAT_ID
npm install
npm run dev # tsx, restarts on changeFor production (Node):
npm run build
npm startFor production (Docker):
docker build -t mimir-telegram-bot .
docker run -d \
--name mimir-bot \
--env-file .env \
-v $(pwd)/data:/app/data \
mimir-telegram-bot.env.example ships with the live Stellar Testnet contract ids, so the only two
values you must supply are BOT_TOKEN and TELEGRAM_CHAT_ID. To split traffic,
set TELEGRAM_MARKET_CHAT_ID and/or TELEGRAM_SQUAD_CHAT_ID; each overrides the
legacy destination for that contract, while an omitted override falls back to
TELEGRAM_CHAT_ID. Every other
variable is documented inline there. A missing or malformed value aborts startup
with all the problems listed at once — the bot never boots into a state where it
looks healthy but notifies nobody.
| Command | What it does |
|---|---|
/start |
What the bot is |
/help |
Same, plus the command list |
/status |
Chain tip, the RPC's retained-history floor, the chain clock skew (newest chain close time the bot has seen, against its own clock), both watched contract ids, the last ledger an event was seen in per contract, the persisted cursor, poll/send/skip/dedup counters (plus messages dropped by a shutdown drain and cursors automatically rewound to the retained floor), and the last error |
/audit |
Operator only. The operator audit report: recent scan failures, send failures, skipped and cap-dropped events, cursor problems — redacted and bounded (see Operator audit trail) |
/contracts |
The two contract ids this bot watches (mimir-market, mimir-squad) and a stellar.expert link for each. Reads only from config, so it answers the same during a cold start, a run of RPC failures, or between restarts — unlike /status, there is nothing here that can be "unhealthy" |
/preview |
Previews channel notification formatting for mimir-market or mimir-squad without affecting cursors or poller state |
/pause |
Operator only. Stops scheduling new poll cycles; a scan already in progress may finish and persist its normal cursor |
/resume |
Operator only. Schedules the next poll cycle immediately, without changing or replaying cursors |
npm run cursor -- backup |
Backs up the validated cursor file without contacting Stellar or Telegram |
npm run cursor -- restore --from PATH |
Restores a validated backup while holding the poller lock; replacing an existing cursor requires --force |
/start, /help, and Telegram's command menu are generated from the same
command metadata used to register handlers. Without OPERATOR_TELEGRAM_USER_ID,
the help and menu omit /audit, /pause, and /resume; configured operator
commands are labelled "Operator only". Help uses static metadata, so it remains
available when the poller or RPC is unhealthy. Adding a command requires its
metadata and handler in src/bot.ts; TypeScript checks that every registry entry
has a handler. This changes no cursor files or deployment settings.
Commands from a user other than OPERATOR_TELEGRAM_USER_ID receive no control
response and cannot mutate poller state — this includes /audit, whose report
is operator-only. Repeated /pause or /resume commands
are idempotent. Control state is process-local: a restart resumes polling and
loads the existing version-1 cursor file.
/status is for a human in the chat. For a supervisor, a dashboard, or a shell
on the box, the poller also writes the same facts as JSON to STATUS_FILE
(default data/status.json) after every cycle, and on start and stop:
npm start -- --status # or: node dist/index.js --status{
"schemaVersion": 1,
"generatedAt": "2026-08-21T10:00:00.000Z",
"uptimeMs": 3600000,
"running": true,
"network": "testnet",
"rpcUrl": "https://soroban-testnet.stellar.org",
"chatId": "-…7890",
"pollIntervalMs": 30000,
"maxNotificationsPerCycle": 20,
"cycles": 120,
"lastPollAt": 1755770400000,
"lastSuccessAt": 1755770400000,
"latestLedger": 4226733,
"oldestLedger": 4105773,
"notificationsSent": 11,
"notificationsFailed": 0,
"eventsSkipped": 3,
"eventsDeduplicated": 0,
"cursorRewinds": 0,
"consecutiveFailures": 0,
"lastError": null,
"targets": [
{
"source": "market",
"contractId": "CDV6JXIJCALSXQELCS6YUEWJWG5DFXQK5PJ5I7MWI6KVMQJBC5DLPKZI",
"cursor": "0018276211125911551-4294967295",
"lastEventLedger": 4226729,
"rewindFromLedger": null,
"lastError": null
}
]
}--status reads the file only — it never contacts Telegram or the RPC — so it is
safe to run from a health check or a cron job while the bot is running. It exits
0 when a snapshot was read and 1 when there is none or it is not valid JSON.
The offline cursor command uses CURSOR_FILE from the environment or .env
(default data/cursor.json). Backups default to a timestamped sibling file;
choose another persistent location with --out. A backup is a point-in-time,
byte-preserving copy of a supported current or legacy cursor file and does not
pause polling. It is written atomically and is never overwritten unless
--force is supplied. Run npm run build before using the command in a local
checkout; the Docker image already includes the compiled CLI.
npm run cursor -- backup
npm run cursor -- backup --out /safe-storage/cursor-before-deploy.json
npm run cursor -- restore --from /safe-storage/cursor-before-deploy.json --forceRestore validates and normalizes the backup to the current version-1 cursor
format, takes INSTANCE_LOCK_FILE (default data/poller.lock) to prevent the
bot from starting or writing concurrently, and atomically replaces the live
file. Stop the bot before restoring. A live process holding the lock makes the
command fail; stale locks are handled by the same PID check used at startup.
Restoring an existing cursor requires --force; malformed or future-version
files are always rejected. The command prints only paths, schema version, and
target count, never opaque cursor contents.
Keep backups on persistent storage separate from the deployment's writable
cursor when possible. On Railway, a one-off command must use the same attached
/app/data volume (or an explicitly mounted backup location). The runtime
Docker image contains the CLI through the normal TypeScript build. A rollback
to an older bot release remains compatible with the version-1 cursor format;
preserve the backup until the restored release has resumed polling. The chain
remains the source of truth: a backup restores reader position and dedup state,
not missed notifications or on-chain events.
What is deliberately not in it. The snapshot is built from an allowlist of
fields, so nothing can leak by accident. It never contains the bot token, a
private key, or a payment proof. The chat id is redacted to its sign and last
four digits (-…7890), and every string that comes from outside the process —
RPC errors, Telegram errors, cursors — is whitespace-collapsed and truncated
(MAX_ERROR_CHARS, 300) so a hostile or chatty endpoint cannot write an
unbounded blob into the file or into a log line. The write is atomic
(write-then-rename), so a reader never sees a half-written document.
Reading it in a health check. running: false means the process stopped
deliberately (SIGINT/SIGTERM) or has not started; consecutiveFailures > 0 with
a fresh lastPollAt means the RPC is failing but the loop is alive; a
generatedAt that stops advancing means the process is wedged or gone. The
chain remains the source of truth — this file reports on the reader, it is not a
substitute for reading the chain.
The chain reader runs standalone. Testnet's Soroban RPC is public and unauthenticated, so this needs nothing but the contract ids:
npm run scan # both contracts, from the RPC's retained floor
npm run scan -- --pages 40 # walk further
npm run scan -- --show 20 # print 20 decoded events per contract
npm run scan -- --from 4226500 # explicit start ledger
npm run scan -- --contract market # scan only the market contract
npm run scan -- --contract squad # scan only the squad contract
npm run scan -- --help # show usage
npm run scan # both contracts, from the RPC's retained floor
npm run scan -- --pages 40 # walk further
npm run scan -- --show 20 # print 20 decoded events per contract
npm run scan -- --from 4226500 # explicit start ledger
npm run scan -- --json # one mimir-scan-v1 JSON document on stdout
npm run scan -- --json --show 20 # JSON including 20 decoded events per contractHuman mode prints the ledger window, an event-name histogram, and the decoded
payloads. With --json, stdout is a single mimir-scan-v1 document (bigints as
decimal strings) and progress goes to stderr, so npm run scan -- --json | jq
stays valid. Each target reports its startLedger and startClamped, so it is
clear when a requested --from was moved up to the retained floor. Neither mode
prints bot tokens or signing keys — the scanner never holds them. This is how the
decoder was verified against the live deployment.
Replay reads a fixed ledger range from the chain and optionally re-posts the events to Telegram. It is one-shot: it exits when the range is exhausted and never writes a cursor file — the live poller's cursor state is untouched.
npm run replay -- --from 4226500 # dry-run: decode only, no send
npm run replay -- --from 4226500 --to 4226800 # bounded range
npm run replay -- --from 4226500 --send # send to Telegram (needs BOT_TOKEN)
npm run replay -- --from 4226500 --contract market # one contract only
npm run replay -- --from 4226500 --json # machine-readable mimir-replay-v1
npm run replay -- --from 4226500 --pages 5 # walk at most 5 pages per contract
npm run replay -- --from 4226500 --cap 10 # cap at 10 notifications per contract
npm run replay:mock -- --from 4226500 # local mock profile, no credentialsFlags:
| Flag | Default | Meaning |
|---|---|---|
--from <ledger|cursor> |
(required) | Start of the range — a ledger number or an opaque RPC cursor |
--to <ledger> |
chain tip | End of the range (inclusive). Clamped to the tip when above it |
--send |
off | Actually post to Telegram; requires BOT_TOKEN and TELEGRAM_CHAT_ID |
--contract market|squad |
both | Scan only the named contract |
--pages <n> |
20 | Page budget per contract |
--cap <n> |
MAX_NOTIFICATIONS_PER_CYCLE |
Maximum notifications per contract per run |
--show <n> |
0 | Include the last n decoded events per target in the report |
--json |
off | Machine-readable mimir-replay-v1 JSON on stdout; progress on stderr |
--mock |
off | MIMIR_PROFILE=mock: local RPC, fixture contracts, no credentials needed |
Default mode is dry-run. Events are decoded and counted; nothing is posted to
Telegram. Add --send to deliver notifications. The run always exits with code
0 on completion, 1 on a fatal RPC or config error, and 2 on a bad flag.
Cursor clamping. --from below the RPC's retained floor is moved up to the
floor with a warning. --to above the chain tip is clamped to the tip. A --to
before --from is a usage error (exit 2). Neither clamp changes the live
poller's cursor.
Bounded output. Admin events (oracle_changed, ownership_transferred, …)
are logged at the progress level and not sent. Unknown or malformed events are
logged and skipped. Send failures are retried on the next cycle; each individual
Telegram send is bounded by TELEGRAM_SEND_TIMEOUT_MS (default 15000) so a
stalled or rate-limited Telegram endpoint cannot wedge the poller. A send that
exceeds the timeout is counted as a failure, logged with the target and event
name (never the token or payload), and the cursor is not advanced past it, so
the event is retried after a restart. Cursor files remain version-1 and
compatible with existing deployments; no configuration change is required to
keep current behavior.ailures are counted as skipped and do not abort the
run. No bot token or private key ever appears in progress output or the JSON
report.
JSON report shape (--json stdout, one document, ends with \n):
{
"format": "mimir-replay-v1",
"network": "testnet",
"rpcUrl": "https://soroban-testnet.stellar.org",
"fromLedger": 4226500,
"toLedger": 4226800,
"dryRun": true,
"targets": [
{
"source": "market",
"contractId": "CDV6JXIJCALSXQELCS6YUEWJWG5DFXQK5PJ5I7MWI6KVMQJBC5DLPKZI",
"fromLedger": 4226500,
"toLedger": 4226800,
"startLedger": 4226500,
"startClamped": false,
"pages": 3,
"events": 2,
"sent": 0,
"skipped": 0,
"capped": 0,
"adminLogged": 0,
"duplicates": 0,
"truncated": false,
"lastEventLedger": 4226729,
"cursor": "0018276211125911551-4294967295",
"eventLog": []
}
],
"totals": {
"events": 2,
"sent": 0,
"skipped": 0,
"capped": 0,
"adminLogged": 0,
"duplicates": 0
}
}bigint amounts are serialized as decimal strings (same convention as
mimir-scan-v1), so npm run replay -- --json | jq is valid. Progress and
warnings always go to stderr.
/status says what the poller is doing right now. The audit trail answers the
question after a week of unattended running: what actually happened — scan
failures and recoveries, failed Telegram sends, skipped admin events, bursts
truncated by the per-cycle cap, cursor loads, stale cursors and cursor write
failures.
It is an append-only JSONL file (data/audit.jsonl by default; AUDIT_FILE
changes it, leaving the value empty disables it). The poller appends after every
cycle, so the trail survives restarts alongside the cursor. Read it two ways:
npm run audit # report from data/audit.jsonl
npm run audit -- --tail 50 # render the 50 most recent lines
npm run audit -- --json # machine-readable stats only
npm run audit -- --file p.jsonlor send /audit in the chat as the operator, which merges the live in-memory
window with the file so entries not yet flushed are still visible.
Everything in the trail is safe to paste into an issue, and this is enforced when an entry is recorded, not by caller discipline:
- Free-text details pass redaction first: bot tokens, secret/seed strkeys, URLs
and any unrecognized long token are replaced. Public
C…contract ids andG…account ids stay readable — they are chain identifiers/statusalready prints. - Details are length-clamped (240 chars). No payloads, payment amounts as log lines, or unbounded remote data are ever stored — the chain is the record.
- The in-memory window and the report are both bounded, and the report says so when older entries were not shown.
- Reading never throws on you: an unreadable or unknown-version line is skipped and counted, never fatal.
MIMIR_PROFILE=mock (or the scanner's --mock flag) fills in any config value
the environment leaves unset with a local, loopback-only Soroban mock:
fixture contract ids, http://127.0.0.1:8420 RPC, and an isolated cursor file
at data/cursor.mock.json so a drill can never touch the real bot's position.
Explicit environment variables always win, any other profile name fails fast at
startup, and nothing here needs a bot token, Telegram credentials, or Testnet.
npm run mock:rpc # serve the fixture scenario on 127.0.0.1:8420
npm run scan:mock # scanner --mock: decode the scenario, no credentials
npm run mock:poll # dry run: mock RPC + real poller, sends are logged
npm run mock:poll -- --fail-events error # inject in-band JSON-RPC failures
npm run mock:rpc -- --stale-cursor # reject cursors once the poller has one
npm run mock:poll -- --malformed # append an undecodable event (must skip, not crash)
npm run mock:poll -- --port 0 # ephemeral port (any entry point accepts it)Failure kinds are error, http-500, rate-limit, and stale-cursor, with
the shorthands --fail-rpc, --rate-limit, --stale-cursor for getEvents
and --fail-health <kind> for getHealth. The mock enforces the real RPC's
request rules — mutually exclusive startLedger/cursor, the retained floor as
an error rather than an empty page, bounded error messages — and the dry run
exercises the poller's cursor-safety, restart, and bounded-log guarantees end to
end. npm test covers all of it (tests/mock-*.test.mjs); run just those with
npm run test:mock.
Each Soroban event is size-checked before native decoding in
src/stellar/decode.ts:
| Cap | Default | Failure mode |
|---|---|---|
Event value XDR |
16 KiB (MAX_DECODED_EVENT_XDR_BYTES) |
Event becomes unknown with a short reason; cursor still advances |
| Per-topic XDR | 1 KiB (MAX_EVENT_TOPIC_XDR_BYTES) |
Same |
| Decoded string field | 2048 chars (MAX_DECODED_STRING_CHARS) |
Same |
Oversized or malformed events never crash the poller and are never re-logged as raw remote bytes — only a truncated reason string is kept. No configuration change is required; existing cursors remain valid.
Soroban's getEvents is not eth_getLogs, and the difference is the whole
design of src/stellar/events.ts:
- Paging is by opaque cursor, not block range, so the walk is inherently sequential — there is no chunk fan-out to parallelise.
startLedger/endLedgerandcursorare mutually exclusive in one request.- The RPC keeps only a rolling window of events (~120,960 ledgers, roughly a
week, on Testnet). A
startLedgerbelow the retained floor is an error, not an empty result, so the floor is clamped fromgetHealth()first. - The window from
getHealth()is validated before the first request: an inverted or malformed window fails with a bounded error, a start ledger below the floor is clamped up to it, andnpm run scan -- --from <future>is refused rather than silently reading a different range. - An empty page does not mean the scan is finished. One request covers a
bounded slice of ledgers and returns whatever was in it — frequently nothing —
plus a cursor to continue from. Terminating on a short page (the correct
instinct for
eth_getLogs) silently yields zero events. Verified against the live deployment: reading the market contract from the retained floor takes 13 pages, 12 of which are empty, to reach the page holding all 11 of its events.
So the walk terminates on the cursor, never on the payload.
A cursor is inclusive of the event it names: the same event can come back from a later page of the same walk, and again from the next cycle that resumes from the persisted cursor — including the first cycle after a restart. Left unguarded, one on-chain event becomes two identical chat messages.
The reader and the poller therefore share a small, bounded dedup window
(src/dedup.ts): the ids of the most recently processed events per contract,
oldest evicted first. A redelivery inside that window is dropped and counted
instead of posted — visible as duplicates= in npm run scan output, as
eventsDeduplicated on GET /health and in status.json, and as deduped in
/status. The window is seeded into every scan from the cursor file, so the
guard survives a restart, and it never grows with chain history: an event older
than the window can legitimately be announced again, which is the accepted
trade-off for O(1) memory and a cursor file that stays small. Set
EVENT_DEDUP_WINDOW=0 to disable suppression.
One event, one key. Window membership is decided by a single canonical key
(eventKey() in src/dedup.ts), derived in strict order:
- The RPC's own paging token
id(<TOID>-<index>) when present — unique per event and stable across pagination; used verbatim so ids persisted by older releases keep working unchanged. - The same token under the name
eventId. - A content-derived composite
v2:<ledger>:<txHash>:<txIndex>:<opIndex>:<digest>when a response omits both — chain positions plus a fixed-length SHA-256 digest of the topic XDR. Two events share this key only when every identity-bearing field agrees, so events emitted by the same transaction (even the same operation) are never falsely merged. - Otherwise
null: notxHash, or topic content that cannot be encoded deterministically. Those events pass through undeduplicated rather than risk suppressing a real one — a duplicate is recoverable, a wrongly dropped event is not.
The scanner, the poller, and npm run replay all route through this one
function, so an event is never identified one way in one place and another way
somewhere else.
This is suppression, not backfilling. A dropped duplicate does not hold the cursor back — the chain remains the record and the walk still advances.
The poller exposes empty-page telemetry for this: cumulative and last-cycle
counts of empty getEvents pages appear in /status and in each scan log line
(N page(s) (M empty)). High empty-page ratios are expected on quiet contracts;
a sudden drop to zero pages, or truncated without progress, is the signal to
investigate RPC or EVENT_MAX_PAGES.
Events are also not a source of truth for current state — a claim's stakes and status come from the contract's own getters. This bot is a timeline, not an index.
A rolling window means a cursor can outlive the RPC. If the bot is stopped long enough (a multi-day outage, a wedged host, a long Telegram outage holding a deploy), the persisted cursor can fall below the retained floor — everything it points at is already gone. Soroban rejects such a read as stale, and holding the cursor would fail that contract's scan forever.
The poller now recovers from exactly that case, without guessing:
- On a stale rejection it asks
getHealth()for a fresh window and only acts when the cursor's own ledger places it strictly belowoldestLedger. - It then drops the doomed cursor and rescans from
oldestLedger(astartLedgerwalk, sincecursorandstartLedgerare mutually exclusive in one request). Everything below the floor was already unreadable, so nothing still retrievable is skipped, and the chain remains the record. - The recovery is bounded: at most
MAX_FLOOR_REWINDS(3) consecutive automatic rewinds per contract, then the poller stops and logs that operator action is required. A misbehaving RPC cannot make it thrash. - It is conservative: an opaque cursor this build cannot place, a cursor ahead of the tip, or a window that cannot be read is left untouched and the bounded RPC error is surfaced. Nothing is rewritten on a hunch.
- A stale rejection immediately marks that target
cursorStaleinstatus.jsonandGET /health, and makes readiness return503until that target completes a successful scan./statusand/healthidentify whether the cursor is unchanged or a retained-floor recovery is underway. The alert is reconstructed after restart from the persisted rewind position or the next RPC rejection; the version-1 cursor schema does not change. - Railway's configured
GET /healthdeployment probe therefore remains unready while a stale cursor is unresolved. Recovery continues in-process; do not delete or replace the persistent cursor volume to force readiness.GET /health/livestays200for supervisors that need process liveness independently of readiness. - A cursor that simply never advances is the other half of the story: the
scan keeps succeeding, so no error counter moves, yet the walk is behind the
tip and never closes the gap. After
CURSOR_STALL_CYCLES(5) such cycles with a lag of at leastCURSOR_STALL_MIN_LAG_LEDGERS(2) ledgers, the poller logs aCURSOR STALLEDwarning once,/statusgains acursor stalled:line, and/healthreportscursorStalledper target and becomesdegraded. Sitting at or within a ledger of the tip is healthy idling and never counts, and a cold start's first cursor is progress, not a stall. Nothing is rewritten: the cursor file is left intact and the bot keeps retrying. - The miss is logged as a bounded ledger count (
cursor is N ledger(s) below the retained floor), never as a raw RPC payload, and/statusandGET /healthexposecursorRewindsplus the per-targetrewindFromLedgerwhile it lasts. - Detection does not wait for the rejection. Before it asks for events, every cycle places a resumed cursor against the floor the last successful scan proved, so a position the RPC answers with an empty page instead of a stale error is caught too - that case used to retry the same ledger forever without a line in the logs. One comparison and no extra request, and it can only fire for a position this build can place strictly below the floor.
A recovery is observable: /status gains a Cursors rewound to the retained floor: N line once cursorRewinds > 0, and status.json reports the same
counter and the active rewindFromLedger. The counter is per process, so it
resets on restart; the position itself is persisted so a restart mid-recovery
resumes from the same floor.
A position that is provably below the floor is a restart gap: the ledgers between it and the floor are gone for good. Main's recovery resumes from the floor, so the retained window is delivered late rather than dropped — but how much was lost, and when, is what an operator needs to see, and it is reported once per occurrence rather than once per cycle:
/statusaddsrestart gap: N ledgers unrecoverable, cursor reset <ago>for each affected target, andRestart gaps detected since start: Nwhen a gap has been seen.GET /healthaddsrestartGaps, alastRestartGapobject (at,source,cursorLedger,oldestLedger,missedLedgers) and per-targetgapLedgers,cursorResetAtandcursorUnreadable.- The
/healthcommand adds the same summary and the per-target gap.
A cursor whose ledger cannot be read out of the opaque token is not a gap: it
is reported as cursorUnreadable and forwarded unchanged, because a token shape
this build does not understand is not evidence that the RPC will reject it. Only
a position this build can place strictly below the floor is a gap.
The decoder is deliberately forward-compatible at the event boundary:
- Soroban event topics are read in declaration order, and non-topic fields are read from the event value map using their deployed snake_case names.
- A known event with a malformed topic, value, address, integer, or XDR value
becomes an
unknownevent.decodeEventnever throws into the poller, so one bad event cannot stop a scan or move a cursor based on a partial payload. - Events that are valid on-chain but unknown to this version are retained as
unknownfor bounded logs and are skipped for Telegram. They are not invented, retried, or treated as current contract state. - Amounts remain
bigintatomic USDC values until formatting; no floating-point conversion is used. Contract strings are clipped at the notification and diagnostic boundaries, and scanner output is bounded.
The compatibility promise is for the deployed event wire shape and the public decoded payload names above, not for arbitrary XDR or future contract fields. Adding an optional field is safe when the existing fields retain their names and types. Renaming a topic or changing a field type is a decoder compatibility change and must be deployed together with a recorded fixture and an operational note. The chain remains authoritative if the bot version cannot decode an event.
RPC configuration is read-only: STELLAR_RPC_URL must point to a Soroban RPC
endpoint, and STELLAR_NETWORK_PASSPHRASE controls network labeling for
explorer links. The client sends no signing material and creates no wallet.
Explorer path components are URL-encoded; custom STELLAR_EXPLORER_BASE_URL
values are supported without changing cursor or decoder compatibility.
The poller writes its resume position to data/cursor.json (write-then-rename,
so a crash mid-write cannot truncate it). The on-disk document is a versioned
schema (version: 1 today), and the poller appends audit entries to
data/audit.jsonl. Both must survive restarts, so give data/ the same
treatment as the cursor:
{
"version": 1,
"updatedAt": "2026-08-21T10:00:00.000Z",
"targets": {
"market": {
"cursor": "0018276211125911551-4294967295",
"lastEventLedger": 4226729,
"recentEventIds": ["0018276211125911551-4294967295", "0018276211125911552-1"]
},
"squad": {
"cursor": "0018276211125911551-4294967295",
"lastEventLedger": 4226733,
"recentEventIds": []
}
}
}recentEventIds is the persisted dedup window (see
Overlapping pages and duplicate events)
and is additive: it is bounded by EVENT_DEDUP_WINDOW (default 256) and older
cursor files without the field load as an empty window. Entries are the
canonical dedup keys: mostly raw RPC paging tokens, but a v2: prefix marks a
content-derived composite for a response that arrived without one. Both forms
round-trip as opaque strings across restarts, and mixed generations coexist
harmlessly — a file written by an older release may still hold its
<ledger>:<txHash>:<count> strings, which simply age out of the bounded window
as new keys are added. No schema bump and no migration are involved. The
version and the cursor / lastEventLedger fields are unchanged, so the
format stays backward-compatible in both directions.
rewindFromLedger is a second additive field, written only while a target is
recovering from a stale cursor (see
Recovering from a stale cursor). It records
the retained floor the next scan resumes from, so a restart in the middle of a
recovery keeps reading from that floor instead of cold-starting. It is dropped
as soon as the floor walk returns a fresh resume cursor, so an ordinary cursor
file never contains it; files written before it existed load as "no rewind
pending", and older builds ignore it.
Compatibility / migration
- Current files (
version: 1) load as-is. - Legacy unversioned envelopes (
{ "targets": … }withoutversion) and flat maps ({ "market": { "cursor": … }, "squad": … }or string cursors) are migrated in-place to schema v1 on startup, then rewritten atomically. - Unknown future
versionvalues are rejected: the file is quarantined and the bot cold-starts rather than guessing, so a downgrade cannot mis-read a newer file. Upgrade the bot before rolling forward again. - Corrupt JSON or unrecognised shapes are also quarantined and cold-started (see failure behaviour below). Logs never include the raw file body.
On a cold start (no usable file) it begins START_LOOKBACK_LEDGERS behind the
chain tip rather than replaying the whole retained window into your chat.
/pause and /resume never edit this file; they only control scheduling, so the
cursor format remains version 1 and a restart does not preserve a pause. A
graceful shutdown flushes any cursor state that is still only in memory before
the process exits — see Graceful shutdown.
Tests never use this directory: they run against an ephemeral data directory created under the OS temp dir and removed afterwards (see docs/contributor-fixtures.md).
For local restart and regression checks without Testnet or Telegram credentials, seed that file with a deterministic fixture:
npm run seed:cursor # writes ./data/cursor.json (refuses overwrite)
npm run seed:cursor -- --force # replace an existing file
npm run seed:cursor -- --empty # null cursors (file present, cold resume)
npm run seed:cursor -- --out /tmp/cursor.jsonThe seeder uses the same write-then-rename discipline as the poller, never reads bot tokens or signing keys, and refuses cursor values that look like secrets.
If the file exists but is corrupt (truncated JSON, wrong version, or a
non-object targets map), the poller renames it to
CURSOR_FILE.corrupt.<ISO-timestamp> and cold-starts. That keeps the bad file
for debugging and prevents the next saveCursors from silently overwriting
evidence of the failure.
The version field is checked on load. A file with a missing or unrecognised
version is treated as a cold start (warning logged) rather than silently
misread. This protects against reading a file written by a future release after
a downgrade.
Deployment note: a flat file is fine for v0 but it must survive restarts. At startup, the poller verifies persistent-volume availability and cursor-file permissions by testing write access to the configured CURSOR_FILE directory and read/write access to any existing cursor file. If persistent-volume or file-permission verification fails, the poller logs an actionable warning and falls back to in-memory cursor management without crashing, ensuring operational continuity while surfacing volume warnings via /status and /health. On an always-on host, put data/ on a persistent volume (or point CURSOR_FILE at one). On an ephemeral filesystem every restart is a cold start, and events that happened while the bot was down are never posted. Swapping this for a real KV store is a deliberate future step, not something this repo does today.
The poller takes an exclusive file lock (data/poller.lock by default, overridable
with INSTANCE_LOCK_FILE) before it loads the cursor or starts Telegram long
polling. The lock records only pid, hostname, and acquiredAt — never the
bot token or any secret.
- A second live process against the same lock exits immediately with a clear error, so two notifiers cannot race the cursor or double-post events.
- If the previous process died without releasing the lock, the next start detects the dead pid, removes the stale file, and continues.
- Point
INSTANCE_LOCK_FILEat the same persistent volume asCURSOR_FILEso the lock survives the same restarts the cursor does.
For detailed operational behavior, deployment assumptions, and configuration tuning, see the Configuration and Operational Reference.
This process is meant to stay up for weeks, so a single failure never ends it:
-
A failed RPC call fails one contract's scan for one cycle. Its cursor is left untouched, so the next cycle resumes exactly where it stopped.
-
A failed Telegram send drops one message; the cursor still advances. That is deliberate: holding the cursor back would turn a revoked token or a chat the bot was removed from into an infinite replay, and recovery would flood the channel. Notifications are lossy on purpose — the chain is the record.
-
A corrupt cursor file is treated as a cold start rather than a crash.
-
A cursor file with a missing or unrecognised
versionfield is treated as a cold start; a warning is logged with instructions to delete the file. -
A stale cursor — one pointing to a ledger more than ~10 % of the retention window behind
oldestLedger— causes a warning that names the cursor file, the gap size, and how to recover. Events in the gap are already gone from the RPC and will not be posted. -
A burst is capped at
MAX_NOTIFICATIONS_PER_CYCLEmessages per cycle. A single warning is logged when the cap is hit, and the cursor still advances. Telegram's rate limiter is never the thing that takes the bot down. -
Consecutive full-cycle failures are counted. A structured warning is emitted the first time the count crosses 5, 10, 25, 50, or 100 consecutive failed cycles, naming the RPC URL, the last error, and confirming the cursor is intact.
-
A partial notification batch commits the opaque RPC cursor after the returned page has been processed. Unknown events, events beyond
MAX_NOTIFICATIONS_PER_CYCLE, and sends that exhaust three bounded retries are counted as skipped or failed. Holding the cursor back would turn a revoked token or removed chat into an infinite replay, so the cursor always advances — but a send that exhausts its retries is parked in a bounded local queue and replayed on a later cycle once Telegram recovers, so a rate limit or a short outage costs a delay rather than the message. See Dead-letter queue. A failed send is isolated to that routed chat and event; other events continue. Skips are lossy on purpose — the chain is the record; the poller logs the sent/failed/skipped/parked commit decision. A rejected inline keyboard (or a malformed MarkdownV2 payload) fails the same way as any other send. Events without a usable transaction hash are still sent, just without the explorer button. -
A corrupt cursor file (invalid JSON or wrong schema) is quarantined to
data/cursor.json.corrupt.<timestamp>beside the live path, then treated as a cold start; the next successful cycle writes a freshcursor.json, and the quarantined copy is kept for operators instead of being overwritten. -
A valid but stale cursor (one below the retained floor) is caught before the scan as well as on the RPC's own rejection, and rewound to that floor in bounded steps: the poller confirms the position against a fresh
getHealth(), drops the unreachable cursor, rescans fromoldestLedger, and records the recovery plus the ledgers it lost in/status,status.jsonandGET /health. It never guesses — an opaque cursor, an ahead-of-tip cursor, or a window that cannot be read is left untouched and the bounded RPC error is surfaced. AfterMAX_FLOOR_REWINDS(3) consecutive rewinds for one contract the poller stops and asks for operator action. See Recovering from a stale cursor. -
A second concurrent instance is refused at startup via the exclusive lock above. Stale locks from crashed processes are cleared automatically.
-
A request outside the retained window never reaches the RPC in one piece. The window (
oldestLedger…latestLedger) is validated fromgetHealth(); a resume cursor that the token itself places above the chain tip is refused before it is sent, with a bounded error in/statusand logs and the stored cursor left unchanged. A cursor below the retained floor is still forwarded — retention is the RPC's call — and when the RPC rejects it the poller rewinds to the floor (see Recovering from a stale cursor). A cursor shape the bot cannot read is forwarded too, so an RPC cursor-format change cannot wedge it. -
A burst is capped at
MAX_NOTIFICATIONS_PER_CYCLEmessages per cycle, spaced out, so Telegram's rate limiter is never the thing that takes the botdown. RPC, Telegram, and poller error text shown in
/statusor logs is compact, bounded, and the configured bot token is redacted. -
An unreadable audit line (or a failed append) is logged and skipped; the audit trail never throws into the poll loop, and a bad line never takes the report down. An audit file that cannot be read at all reports as empty.
-
A duplicate event — the same id from an overlapping page, a resumed cursor, or a restart — is suppressed and counted (
eventsDeduplicated), never posted twice. It does not hold the cursor back. Bounded per contract byEVENT_DEDUP_WINDOW(default256;0disables). -
A long Stellar outage freezes the chain clock at the newest close time the RPC actually reported.
/statusandGET /healththen show a growing skew rather than a clock that keeps time on its own, so a stalled chain and a wrong local clock stay distinguishable. The clock is saved with the cursors, so a restart resumes it instead of reportingunknown, and an event without aledgerClosedAtnever counts as a chain time. -
A repeating failure is logged through a sampler: the first
LOG_SAMPLE_MAX_PER_WINDOWidentical errors in aLOG_SAMPLE_WINDOW_MSwindow print in full, later repeats are counted, and the line that opens the next window reports how many were held back. An RPC that is down for hours therefore costs a handful of lines instead of one per cycle. The running total is exposed aspoller.suppressedLogson/health;/statusstill shows the latest error verbatim. -
An operator pause prevents new cycles but cannot cancel a bounded scan or Telegram retry loop already in progress. That cycle follows the normal cursor rules above;
/resumestarts the next cycle immediately. -
A status file that cannot be written is logged and ignored; it is an observability signal, never a reason to stop notifying. A corrupt snapshot makes
--statusexit1rather than print garbage. -
A shutdown stops scheduling, drops what has not been sent yet, waits at most
SHUTDOWN_TIMEOUT_MSfor the cycle in progress, and flushes any cursor state that is still only in memory — see Graceful shutdown.
A send that fails every bounded retry is parked in data/dead-letter.json
(DEAD_LETTER_FILE) instead of being dropped. The two failure modes are
different: a rate limit or a short Telegram outage is a delay, while a revoked
token is permanent, and the send site cannot tell them apart. Parked messages are
replayed oldest-first at the start of a later cycle, under the same per-cycle
budget as a normal burst, and the file is bounded on every axis:
| Bound | Setting | Default | Behaviour when reached |
|---|---|---|---|
| Depth | DEAD_LETTER_MAX |
100 |
the oldest parked send is dropped |
| Attempts | DEAD_LETTER_MAX_ATTEMPTS |
10 |
the entry is dropped as poison |
| Message body | — | 4,000 chars | the stored text is truncated |
| Recorded error | — | 200 chars | redacted (bot token), then truncated |
Identity is the event, not the attempt: the same event failing twice is one entry with two attempts, so a failing cycle cannot fill the queue with duplicates. Only the already-formatted message and short operational metadata (source, ledger, event name, redacted error) are stored — never a signing key — and the file is written write-then-rename, so a crash cannot corrupt it.
Two things it deliberately does not do:
- It never holds the cursor back. The cursor advances exactly as it did before, because a permanently broken destination must not stop the poller or replay the chain forever.
- It never replays what a shutdown dropped. A message that was never
attempted at
SIGTERMis counted as dropped and left to the chain, so a restart cannot resurrect a stale backlog into the channel.
/status and GET /health report the queue's depth, replays and drops. Omitting
DEAD_LETTER_FILE disables the queue entirely, and the poller then behaves
exactly as it did before the queue existed.
SIGINT/SIGTERM starts a bounded drain rather than a hard stop:
- The poller stops scheduling cycles and reports itself as
stopping. - Notifications that have not been sent yet are dropped: counted in
/status, logged once with a bounded line, and left to the chain. A send already in flight is allowed to finish, but it does not start another retry/backoff step. - The cycle in progress gets
SHUTDOWN_TIMEOUT_MS(default10000,0skips the wait) to finish and write its cursors. - Any cursor state still only in memory is flushed to
CURSOR_FILE, then the health endpoint and the Telegram long-poll are closed and the process exits0.
Cursors only ever advance after their events have been handed to Telegram, so flushing at any point is safe: the file a restart resumes from never skips an event the chain still has to show. What the drain gives up is delivery of the messages it had not started — notifications are lossy by design and the chain is the record, exactly as for a failed Telegram send.
Why drop rather than finish the burst? Finishing means up to
MAX_NOTIFICATIONS_PER_CYCLE messages × 1.5s spacing plus retry backoff —
minutes that would hold a deploy open. Worse, hitting the deadline halfway
would leave the cursor behind messages that were already sent, replaying them
on restart. Dropping keeps the drain bounded and the resume exact.
A second SIGINT/SIGTERM exits immediately (130/143) if a drain ever
gets stuck. That skips the flush but never corrupts the file: the cursor is
written to a temporary file and renamed, so the worst case is resuming from the
last completed cycle. The teardown after the flush is capped too —
SHUTDOWN_TIMEOUT_MS + 10000 ms, then the process exits 1 with the cursor
file already written.
Where the drain is visible:
| Where | Field |
|---|---|
GET /health |
poller.stopping, poller.pendingFlush, poller.lastFlushAt, poller.notificationsDropped. A deliberate drain reports ok, not degraded |
/status |
stopping in the headline, dropped during shutdown N in the counters, and a drain line while it lasts |
Configuration is additive: SHUTDOWN_TIMEOUT_MS is optional (see
.env.example), no existing variable is renamed, and the version-1 cursor
format is unchanged — a deployment that omits the new key gets the 10000 ms
default.
Operational errors are sampled so a long outage cannot drown the log. Per key —
scan:market, scan:squad, and cycle — the first LOG_SAMPLE_MAX_PER_WINDOW
identical lines inside a LOG_SAMPLE_WINDOW_MS window are printed in full.
Further repeats are counted, and the line that opens the next window reports how
many it held back (suppressed N identical repeats in the previous 300s). The
running total is exposed as poller.suppressedLogs on /health, so a sampled
log still accounts for every failure. Keys are independent: a noisy market
contract cannot silence the squad contract's errors, and the text itself is
still bounded and redacted exactly as before.
Configuration (see .env.example):
LOG_SAMPLE_MAX_PER_WINDOW— full lines per key per window (default3;1logs only the first line of a run)LOG_SAMPLE_WINDOW_MS— window length in milliseconds (default300000; minimum1000)
Sampling is in-memory and resets on restart. It changes log volume only — never
cursors, sends, retry counts, or the /status error text.
Failure modes: the sampler is pure bookkeeping keyed by string, so it cannot
throw into the poll loop. Rollback: omit both keys to keep defaults, or raise
LOG_SAMPLE_MAX_PER_WINDOW until nothing is suppressed.
The notifier is meant to run for weeks through Stellar RPC and Telegram outages. Everything it keeps in memory is fixed-size or capped:
- Per-target state is one small fixed record (cursor, last event ledger, an optional pending floor-rewind ledger, last error).
- The chain clock is one timestamp (the newest observed close time) plus its derived skew; it never accumulates history.
- A scan walks at most 20 event pages, and each cycle sends at most
MAX_NOTIFICATIONS_PER_CYCLEmessages; the rest are counted as skipped. - Error text is redacted (bot token) and clipped before it reaches
/status,/health, or logs; unknown or malformed events are logged as one bounded line. - At most one poll timer is pending, and
stop()leaves none behind.
tests/soak.test.mjs enforces this offline: it drives about 1,700 poll cycles
through a scripted fake RPC (outages, stale-cursor rejections, malformed and
unknown events) with every Telegram send failing, under mocked timers. It asserts
that heap growth after a forced GC stays under 4 MB, that status and every log
line stay bounded and token-free, and that timers do not accumulate. A control
test deliberately leaks per send and must trip the same threshold, so the check
cannot silently stop working. It needs no Testnet, Telegram credentials, or keys.
Deployment assumptions: one process per chat and cursor file (two writers
would race on CURSOR_FILE), the cursor path on persistent storage, and a
supervisor that restarts the process and probes GET /health. If you suspect a
leak in production, watch the process RSS over days; a restart is always safe.
Rollback: deploy the previous build and start it against the same
CURSOR_FILE. The cursor format is unchanged (version 1): chainClockAt and
rewindFromLedger are optional additive fields that older builds ignore and
newer builds drop when absent or malformed, and the chain is the source of
truth, so nothing is replayed beyond the last saved cursor and nothing needs
migrating. An older build that meets a mid-recovery file simply cold-starts that
contract from START_LOOKBACK_LEDGERS rather than mis-reading it. Keep a copy of
the cursor file if you want an exact resume point.
- A message over Telegram's 4096-character limit is split on newlines (or
hard-split when a single line is oversized) and sent as ordered chunks. The
poller's send path and
/status(and other MarkdownV2 replies) share this splitter so a long status or a rare oversized event never fails with "message is too long".
Before the poller starts, boot calls Soroban RPC getHealth() with bounded
retries (STARTUP_HEALTH_DEADLINE_MS / STARTUP_HEALTH_RETRY_MS) so a brief
RPC outage does not abort startup, while a bad URL still fails within the
deadline.
The process exposes a loopback HTTP probe for supervisors and deploy
checks (default http://127.0.0.1:8787):
| Path | Meaning |
|---|---|
GET /health (alias /healthz) |
Readiness-style status. 200 when the poller is running and healthy, including an intentional operator pause; 503 when stopped or degraded (repeated RPC failures or a stale success window). The response includes poller.paused and poller.suppressedLogs. |
GET /health/live (alias /livez) |
Liveness only — the process and HTTP server are up. Always 200 while listening. |
GET /health/diag |
The same report /export produces in the chat, as plain text. 404 when log capture is off. For an operator with host access but no seat in the chat. |
Container Healthcheck: A dedicated CLI probe is available for Docker HEALTHCHECK or Kubernetes exec probes. It reads the same environment variables and exits 0 on success:
npm run healthcheck
# or: node dist/healthcheck.jsThe JSON body is operational status only: poller counters, ledgers, truncated
cursors, whether a target has an error, automatic floor rewinds
(poller.cursorRewinds plus each target's rewindFromLedger), restart gaps
(poller.restartGaps and poller.lastRestartGap, plus each target's gapLedgers,
cursorResetAt and cursorUnreadable), and the chain
clock (poller.chainClockAt plus poller.chainClockSkewMs, the signed difference
in milliseconds between the bot's clock and the newest chain close time it has
observed — positive while the bot is ahead). Each target's cursorStale boolean
indicates an unresolved RPC rejection; resolving it requires a successful scan,
not a health-probe retry or local cursor-age guess. It never includes
BOT_TOKEN, chat ids, private keys, or unbounded remote payloads.
GET /health also answers where each setting's value came from, and never what
it is. That distinction is the difference between "the bot is configured" and
"the bot is configured the way I think it is": a placeholder token inherited from
a profile, a .env the process never found because it started from another
directory, and a variable exported empty all look identical from the outside.
{
"config": {
"profile": null,
"envFile": { "present": true, "suppliedKeys": 12 },
"entries": [
{ "key": "BOT_TOKEN", "source": "env-file", "secret": true },
{ "key": "HEALTH_PORT", "source": "derived", "derivedFrom": "PORT", "secret": false }
],
"counts": {
"process-env": 3,
"env-file": 12,
"profile-default": 0,
"built-in-default": 6,
"derived": 1,
"unset": 4
},
"warnings": []
}
}source is one of process-env, env-file, profile-default,
built-in-default, derived (another setting supplies it, named by
derivedFrom), or unset — absent and optional, which is the normal state for
ALLOWED_CHAT_IDS and OPERATOR_TELEGRAM_USER_ID. No value — token, chat id, or
anything else — is ever part of the report, so it can be pasted into a ticket
as-is; secret: true marks the settings that are sensitive for exactly that
reason. warnings names what is worth acting on: a variable set but empty, a
.env that supplies none of the known settings (usually a working-directory
bug), an unknown MIMIR_PROFILE, or the mock profile being active.
Boot logs the same information as one line, followed by any warnings:
[boot] config profile=none env-file=present(12 keys) process-env=3 env-file=12 built-in-default=6 derived=1 unset=4 secret-keys=6/26
[boot] config HEALTH_STALE_MS is set but empty; the built-in default supplies the value
/status ends with the same one-line summary, so an operator can confirm which
.env a deployment actually read without opening a shell.
Configuration (see .env.example):
HEALTH_HOST— bind address (default127.0.0.1; set to0.0.0.0for Docker)HEALTH_PORT— TCP port (default8787;0disables)HEALTH_STALE_MS— degraded if no successful poll within this window after the first success (default90000;0disables)- A stale cursor rejection independently makes
GET /healthreturn503until its target scans successfully;HEALTH_STALE_MSdoes not disable this cursor alert. STARTUP_HEALTH_DEADLINE_MS— wall-clock budget for retrying the boot RPCgetHealth()probe (default30000;0= single attempt)STARTUP_HEALTH_RETRY_MS— delay between failed boot RPC health attempts (default1000)
Each poll cycle gets a fresh UUID correlation ID. The ID is included in poller
logs, /status, and /health so an RPC failure, malformed event, rate-limited
send, or cycle error can be tied to one bounded cycle. It is not included in
Telegram messages, persisted cursor data, or carried across a restart; the
chain remains the source of truth.
Rollback: set HEALTH_PORT=0 (or omit the new env keys to keep defaults) and
redeploy the previous image — the target alert is additive, does not change the
version-1 cursor format or Telegram delivery, and the previous build safely
ignores the new status field.
Failure modes: binding fails only if the port is already taken (process exits via the listen error path after logging). Client disconnects and probe errors are logged and ignored so they cannot stop the notifier.
- Notification feature flags (
NOTIFY_ENABLED,NOTIFY_MARKET,NOTIFY_SQUAD) are coarse kill switches for Telegram posts. Disabled events are skip-logged and the cursor still advances; unset defaults keep prior always-on behavior.
On startup, the bot verifies that the configured STELLAR_NETWORK_PASSPHRASE
matches the RPC's own network passphrase (from getNetwork()). A mismatch is a
configuration error and the process exits fail-fast, rather than silently reading
events from the wrong chain. The error message names both the configured and
actual passphrases so the mismatch is immediately actionable.
Telegram can preview URLs in messages. The bot controls link previews per notification:
- Enabled for events with transaction hashes in the footer, since explorer links provide useful context about on-chain activity.
- Disabled for command responses (status, help) and events without transaction data, to keep the UI clean and avoid loading preview metadata.
This is driven by the presence of a non-empty txHash field on each event, not
configuration. All persisted state and cursor format remain unchanged.
src/
index.ts entry point: config -> RPC -> bot -> poller -> health HTTP
mock-run.ts dry run: in-process mock RPC + real poller, log-only sends
health.ts local loopback GET /health for supervisors
config.ts env loading and validation, fails fast (MIMIR_PROFILE profiles)
bot.ts grammy setup: /start, /help, /status, /audit, /contracts, /health, /preview, operator pause/resume
dedup.ts canonical event dedup keys (eventKey) + bounded window
poller.ts the loop: scan, notify, persist the cursor, flush audit
log-sampler.ts windowed suppression of repetitive error log lines
audit.ts redaction, bounded audit log, JSONL persistence, report renderer
audit-cli.ts entrypoint for `npm run audit`
replay-cli.ts entrypoint for `npm run replay` (cursor-range replay)
instanceLock.ts exclusive process lock for the cursor owner
status.ts machine-readable status snapshot (allowlisted, bounded)
dev/
seedCursor.ts credential-free local cursor seeder (npm run seed:cursor)
stellar/
client.ts Soroban RPC client + explorer links (tx + contract)
events.ts cursor-paginated getEvents (+ the standalone CLI)
decode.ts typed decoding of both contracts' events
replay.ts cursor-range replay engine + runReplayCli() (npm run replay)
mock-rpc.ts local Soroban mock: scenario, pagination, failure injection
mock-constants.ts mock profile fixture ids, ports, placeholder credentials
notifications/
format.ts decoded event -> MarkdownV2 message
tests/
format.test.mjs notification formatting (incl. deterministic fuzz)
audit.test.mjs redaction, entries, persistence, report rendering
deadLetter.test.mjs queue bounds, poison entries, redacted stored errors
poller-dlq.test.mjs parking, replay after recovery, restart, disabled path
replay.test.mjs cursor-range replay: dry-run, send, clamp, security, cursor-safety
The repo ships railway.json — Railway's Config-as-Code — that wires the
deployment to the rest of this repo:
| Setting | Value | Why |
|---|---|---|
build.buildCommand |
npm run build |
dist/ is gitignored; the image compiles it. |
deploy.startCommand |
npm start |
Run the built poller. |
deploy.healthcheckPath |
/health |
The same endpoint the local health module serves (GET /health). |
deploy.requiredMountPath |
/app/data |
Refuse to start unless a volume is attached where the cursor lives. |
deploy.restartPolicyType |
ON_FAILURE |
Restart on crash, bounded retries. |
deploy.numReplicas |
1 |
One poller owns the cursor; Railway volumes cannot be used with replicas. |
Volume is the one manual step — Railway never creates one from config:
railway volume add --mount-path /app/data # or attach it from the dashboardSet as Railway variables (secrets): BOT_TOKEN, TELEGRAM_CHAT_ID, and
HEALTH_HOST=0.0.0.0. Everything else keeps its repo default:
CURSOR_FILE=./data/cursor.json resolves to /app/data/cursor.json in
Railway's /app working directory, and the health endpoint binds the PORT
that Railway injects (the HEALTH_PORT fallback, see .env.example). No public
domain is needed — healthchecks run from Railway's probe host on the container
network, which is why HEALTH_HOST must not stay loopback-only here.
On deployed failure:
- A crash restarts under
ON_FAILURE; the volume keeps the cursor so there is no notification replay. A stale or corrupt cursor is already handled as a cold start, never a crash — see Failure behaviour. - A redeploy of a volume-backed service has a short downtime window (Railway keeps only one active deployment per volume); roll back to the previous revision and the cursor is still there.
Config-as-Code is deprecated by Railway in favour of Infrastructure as Code
(.railway/railway.ts with the Railway CLI), with a hard cutoff of 2026-12-01.
This file captures the current, working behaviour and is the migration source of
truth for IaC; follow Railway's migration guide when the time comes.
Run npm run typecheck for a no-emit TypeScript check, npm test for the build and all
tests (format, config validation, and poller behaviour — no live Testnet or Telegram
credentials required), or npm run build to produce the production output.
Run npm run typecheck for a no-emit TypeScript check, npm test for the build plus the deterministic command, poller, format, fixture, mock-profile, config-provenance, health, lockfile and audit-trail suites (including deterministic fuzz cases; npm run test:mock for just the local-mock suites), or npm run build to produce the production output. CI runs typecheck, build, and all tests without network credentials.
Run npm run seed:cursor to write a local cursor fixture.
package-lock.json is the install of record: deployments rebuild with npm ci,
so the committed lockfile must stay in sync with package.json and pin exactly
what it claims. Two checks enforce that, and CI runs both after npm ci:
npm run lockfile:check— offline. The lockfile islockfileVersion3, its root entry matchespackage.json's dependency ranges exactly, every package resolves to aregistry.npmjs.orgtarball with asha512integrity hash, and every direct dependency is pinned at the top level. Drift is reported by package name instead of being silently re-resolved.npm run lockfile:reproduce— asks npm to regenerate the lockfile from itself in a scratch directory and fails if the resolved package set changes, so a hand-edited or partially-resolved lockfile cannot land. The repository working tree is never written to.
The offline suite runs as part of npm test (tests/lockfile.test.mjs), so
drift is caught locally without network access. To change dependencies, edit
package.json, run npm install to regenerate the lockfile, and commit both
files together — a lockfile that no longer matches package.json fails
npm ci, npm run lockfile:check, and CI.
Run npm run typecheck for a no-emit TypeScript check, npm test for the build plus the deterministic command, poller, ledger-window, format, fixture, mock-profile, config-provenance, health, lockfile and audit-trail suites (including deterministic fuzz cases; npm run test:mock for just the local-mock suites), or npm run build to produce the production output. CI runs typecheck, build, and all tests without network credentials.
Contributor workflow for credential-free fixtures (event catalogs, cursor samples, failure-mode expectations) lives in docs/contributor-fixtures.md. Issue and pull-request templates (.github/ISSUE_TEMPLATE/, .github/pull_request_template.md) capture the failure-mode, cursor-safety, and deployment-impact details every change must describe.
AGPL-3.0-or-later, matching the rest of Mimir.