Skip to content

Commit a6a1de4

Browse files
fix(plugin-audit): read-audit reports once per CAUSE, not once per process (#18595)
Fixes #18247 `reportReadAuditWriteFailure` (`packages/plugins/plugin-audit/src/read-audit.ts`) carried the THIRD independent copy of one defect pair — the pair #15166 fixed in `audit-writers.ts` and #17452 fixed in `auth-event-audit.ts`. This PR **removes the duplicate**; it does not write a fourth implementation. The two helpers #18246 exported for exactly this purpose are imported. Clause-②: no — no export is added, no error code is added, and nothing is relaxed. Two already-exported helpers are imported and one existing message literal becomes conditional; `git diff` adds no `export` in `packages/`. ## The two defects, both live on a seam the repo already declared durability-critical 1. **Its own process-level `failureReported` boolean.** The first failure of ANY cause silenced every later failure of every OTHER cause for the life of the process. Record-view rows are written from a BUFFER off the request path, so there is no in-flight request left to notice, and the shipped `record_views` list view answers "who viewed this record" with a confident, wrong, SHORT list. 2. **Its own fixed message literal**, printing the ADR-0057 §3.6 / `OS_TELEMETRY_DB` datasource guidance unconditionally — so a fault with nothing to do with datasource routing (an `ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED` refusal, say) sent the operator to check something that was working. Its callee `persistReadAuditRows` is registered in `DURABILITY_CRITICAL_CALLEES` (`scripts/check-durability-degradation-log-level.mjs`), so the repo has already declared this write durability-critical. A reporter that switches itself off after one cause is exactly the failure that declaration cannot afford. ## What changed - The dedupe key is now `auditFailureCauseKey`, **imported** from `audit-writers.ts` rather than re-spelled — a second copy of the key is how this defect reached the second file. The counting unit is a DEGRADATION, and a second cause is a second degradation. - The object dimension of the shared key is the **ledger** (`sys_audit_log`), not the viewed object. One flush is ONE write carrying rows about MANY audited objects, so there is no single viewed object to name, and picking whichever landed first in the batch would make the key depend on traffic — the one property `auditFailureCauseKey` exists to deny. The key therefore reduces to the driver's code vocabulary: bounded by construction. (`audit-writers.ts` passes `ctx.object` and `auth-event-audit.ts` its constant `sys_session` because on those two seams the audited object IS single-valued per write.) - The first line now leads with `auditFailureCauseSummary(err, detail)`, so the driver's code and message reach the operator instead of being computed and dropped. - The ADR-0057 §3.6 remedy is asked for through the shared `isMissingTableError` predicate and printed for exactly the missing-table cause it was written for. `persistReadAuditRows` writes ONE table, so the question is asked about that one — unlike `persistAuditTrailRow`, which writes the ledger row and its `sys_activity` mirror and asks about both. Every other cause now gets the driver's own verdict plus the fix that matches it. ## What deliberately did NOT change - The once-per-degradation anti-noise rule itself. A repeat of an already-reported cause still degrades to `debug`, and AGENTS.md names "log every failure at `error`" as this rule's falsifier. - The `error`-then-`warn` sink fallback (#9657) — and its dedupe is per-cause too, so a host that injected a logger without `error` hears the second fault as well. - `audit-writers.ts` is untouched: it is the source imported FROM. - The rule that an audit failure never reaches the read. ## Evidence ### Tests — 7 new pins, `packages/plugins/plugin-audit/src/read-audit.test.ts` Four pin the defects, three are the discriminating controls that must NOT move: ``` pnpm --filter @objectstack/plugin-audit test -> exit 0 24 files / 353 tests passed pnpm --filter @objectstack/plugin-audit typecheck -> exit 0 (tsc, tsconfig.scripts.json, check:test-typecheck: 0 files / 0 errors) ``` ### Ablation — the pins are capable of failing Both defects put back (`git show` of the pre-fix blob onto the path, proven on disk: `reportedReadAuditFailureCauses` 3 to 0, `missingTable` 2 to 0, `let failureReported = false;` 0 to 1, `SHORT answer. Fix: confirm` 0 to 1; mutated blob `f202954d` vs HEAD blob `ae1e1b9a`), then: ``` Tests 4 failed | 29 passed (33) x reports a SECOND, DIFFERENT cause at error - a new cause is a new degradation AssertionError: expected [ { level: 'error', ... } ] to have a length of 2 but got 1 x carries the underlying code and message in the first line it prints AssertionError: expected 'Read-audit write FAILED - 1 record-vi...' to match /ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED/ x prints the datasource remedy for the cause it is the remedy FOR, and not for others AssertionError: expected 'Read-audit write FAILED - 1 record-vi...' not to match /OS_TELEMETRY_DB/ x [#9657] the `warn` fallback is per-cause too - a sink with no `error` hears the second fault AssertionError: expected [ { level: 'warn', ... } ] to have a length of 2 but got 1 ``` The three controls stayed GREEN on both sides, which is the half that matters: "still degrades a REPEAT of an already-reported cause to debug", "keys on the error CODE, never its message" (200 batches, 200 distinct messages, one code, one line) and "folds a fault carrying NO code into ONE bucket". Deleting the dedupe outright would redden those three. Restore leg: `git checkout HEAD -- PATH` (naming HEAD, never a bare checkout), blob back to `ae1e1b9a`, `git diff HEAD` for the path empty, `git status --porcelain` clean. No ablation artifact is left in the tree. ### Gates — `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --commands`, reconciled with `--ran` 63 derived families, all run with the exit code captured to disk before reading; reconciliation reports `63 derived, 60 run, 3 NOT-MEASURED, 0 UNRUN`. - **60 green.** - **3 NOT MEASURED** — `check:dual-build-cjs-loads`, `check:i18n`, `check:type-check-debt` all exited **3**, the code these gates use for PREREQUISITE NOT MET: each needs a full `pnpm build` (57 packages have no `dist/` in this worktree, which built only plugin-audit's dependency closure). CI builds, so CI measures them. Not read as green and not as red. - **`pnpm check:durability-log-level` — run although this card's derivation does not name it**, because the diff sits in a `catch` guarding a registered durability-critical callee. Green: `36 durability-critical catch seam(s), all loud, rethrowing or propagating to the caller`. The gate has no objection to this change. - **`pnpm check:cross-package-test-inputs` exited 1, and the finding is not this diff's.** It names `packages/cli/test/init-created-files-summary.e2e.test.ts` descending `packages/spec/dist/` — no path of mine. Mechanism: `coversDirectory`/the walked-root radius answer with `readdirSync` against the real filesystem, and `packages/spec/dist/` is a gitignored build artifact that exists here only because the dependency-closure build created it. Control: the same gate on a checkout with no `packages/spec/dist` exits 0 (`OK: 29 package(s) read outside themselves, all declared`). Reported below rather than ridden in. ### Lint — the whole population, not a narrowing ``` node --stack-size=4000 node_modules/eslint/bin/eslint.js . --no-inline-config --format json -> exit 0 files in eslint population: 6803 files with findings: 0 ``` Run at `6195b00` (the final commit), on a clean tree. ## Acceptance notes - **`reportOverflow` in the same file was examined and is NOT this class.** Its report has no cause dimension at all — the buffer overflowing is one condition, it takes no `err`, and its remedy text is already cause-agnostic. A cause key there would key on nothing. Noted, not filed; successor: whoever next touches this batcher. - **`packages/services/service-settings/src/config-change-audit.ts:157` stays out, and my reading agrees with the card's.** Its callee is a bare `eng.insert` that no register names, its first line already carries `Cause: ` plus the real detail, and its remedy text is already cause-agnostic. An observation, not a contract violation. - **`check:cross-package-test-inputs` reverses its verdict on a gitignored build artifact** (see Gates above) — reported to the PM with dedupe words for the filing seat, not filed from here and not fixed here. --- _Generated by [Claude Code](https://claude.ai/code/session_01WmBwEiWPff9JZPd5BSGNeH)_ Co-authored-by: Claude <noreply@anthropic.com>
1 parent ad067ad commit a6a1de4

3 files changed

Lines changed: 305 additions & 17 deletions

File tree

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,16 @@
1+
---
2+
'@objectstack/plugin-audit': patch
3+
---
4+
5+
**The read-audit failure report now speaks once per CAUSE instead of once per PROCESS, and prints the telemetry-datasource remedy only for the cause it is the remedy for.**
6+
7+
`installReadAuditWriter`'s `reportReadAuditWriteFailure` (`read-audit.ts`) carried its own process-level `failureReported` boolean and its own fixed message literal — the third independent copy of the pair #15166 fixed in `audit-writers.ts` and #17452 fixed in `auth-event-audit.ts`. Both defects were live on a seam the repo has already declared durability-critical (`persistReadAuditRows` is registered in `DURABILITY_CRITICAL_CALLEES`):
8+
9+
- **The first failure of any cause silenced every later failure of every other cause for the life of the process.** A server could keep losing record-view batches for hours to a second, unrelated fault with one `error` line at the top of the log describing the first — and record-view rows are written from a buffer off the request path, so no in-flight request is left to notice. The dedupe key is now the failure's identity, `auditFailureCauseKey`, imported from `audit-writers.ts` rather than re-spelled. A repeat of an already-reported cause still degrades to `debug`; a NEW cause gets its own `error` line, once.
10+
- **The ADR-0057 §3.6 / `OS_TELEMETRY_DB` datasource guidance printed unconditionally**, so a fault with nothing to do with datasource routing (an `ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED` refusal, say) sent the operator to check something that was working. The guidance is not deleted and not weakened — it is asked for through the shared `isMissingTableError` predicate and printed for exactly the missing-table cause it was written for; every other cause now gets the driver's own verdict quoted at the head of the line plus the fix that matches it.
11+
12+
**Behaviour that deliberately does not change:** the once-per-degradation anti-noise rule itself (a repeat of the same cause is still one line), the `error`-then-`warn` sink fallback (#9657), and the rule that an audit failure never reaches the read.
13+
14+
No API, option or type moves; nothing an author writes changes.
15+
16+
Clause-②: no

‎packages/plugins/plugin-audit/src/read-audit.test.ts‎

Lines changed: 192 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -597,3 +597,195 @@ describe('#8992 shutdown drains the tail', () => {
597597
expect(writer.pending()).toBe(0);
598598
});
599599
});
600+
601+
describe('read-audit failure reporting — once per CAUSE, not once per process (#18247)', () => {
602+
interface LogLine {
603+
level: string;
604+
message: string;
605+
meta?: Record<string, any>;
606+
}
607+
608+
const driverError = (message: string, code?: string): Error => {
609+
const e = new Error(message) as Error & { code?: string };
610+
if (code !== undefined) e.code = code;
611+
return e;
612+
};
613+
614+
const NO_SUCH_TABLE = () => driverError('no such table: sys_audit_log', 'SQLITE_ERROR');
615+
const ORG_REQUIRED = () =>
616+
driverError('system write requires an organization', 'ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED');
617+
618+
/**
619+
* A REAL engine whose ledger write fails with a caller-chosen error.
620+
*
621+
* The engine stays real for the reason this file's header gives — the hook
622+
* dispatch and the field-presence probe are the engine's behaviour, not
623+
* ours. Only `insert` is replaced, and only AFTER install, so the probe has
624+
* already run against the real registry exactly as it does in production.
625+
*
626+
* `omitError: true` drops the OPTIONAL `error` sink (#9657) so the degrade
627+
* path can be exercised on its own.
628+
*/
629+
async function makeCauseHarness(
630+
nextError: (n: number) => unknown,
631+
omitError = false,
632+
): Promise<{ view: () => Promise<void>; at: (level: string) => LogLine[] }> {
633+
const engine = await makeEngine();
634+
await engine.insert(
635+
'contact',
636+
{ id: 'c1', full_name: 'Wei Zhang', organization_id: 'org_a' },
637+
{ context: { isSystem: true } },
638+
);
639+
const logs: LogLine[] = [];
640+
const logger: Record<string, unknown> = {
641+
warn(message: string, meta?: Record<string, any>) {
642+
logs.push({ level: 'warn', message, meta });
643+
},
644+
debug(message: string, meta?: Record<string, any>) {
645+
logs.push({ level: 'debug', message, meta });
646+
},
647+
};
648+
if (!omitError) {
649+
logger.error = (message: string, _err?: Error, meta?: Record<string, any>) => {
650+
logs.push({ level: 'error', message, meta });
651+
};
652+
}
653+
const writer = installReadAuditWriter(engine, {
654+
objects: ['contact'],
655+
timers: makeManualTimers(),
656+
logger: logger as any,
657+
})!;
658+
let n = 0;
659+
(engine as any).insert = async () => {
660+
throw nextError(n++);
661+
};
662+
/** One record view, drained immediately — one failed batch per call. */
663+
const view = async (): Promise<void> => {
664+
await engine.findOne('contact', { where: { id: 'c1' }, context: viewerCtx });
665+
await writer.flush();
666+
};
667+
return { view, at: (level: string) => logs.filter((l) => l.level === level) };
668+
}
669+
670+
it('reports a SECOND, DIFFERENT cause at error — a new cause is a new degradation', async () => {
671+
// THE DEFECT. On the process-wide boolean this file carried, a second,
672+
// unrelated fault produced one `debug` and no `error` at all: the first
673+
// cause of the process had silenced every later cause for the life of the
674+
// process, on a seam `DURABILITY_CRITICAL_CALLEES` already names.
675+
let phase = 0;
676+
const { view, at } = await makeCauseHarness(() => (phase === 0 ? NO_SUCH_TABLE() : ORG_REQUIRED()));
677+
678+
await view();
679+
phase = 1;
680+
await view();
681+
682+
const errors = at('error');
683+
expect(errors).toHaveLength(2);
684+
expect(errors[0].message).toMatch(/no such table: sys_audit_log/);
685+
expect(errors[1].message).toMatch(/ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED/);
686+
expect(at('debug')).toHaveLength(0);
687+
expect(at('warn')).toEqual([]);
688+
});
689+
690+
it('still degrades a REPEAT of an already-reported cause to debug', async () => {
691+
// ⚠️ THE DISCRIMINATING CONTROL. Deleting the boolean outright would also
692+
// make the test above pass, and would be the falsifier AGENTS.md names
693+
// ("log every failure at `error`"). This is the half that must NOT change:
694+
// the same cause on the same ledger still says it once.
695+
const { view, at } = await makeCauseHarness(() => NO_SUCH_TABLE());
696+
697+
for (let i = 0; i < 5; i += 1) await view();
698+
699+
expect(at('error')).toHaveLength(1);
700+
expect(at('debug')).toHaveLength(4);
701+
// The repeats name the cause they were folded into, so a `debug` sweep can
702+
// tell "the same fault, 4 more times" from "four different faults".
703+
expect(at('debug')[0].meta?.cause).toBe(at('debug')[3].meta?.cause);
704+
});
705+
706+
it('keys on the error CODE, never its message, so a per-row fault cannot flood `error`', async () => {
707+
// ⚠️ THE ANTI-NOISE PIN (AGENTS.md; #4420). A driver names the offending
708+
// ROW in its message, so a message-keyed dedupe would grow one `error`
709+
// line per lost batch. 200 batches, 200 distinct messages, ONE code ⇒ one
710+
// line.
711+
const BATCHES = 200;
712+
const { view, at } = await makeCauseHarness((i) =>
713+
driverError(`UNIQUE constraint failed: sys_audit_log.id (row aud_${i})`, 'SQLITE_CONSTRAINT_UNIQUE'),
714+
);
715+
716+
for (let i = 0; i < BATCHES; i += 1) await view();
717+
718+
expect(at('error')).toHaveLength(1);
719+
expect(at('debug')).toHaveLength(BATCHES - 1);
720+
});
721+
722+
it('folds a fault carrying NO code into ONE bucket rather than growing one', async () => {
723+
// The other half of the bound: "the code, or its ABSENCE" is a single key
724+
// value, so an uncoded driver — the shape with nothing bounded to key on —
725+
// still says it once instead of once per flush.
726+
const BATCHES = 200;
727+
const { view, at } = await makeCauseHarness((i) => driverError(`insert failed for batch ${i}`));
728+
729+
for (let i = 0; i < BATCHES; i += 1) await view();
730+
731+
expect(at('error')).toHaveLength(1);
732+
expect(at('debug')).toHaveLength(BATCHES - 1);
733+
});
734+
735+
it('carries the underlying code and message in the first line it prints', async () => {
736+
// The information was computed one line above the branch and dropped on
737+
// the floor: the `error` path built a fixed string and never read `err`.
738+
const { view, at } = await makeCauseHarness(() => ORG_REQUIRED());
739+
740+
await view();
741+
742+
const msg = at('error')[0].message;
743+
expect(msg).toMatch(/ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED/);
744+
expect(msg).toMatch(/system write requires an organization/);
745+
// The consequence half is unchanged — it is still owed, and still first.
746+
expect(msg).toMatch(/compliance trail is now INCOMPLETE/);
747+
expect(msg).toMatch(/who viewed this record/);
748+
});
749+
750+
it('prints the datasource remedy for the cause it is the remedy FOR, and not for others', async () => {
751+
// ⛔ Not a deletion: the ADR-0057 §3.6 routing text is genuinely correct
752+
// for the "no such table" cause it was written for, so it must still print
753+
// there. What is fixed is that it used to print for EVERY cause.
754+
const missing = await makeCauseHarness(() => NO_SUCH_TABLE());
755+
await missing.view();
756+
const forMissingTable = missing.at('error')[0].message;
757+
expect(forMissingTable).toMatch(/telemetry/);
758+
expect(forMissingTable).toMatch(/OS_TELEMETRY_DB=0/);
759+
760+
// The misdirection: the cause is an organization refusal and the text sent
761+
// the operator to check a datasource that was working.
762+
const refused = await makeCauseHarness(() => ORG_REQUIRED());
763+
await refused.view();
764+
const forRefusal = refused.at('error')[0].message;
765+
expect(forRefusal).not.toMatch(/OS_TELEMETRY_DB/);
766+
expect(forRefusal).not.toMatch(/telemetry/i);
767+
// It still owes a fix — it just owes the RIGHT one.
768+
expect(forRefusal).toMatch(/Fix:/);
769+
});
770+
771+
it('[#9657] the `warn` fallback is per-cause too — a sink with no `error` hears the second fault', async () => {
772+
// `ReadAuditLogger.error` is OPTIONAL, so the degrade path is the only
773+
// channel a host without one ever gets. Fixing the dedupe on the `error`
774+
// branch alone would leave that host exactly where it started.
775+
let phase = 0;
776+
const { view, at } = await makeCauseHarness(
777+
() => (phase === 0 ? NO_SUCH_TABLE() : ORG_REQUIRED()),
778+
true,
779+
);
780+
781+
await view();
782+
phase = 1;
783+
await view();
784+
785+
const warns = at('warn');
786+
expect(warns).toHaveLength(2);
787+
expect(warns[0].message).toMatch(/no such table: sys_audit_log/);
788+
expect(warns[1].message).toMatch(/ERR_SYSTEM_WRITE_ORGANIZATION_REQUIRED/);
789+
expect(at('error')).toEqual([]);
790+
});
791+
});

‎packages/plugins/plugin-audit/src/read-audit.ts‎

Lines changed: 97 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -88,12 +88,26 @@
8888

8989
import type { HookContext } from '@objectstack/spec/data';
9090
import type { IDataEngine } from '@objectstack/spec/contracts';
91+
import { isMissingTableError } from '@objectstack/types';
9192
// DERIVED, never re-typed — the same rule `audit-writers.ts` states for its own
9293
// two faces. An object excluded from write auditing (recursion, auth/session
9394
// noise, ADR-0057 telemetry plumbing) is excluded from read auditing for the
9495
// identical reasons, and a second hand-kept list would disagree on the day
9596
// either is fixed.
96-
import { AUDIT_EXCLUDED_OBJECTS, createFieldPresenceProbe } from './audit-writers.js';
97+
//
98+
// ⚠️ [#18247] `auditFailureCauseKey` / `auditFailureCauseSummary` join that
99+
// derivation for the same reason and on the same seam. #15166 replaced a
100+
// process-wide boolean with a cause key in `audit-writers.ts`; #17452 found a
101+
// SECOND copy of the boolean in `auth-event-audit.ts` and imported the key
102+
// rather than re-spelling it, and exported it for exactly that purpose. This
103+
// file was the THIRD copy. ⛔ A third spelling of the key is how the defect
104+
// travelled the first two times — import it.
105+
import {
106+
AUDIT_EXCLUDED_OBJECTS,
107+
auditFailureCauseKey,
108+
auditFailureCauseSummary,
109+
createFieldPresenceProbe,
110+
} from './audit-writers.js';
97111

98112
/**
99113
* The ledger action this writer emits.
@@ -481,28 +495,94 @@ export function installReadAuditWriter(
481495
);
482496
};
483497

484-
let failureReported = false;
498+
/**
499+
* Report a lost batch of record-view rows — once per CAUSE, not once per
500+
* failed flush.
501+
*
502+
* Same discipline, and the same reason, as `reportAuditWriteFailure` in
503+
* `audit-writers.ts` and `reportAuthEventWriteFailure` in
504+
* `auth-event-audit.ts`: a systemic cause (the ledger is unreachable from
505+
* this connection) would otherwise emit one `error` per flush and train
506+
* everyone to skim the channel. That much is unchanged, and ⛔ must stay —
507+
* AGENTS.md records the once-per-degradation rule as a deliberate anti-noise
508+
* choice and names 「log every failure at `error`」 as its falsifier.
509+
*
510+
* [#18247] What changed is the COUNTING UNIT, and it changed here for the
511+
* THIRD time in this package: this file carried its OWN copy of the
512+
* process-wide boolean and its OWN fixed message literal, so neither
513+
* #15166's fix to `audit-writers.ts` nor #17452's to `auth-event-audit.ts`
514+
* reached it. All three copies had the same two defects.
515+
*
516+
* 1. One process-wide boolean means the first failure of ANY cause silences
517+
* every later failure of every OTHER cause for the life of the process.
518+
* The rule's unit is a DEGRADATION and a second cause is a second
519+
* degradation, so the key is now the failure's identity —
520+
* {@link auditFailureCauseKey}, imported rather than re-spelled. A repeat
521+
* of an already-reported cause still degrades to `debug`, exactly as
522+
* before; a NEW cause gets its own `error` line, once.
523+
* 2. The fixed literal printed the ADR-0057 §3.6 telemetry-datasource
524+
* remedy for every cause, so a fault that had nothing to do with
525+
* datasource routing sent its operator to check something that was
526+
* working. ⛔ The guidance is not deleted and not weakened — it is the
527+
* right remedy for the missing-table cause it was written for, and is now
528+
* printed for exactly that cause, asked through the shared
529+
* `isMissingTableError` predicate.
530+
*
531+
* ⭐ Why this seam is not merely the third repetition: `persistReadAuditRows`
532+
* is registered in `DURABILITY_CRITICAL_CALLEES`
533+
* (`scripts/check-durability-degradation-log-level.mjs`), so the repo has
534+
* already declared this write durability-critical. A reporter that switches
535+
* itself off after one cause is exactly the failure that declaration cannot
536+
* afford — the board's green looks identical to a real one.
537+
*
538+
* ⛔ The key is built from the error's `code`, NEVER its message — see
539+
* {@link auditFailureCauseKey} for why that is what keeps the cause set
540+
* bounded by boot-declared vocabularies instead of by traffic.
541+
*/
542+
const reportedReadAuditFailureCauses = new Set<string>();
485543
const reportReadAuditWriteFailure = (count: number, err: unknown): void => {
486544
const detail = String((err as any)?.message ?? err);
487545
try {
488-
if (failureReported) {
489-
logger?.debug?.('Read-audit write failed (already reported)', { count, err: detail });
546+
// The object dimension of the shared key is the LEDGER here, not the
547+
// object whose record was viewed. ⛔ Not an arbitrary choice between the
548+
// two: one flush is ONE write carrying rows about MANY audited objects,
549+
// so there is no single viewed object to name — picking whichever landed
550+
// first in the batch would make the key depend on traffic, which is the
551+
// one property {@link auditFailureCauseKey} exists to deny. The failure
552+
// is a property of the write against `sys_audit_log`, so that is the
553+
// object named, and the key reduces to the driver's code vocabulary:
554+
// bounded by construction, and still the same key shape rather than a
555+
// second one. (`audit-writers.ts` passes `ctx.object` and
556+
// `auth-event-audit.ts` its constant `sys_session` because on those two
557+
// seams the viewed/audited object IS single-valued per write.)
558+
const cause = auditFailureCauseKey('sys_audit_log', err);
559+
if (reportedReadAuditFailureCauses.has(cause)) {
560+
logger?.debug?.('Read-audit write failed (already reported)', { count, err: detail, cause });
490561
return;
491562
}
492-
failureReported = true;
563+
reportedReadAuditFailureCauses.add(cause);
564+
// `persistReadAuditRows` writes ONE table, so the missing-table question
565+
// is asked about that one — unlike `persistAuditTrailRow`, which writes
566+
// the ledger row and its `sys_activity` mirror and asks about both.
567+
const missingTable = isMissingTableError(err, 'sys_audit_log');
493568
const message =
494-
`Read-audit write FAILED — ${count} record-view row(s) were LOST and the compliance trail is now ` +
495-
'INCOMPLETE. The reads themselves SUCCEEDED and returned 200, so the API, the screens and every ' +
496-
'counter read clean; only the `sys_audit_log` rows recording WHO opened those records never landed, ' +
497-
'and nothing retries them. Every subsequent batch is likely lost the same way (this is reported ONCE ' +
498-
'— raise the log level to `debug` to see the rest). The whole point of this capability is answering ' +
499-
'"who viewed this record" for an auditor, so the failure mode is a query that returns a confident, ' +
500-
'wrong, SHORT answer. Fix: confirm `sys_audit_log` is reachable from the connection this write ran ' +
501-
'on — its ADR-0057 §3.6 lifecycle class routes it to the dedicated `telemetry` datasource whenever ' +
502-
'one is registered (`os dev` provisions one by default as a SIBLING SQLite file), so a "no such ' +
503-
'table" here usually means the write executed against a DIFFERENT datasource than the one the table ' +
504-
'was created in. Set `OS_TELEMETRY_DB=0` to keep every lifecycle-classed object on the primary ' +
505-
'datasource.';
569+
`Read-audit write FAILED (${auditFailureCauseSummary(err, detail)}) — ${count} record-view row(s) ` +
570+
'were LOST and the compliance trail is now INCOMPLETE. The reads themselves SUCCEEDED and returned ' +
571+
'200, so the API, the screens and every counter read clean; only the `sys_audit_log` rows recording ' +
572+
'WHO opened those records never landed, and nothing retries them. Every subsequent batch failing ' +
573+
'THIS WAY is lost the same way (this CAUSE is reported ONCE — raise the log level to `debug` to see ' +
574+
'the rest; a DIFFERENT cause gets its own `error` line). The whole point of this capability is ' +
575+
'answering "who viewed this record" for an auditor, so the failure mode is a query that returns a ' +
576+
'confident, wrong, SHORT answer. ' +
577+
(missingTable
578+
? 'Fix: confirm `sys_audit_log` is reachable from the connection this write ran on — its ADR-0057 ' +
579+
'§3.6 lifecycle class routes it to the dedicated `telemetry` datasource whenever one is ' +
580+
'registered (`os dev` provisions one by default as a SIBLING SQLite file), so a "no such table" ' +
581+
'here usually means the write executed against a DIFFERENT datasource than the one the table was ' +
582+
'created in. Set `OS_TELEMETRY_DB=0` to keep every lifecycle-classed object on the primary ' +
583+
'datasource.'
584+
: 'Fix: resolve the driver fault quoted at the head of this line on the connection this write ran ' +
585+
'on — every batch that hits it loses its rows until it is resolved.');
506586
// `error` is OPTIONAL on this sink, so `logger?.error?.(…)` printed
507587
// NOTHING when the host injected one without it — the durability
508588
// degradation this text describes would then be reported by nobody at

0 commit comments

Comments
 (0)