Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -291,6 +291,24 @@ ORCAROUTER_ENDPOINT=https://api.orcarouter.ai/v1/chat/completions
# DESCRIPTION: Per-request timeout in ms.
# CODEX_TIMEOUT=120000

# ------------------------------------------------------------------------------
# Cursor (uses your Cursor Pro subscription via local cursor-agent CLI)
# ------------------------------------------------------------------------------
# DESCRIPTION: Enable the Cursor local provider (requires `cursor-agent` CLI installed + `agent login`).
# Values: true | false
# CURSOR_ENABLED=true
# DESCRIPTION: Cursor model id (run `cursor-agent models` to list).
# CURSOR_MODEL=composer-2.5
# DESCRIPTION: Path to the `cursor-agent` binary; auto-detected if unset.
# CURSOR_BINARY_PATH=cursor-agent
# DESCRIPTION: Per-request timeout in ms.
# CURSOR_TIMEOUT=120000
# DESCRIPTION: Auto-approve ALL cursor-agent permission requests, shell execution
# included (passes --force). The sandbox cwd is not a security boundary while this
# is on. Set false to approve MCPs only and reject edit/execute/delete requests.
# Values: true | false
# CURSOR_AUTO_APPROVE=true

# ------------------------------------------------------------------------------
# Embeddings provider override
# ------------------------------------------------------------------------------
Expand Down
12 changes: 12 additions & 0 deletions config/difficulty-anchors.json
Original file line number Diff line number Diff line change
Expand Up @@ -28,5 +28,17 @@
"design a horizontally scalable architecture for this service with failure-mode analysis",
"Compare how databricks.js and openai-format.js each handle non-2xx upstream responses, then propose a unified error-normalization helper both could share",
"Code review the PR #84 routing hardening changes"
],
"frontier": [
Comment on lines +31 to +32

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
Added "frontier" key to config/difficulty-anchors.json. Please verify the key name is intentionally "frontier" and not a typo (e.g., "frontier-level" or "frontier-tier") to maintain consistency with existing keys like "design", "code", etc., or ensure it aligns with the intended schema.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intentional. frontier is the anchor class src/routing/intent-score.js reads (sims.frontier, FRONTIER_MIN_SIM, the REASONING band) alongside trivial/substantive/heavyweight. No change.

"Prove the correctness of this lock-free queue implementation and identify any ABA hazards",
"Design a novel conflict-free replicated data type for collaborative rich-text editing and argue its convergence from first principles",
"Derive the optimal cache eviction policy for this access distribution and prove its competitive ratio",
"Diagnose this heisenbug: a race that only reproduces under load across three services, given these interleaved logs",
"Design a Byzantine-fault-tolerant consensus protocol variant that tolerates f faults with 2f+1 replicas by weakening liveness, and analyze exactly which guarantees are lost",
"Formally verify that this state machine can never deadlock, or produce a counterexample trace",
"Given these profiler traces, determine whether the tail latency is queueing-theoretic or GC-driven, and prove which by constructing a discriminating experiment",
"Reason through the security of this key-rotation scheme under an adversary who can observe but not modify traffic, and find the weakest assumption it relies on",
"Work out the exact memory ordering constraints this concurrent hashmap needs on ARM, and justify each barrier from the C++ memory model",
"Given this failing distributed transaction trace, determine whether the anomaly is write skew or lost update, and design the minimal isolation-level change that eliminates it"
]
}
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@
"start:supervised": "while true; do node index.js 2>&1 | npx pino-pretty --sync; code=$?; echo \"[supervisor] lynkr exited ($code) — restarting in 3s\"; sleep 3; done",
"lint": "eslint src index.js",
"test": "npm run test:unit && npm run test:performance",
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js",
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js",
"test:memory": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com node --test test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js",
"test:new-features": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com node --test test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js",
"test:performance": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com node test/hybrid-routing-performance.test.js && DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com node test/performance-tests.js",
Expand Down
46 changes: 45 additions & 1 deletion src/api/router.js
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ const providersRouter = require("./providers-handler");
const claudeDesktopGatewayRouter = require("./claude-desktop-gateway");
const { getRoutingHeaders, getRoutingStats, analyzeComplexity, getModelTierSelector, analyzeRisk, checkSessionPin, writeSessionPin, checkPinScoreDrift } = require("../routing");
const { resolveTierForModelId } = require("../routing/model-slots");
const { stripHarnessEnvelope } = require("../routing/harness-envelope");

// Upstream streams can die without a clean end (reader.read() never
// resolves on a dropped socket), hanging the client forever. Every
Expand Down Expand Up @@ -1409,7 +1410,11 @@ router.post("/v1/messages", rateLimiter, async (req, res, next) => {
}
return '';
})();
// Suggestion-mode detection reads the RAW text (the marker is itself
// wrapper text); the force/risk probes below get envelope-stripped text
// so Cursor's <user_info>/<rules> blocks can't fire triggers on "Hi".
const isSuggestionMode = _lastUserText.includes('[SUGGESTION MODE:');
const _lastUserAskClean = stripHarnessEnvelope(_lastUserText);
// Tool-lessness alone is NOT harness evidence: generic API clients
// (curl, benchmarks, SDKs) legitimately send bare messages and must get
// full routing — live regression 2026-07-08: a benchmark's security-
Expand Down Expand Up @@ -1535,7 +1540,7 @@ router.post("/v1/messages", rateLimiter, async (req, res, next) => {
let _pinForceBypass = null;
if (!sideTier && pinCheck.serve && pinCheck.reason === 'guards_passed' && !isSideRequest) {
try {
const _probe = { messages: [{ role: 'user', content: _lastUserText || '' }] };
const _probe = { messages: [{ role: 'user', content: _lastUserAskClean || '' }] };
const ca = require("../routing/complexity-analyzer");
if (_lastUserText && ca.shouldForceReasoning(_probe)) _pinForceBypass = 'force_reasoning';
else if (_lastUserText && ca.shouldForceCloud(_probe)) _pinForceBypass = 'force_cloud';
Expand Down Expand Up @@ -1883,6 +1888,15 @@ router.post("/v1/messages", rateLimiter, async (req, res, next) => {
tierMethod: tier?.method || null,
tierPinned: tier?.pinned ?? null,
cacheState: _pinCacheState,
// Mid tool-exchange frames must keep serving the in-flight model
// (same invariant as the orchestrator's tool_history pin serves).
hasToolHistory: (() => {
try { return require("../routing/session-affinity").payloadHasToolHistory(req.body); }
catch { return false; }
})(),
// This branch IS the flat-fee subscription: the gate's dollar
// break-even leg has no premium to amortize here.
flatRate: true,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Hardcoded 'true' for flatRate may be misleading; ensure this is always accurate for passthrough routes or derive it from tier configuration.

Suggestion:

Suggested change
flatRate: true,
flatRate: tier?.flatRate ?? true,

// Quota pressure shortens the hold horizon inside the gate (high
// sustained burn → descents clear sooner). Zero when the ledger has
// no history — the gate behaves exactly as before.
Expand Down Expand Up @@ -1919,6 +1933,36 @@ router.post("/v1/messages", rateLimiter, async (req, res, next) => {
logger.debug({ err: err.message }, '[Routing] Pin-hold restore failed (non-fatal)');
}
}
// Upgrades must persist the SERVED model into the pin, mirroring the
// pin_hold restore above — the write at the fresh-decision site stored
// the pre-rewrite tier model, and a pin that forgets the upgrade lets
// later frames descend below what actually served. Unlike pin_hold,
// an upgrade usually happens with no prior pin, so the payload comes
// from the fresh tier object. Floored (+continuation_inherit) and
// risk-lifted tiers stay unpinned (writeSessionPin re-checks, but the
// method is overridden to 'session_pin' here, so guard at the source).
if (
_passthroughRoute.action === 'upgrade'
&& _passthroughRoute.model
&& pinCheck?.sessionId
&& tier?.provider
&& !String(tier?.method || '').includes('+continuation_inherit')
&& tier?.escalation_source !== 'risk'
) {
try {
const { writeSessionPin } = require("../routing/index");
writeSessionPin(pinCheck.sessionId, {
provider: tier.provider,
model: _passthroughRoute.model,
tier: tier.tier,
score: tier.score ?? null,
method: 'session_pin',
reason: 'passthrough_upgrade',
}, req.body);
} catch (err) {
logger.debug({ err: err.message }, '[Routing] Upgrade pin persist failed (non-fatal)');
}
}
if (_passthroughRoute.action !== 'verbatim' && _passthroughRoute.model) {
logger.info({
from: _passthroughClientModel,
Expand Down
33 changes: 30 additions & 3 deletions src/cache/embeddings.js
Original file line number Diff line number Diff line change
Expand Up @@ -145,6 +145,9 @@ const STRICT = process.env.LYNKR_EMBEDDINGS_STRICT === 'true';
function _noteFallback(providerName, err) {
fallbackCount += 1;
lastProviderError = err?.message || String(err);
// Cooldown counts from the LAST failed dial, not the first attempt's start —
// the in-place retry delay must not eat into the cooldown window.
lastProviderAttempt = Date.now();
if (embeddingProviderAvailable !== false) {
embeddingProviderAvailable = false;
degradedSince = Date.now();
Expand All @@ -168,6 +171,14 @@ function _noteRecovery(providerName) {
degradedSince = null;
}

// One in-place retry before tripping degradation: a single transient failure
// (Ollama cold-loading the model at boot, a blip mid-restart) was flipping
// the WHOLE provider to hash embeddings for a full cooldown after every
// server start (recurring live incident, last 2026-09-26). A short backoff
// absorbs the blip; a provider that is actually down fails twice and
// degrades exactly as before.
const TRANSIENT_RETRY_DELAY_MS = 1500;

function _wrapProvider(providerName, providerFn) {
return async (text) => {
// While degraded, only re-attempt the provider after the cooldown; serve
Expand All @@ -184,9 +195,25 @@ function _wrapProvider(providerName, providerFn) {
const result = await providerFn(text);
_noteRecovery(providerName);
return result;
} catch (err) {
_noteFallback(providerName, err);
if (STRICT) throw err;
} catch (firstErr) {
// Already-degraded providers get no second chance (this attempt WAS the
// post-cooldown probe); healthy ones earn one retry before the flip.
if (embeddingProviderAvailable !== false) {
await new Promise((r) => setTimeout(r, TRANSIENT_RETRY_DELAY_MS));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The retry delay uses Date.now() for timing. If the system clock is adjusted (e.g., NTP sync), the retry may execute sooner or later than expected. For production-grade code, consider using a monotonic clock or document this behavior.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Leaving as-is. The retry delay itself is a setTimeout; only the degraded-provider cooldown gate reads Date.now(), and a clock step there just shifts one re-probe slightly earlier or later, with no correctness impact.

Comment on lines +201 to +202

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The code uses a Promise-based setTimeout (await new Promise((r) => setTimeout(r, TRANSIENT_RETRY_DELAY_MS))) but the project prefers async/await patterns. While function is not a security issue, it's inconsistent with async best practices.

Suggestion: Use a helper function like const delay = (ms) => new Promise((r) => setTimeout(r, ms)); and then call await delay(TRANSIENT_RETRY_DELAY_MS);

Suggestion:

Suggested change
if (embeddingProviderAvailable !== false) {
await new Promise((r) => setTimeout(r, TRANSIENT_RETRY_DELAY_MS));
function delay(ms) {
return new Promise((r) => setTimeout(r, ms));
}
// In _wrapProvider:
await delay(TRANSIENT_RETRY_DELAY_MS);

try {
const result = await providerFn(text);
_noteRecovery(providerName);
logger.debug({ provider: providerName, error: firstErr?.message },
'[Embeddings] Transient provider failure absorbed by retry');
return result;
} catch (secondErr) {
_noteFallback(providerName, secondErr);
if (STRICT) throw secondErr;
return generateHashEmbedding(text);
}
}
_noteFallback(providerName, firstErr);
if (STRICT) throw firstErr;
return generateHashEmbedding(text);
}
};
Expand Down
Loading
Loading