Skip to content

[LightRAG][Step 7] Prototype Knowledge Demand Profiles and knowledge-compute co-routing #169

Description

@rickisba

Parent Epic: #162
Depends on: #168

Goal

Build a research prototype that converts graph-augmented retrieval results into a structured Knowledge Demand Profile (KDP) and evaluates whether this representation improves cross-pool knowledge and compute routing.

This step is research-oriented. It must not be merged into the default production path without measured evidence and an explicit review of the algorithm, failure modes, and observability.

Research Hypothesis

A flat Top-K knowledge list does not express which knowledge is necessary, optional, redundant, mutually dependent, or expensive to serve. A structured demand profile may allow CacheRoute to jointly select a sufficient knowledge subset, KDN, Proxy, and injection policy under quality and latency constraints.

Proposed KDP

Define a versioned, bounded representation containing fields such as:

  • query/retrieval type and confidence;
  • canonical candidate KIDs;
  • relevance and provenance;
  • required/optional role or necessity estimate;
  • token/byte cost;
  • entity and relationship coverage summaries;
  • bounded dependency edges between canonical KIDs;
  • KDN replica/KV availability annotations;
  • uncertainty and missing-data markers.

LightRAG supplies semantic entities, relationships, chunks, and retrieval evidence. KDN and CacheRoute supply canonical identity, replica location, KV readiness, token/byte cost, network state, and pool state.

Scope

  • Define KnowledgeDemandProfile schema and versioning with strict size bounds.
  • Design deterministic mapping and aggregation from LightRAG document/chunk/entity/relationship results to canonical KIDs.
  • Separate observed, measured, inferred, predicted, and unknown values.
  • Implement one or more interpretable baseline policies for:
    • knowledge subset selection;
    • KDN selection;
    • Proxy selection;
    • text/KV injection recommendation or hint.
  • Preserve the production hierarchy: Scheduler produces pool-level decisions; Proxy remains authoritative for concrete Instance selection and final executable injection behavior.
  • Provide a quality-constrained objective or utility model combining retrieval/answer quality with retrieval latency, network transfer, injected tokens, prefill cost, queue cost, KV reuse, and SLO risk.
  • Add offline replay or simulation support before online activation.
  • Add ablation modes: flat embedding Top-K, LightRAG flat Top-K, graph-aware subset without system cost, system-cost-aware routing without graph structure, and full KDP co-routing.
  • Add trace fields sufficient to explain each decision and attribute predicted versus actual cost.

Evaluation Requirements

Evaluate both semantic and system effects:

  • Recall@K, MRR, NDCG, or task-specific answer quality using labeled/reproducible data;
  • minimum sufficient knowledge coverage or an explicitly defined proxy metric;
  • redundant injected tokens;
  • TTFT and E2E latency;
  • SLO attainment;
  • cross-KDN bytes and transfer time;
  • KDN queue time;
  • text/KV selection and KV reuse;
  • prediction error for cost estimates;
  • scheduling overhead;
  • behavior under KDN/Proxy load shifts, stale indexes, missing KV artifacts, and LightRAG degradation.

Do not claim quality preservation from retrieval metrics alone. The evaluation must include final answer quality or a defensible task-level metric.

Safety and Compatibility Rules

  • KDP is optional and versioned.
  • Unknown or missing semantic/system fields must not be interpreted as positive support.
  • The embedding backend and non-KDP routing remain available as baselines and fallbacks.
  • KDP payloads are bounded and do not contain full documents, credentials, embeddings, KV bytes, or LMCache-private objects.
  • No new functional root directory is allowed without separate architecture review.
  • Do not replace existing README content; publish research design and methodology in dedicated documentation.

Non-Goals

  • Do not replace Proxy Instance selection with LightRAG.
  • Do not train an opaque end-to-end policy before interpretable baselines are established.
  • Do not make inferred graph relationships authoritative for artifact existence or compatibility.
  • Do not make the research policy the default production strategy in this Issue.

Acceptance Criteria

  • A versioned KDP schema and clear provenance model exist.
  • LightRAG outputs are mapped to canonical KIDs with bounded dependency information.
  • At least one interpretable knowledge-compute co-routing baseline is implemented.
  • Offline replay/simulation compares all required ablations.
  • Experiments report semantic quality and system performance together.
  • Failure/degradation experiments demonstrate safe fallback.
  • CPU-only unit tests cover schema, aggregation, cost calculation, and policy behavior using synthetic data.
  • Any online experiment is gated behind explicit configuration and documented as experimental.

Validation

Run schema/policy unit tests, offline replay experiments, and controlled end-to-end evaluation. The PR must report datasets/workloads, environment, commands, raw result locations, limitations, and use Closes #169.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions