Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ This document specifies mandatory rules, design patterns, prompt caching optimiz

## 0. Source of Truth (Read First)

- **`blueprint.md`** — research design, scientific claims, scope. **Version A (course) is the ONLY active scope.** Do NOT implement Version B (publication) features during the course.
- **`blueprint.md`** — research design, scientific claims, scope. **Version A is frozen on branch `version-A`; Version B (publication) work runs on branch `version-B` only.** Do not mix scopes or alter Version A artifacts after the gate push without a documented reason.
- **`roadmap.md`** — phased implementation plan. Treat as guidance, NOT gospel: version numbers, URLs, and library claims in it can be stale. Always prefer the **latest stable versions actually available/installed** (check `.venv` and PyPI before pinning anything).
- **`AGENTS.md`** (this file) — binding agent rules. On conflict with `roadmap.md`, this file and `blueprint.md` win.

Expand Down Expand Up @@ -127,7 +127,7 @@ HaluRISC/
- **Security Rule:** Server-only variables (`OPENAI_API_KEY`) must NEVER start with `NEXT_PUBLIC_`. They are strictly accessed in `app/api/chat/route.ts` (server side).
- **FastAPI Python Backend Environment:**
- **Location:** Root `.env` or system environment variables loaded via `python-dotenv`.
- **Keys:** `FASTAPI_HOST`, `FASTAPI_PORT`, `FASTAPI_DEBUG`, `OPENAI_API_KEY` (for `/judge`), `OPENAI_MODEL`, `DEEPSEEK_API_KEY` (optional fallback judge), `HALU_API_DEVICE` (`cuda`|`cpu`, default `cpu`; `cuda` auto-falls back to CPU if torch has no CUDA, models load fp16 on CUDA), `HALU_API_PRELOAD` (default `1`; `0` skips the startup preload of heavy spaCy/NLI/SBERT models).
- **Keys:** `FASTAPI_HOST`, `FASTAPI_PORT`, `FASTAPI_DEBUG`, `OPENAI_API_KEY` (for `/judge`), `OPENAI_MODEL`, `DEEPSEEK_API_KEY` (optional fallback judge), `HALU_API_DEVICE` (`cuda`|`cpu`, default `cpu`; `cuda` auto-falls back to CPU if torch has no CUDA, models load fp16 on CUDA), `HALU_API_PRELOAD` (default `1`; `0` skips the startup preload of heavy spaCy/NLI/SBERT models), `HALU_XGB_DEVICE` (`cuda`|`cpu`|`auto`; set `cpu` in Colab so saved XGBoost models are portable across platforms — CUDA-trained boosters do not unserialize cross-platform).
- **Git Security Rule:** Neither `.env` nor `.env.local` are ever committed to Git (`.gitignore` protects both).

### 5.1 Dependency Pinning Rule
Expand Down
24 changes: 12 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,21 +29,21 @@

Mean over seeds 42/123/456 (real results from `artifacts/results/final_results.json`):

| Model Architecture | Precision | Recall | F1-Score | AUROC | PR-AUC | MCC |
| ------------------------------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- |
| **Heuristic (1 - overlap)** | 0.9392 | 0.9467 | 0.9429 | 0.9148 | 0.8117 | 0.8854 |
| **Logistic Regression** | 0.9804 | 0.9693 | 0.9749 | 0.9943 | 0.9948 | 0.9501 |
| **Random Forest** | 0.9915 | 0.9858 | 0.9886 | 0.9982 | 0.9987 | 0.9774 |
| **XGBoost + Platt (ours)** | **0.9919** | **0.9853** | **0.9886** | **0.9980** | **0.9987** | **0.9774** |
| Model Architecture | Precision | Recall | F1-Score | AUROC | PR-AUC | MCC |
| --------------------------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- |
| **Heuristic (1 - overlap)** | 0.9392 | 0.9467 | 0.9429 | 0.9148 | 0.8117 | 0.8854 |
| **Logistic Regression** | 0.9804 | 0.9693 | 0.9749 | 0.9943 | 0.9948 | 0.9501 |
| **Random Forest** | 0.9915 | 0.9858 | 0.9886 | 0.9982 | 0.9987 | 0.9774 |
| **XGBoost + Platt (ours)** | **0.9919** | **0.9853** | **0.9886** | **0.9980** | **0.9987** | **0.9774** |

**Calibration:** Platt ECE 0.0116 / Brier 0.0092 · Isotonic ECE 0.0051 / Brier 0.0089 (calibrators fit on validation only).

### LLM-as-Judge comparison (200 test samples, measured)

| Model | Accuracy | Precision | Recall | F1 | Latency p50 | Cost / 1K |
| ------------------------ | -------- | --------- | ------ | ------ | ----------- | --------- |
| GPT 5.6 Luna judge | 0.8400 | 0.9474 | 0.7200 | 0.8182 | 1,293 ms | $0.101 |
| **XGBoost (ours)** | **0.9900** | **1.0000** | **0.9800** | **0.9899** | ~5 ms | ~$0.001 |
| Model | Accuracy | Precision | Recall | F1 | Latency p50 | Cost / 1K |
| ------------------ | ---------- | ---------- | ---------- | ---------- | ----------- | --------- |
| GPT 5.6 Luna judge | 0.8400 | 0.9474 | 0.7200 | 0.8182 | 1,293 ms | $0.101 |
| **XGBoost (ours)** | **0.9900** | **1.0000** | **0.9800** | **0.9899** | ~5 ms | ~$0.001 |

Agreement between judge and XGBoost: 0.84.

Expand Down Expand Up @@ -96,7 +96,7 @@ differ from real-world generation, motivating domain adaptation (Version B direc

```powershell
# 1. Clone the repository
git clone https://github.com/Atik203/HaluLens.git
git clone https://github.com/Atik203/HaluRISC.git
cd HaluRISC

# 2. Activate virtual environment (or create one)
Expand Down Expand Up @@ -165,7 +165,7 @@ Open `http://localhost:3000` in your browser.

Run the **full training pipeline** (feature extraction → XGBoost tuning → calibration → SHAP → RAGTruth validation) on Google Colab with a GPU, then download the artifacts back into this repo:

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1aTlrAcIx5FqAaDiYzsLTiMxbMVw1RsRR?usp=sharing)
[![Open In Colab](https://colab.research.google.com/drive/124wjKFVDyZkDNIjs1WgHW8N7vO7XyY6G?usp=sharing)

1. Open the notebook (viewable by anyone with the link), select **GPU → T4** as the runtime.
2. Run cells in order; cell 3 prompts for `colab/halurisc_src.zip` (from this repo).
Expand Down
Loading