Maintainer & Security Contact: Eshwar Desetty (eshwar.desetty03@gmail.com)
Can we trust an AI model with critical financial decisions?
As security professionals, we are trained to never trust a "black box." Yet, we often deploy LLMs without seeing what's happening inside. This project challenges that norm.
We are not just predicting housing prices; we are auditing the "brain" of Llama-3.3-70b. By attaching a forensic "flight recorder" (OpenTelemetry), we capture every thought, every token, and every latency spike to answer the ultimate security question: Is this model hallucinating, or is it reliable?
This project demonstrates a complete AI Observability & Integrity Audit pipeline. We treat the LLM as a suspect and the traces as our evidence.
By instrumenting the model with Arize Phoenix, we gained X-Ray vision into its decision-making process.
Figure 1: The Red Line (AI) attempts to track the Blue Line (Reality). The gaps represent model hallucinations and drift.
This project implements core AI Security (AISec) concepts:
We utilize Arize Phoenix to capture full traces of every execution. In a security context, this is critical for:
- Incident Response: Reconstructing exactly what input caused a harmful output.
- Prompt Injection Detection: Analyzing raw inputs to identify adversarial patterns.
- Data Leakage Auditing: Verifying that model responses do not contain PII (Personally Identifiable Information).
The performance_curve.html visualization serves as an Integrity Audit. Significant deviations between AI predictions and ground truth indicate:
- Model Drift: The model losing alignment with reality.
- Hallucinations: Fabrication of facts, posing a reliability and integrity risk.
- Adversarial Susceptibility: Identifying inputs that cause the model to fail catastrophically.
The run_experiment.py script demonstrates an automated framework for Batch Auditing. This same infrastructure can be repurposed to:
- Batch-test the model against Jailbreak Prompts.
- Verify compliance with Safety Policies.
- Stress-test the model for Denial of Service (DoS) resilience.
- File:
data/RA_Application_Task.csv - Scope: 38 records used as the "Ground Truth" for the integrity audit.
- Model: llama-3.3-70b-versatile (Groq)
- Script:
run_experiment.py - Action: 38 AI-generated valuations ($180K-$275K range).
- Visualization:
outputs/performance_curve.html - Data:
outputs/experiment_results.csv - Finding: Differences range from -$98K to +$93K, highlighting specific areas of model uncertainty.
- Tool: Phoenix Dashboard (http://localhost:6006/)
- Status: ✅ 38/38 Traces captured. Full visibility into latency, token usage, and prompt chains.
This repository does NOT include API keys. You must create your own .env file locally:
- Copy the template:
cp .env.example .env - Add your Groq API key to
.env - Get your key from: https://console.groq.com/keys
Note: The .env file is gitignored and will never be committed to this repository.
# Use Python 3.9 or higher (required for arize-phoenix)
python3.9 -m venv .venv39
source .venv39/bin/activate # On Windows: .venv39\Scripts\activatepip install --upgrade pip
pip install -r requirements.txt# Copy the example file
cp .env.example .env
# Edit .env and add your Groq API key
# Get your key from: https://console.groq.com/keys# Run the experiment
.venv39/bin/python run_experiment.py
# View Phoenix Dashboard (Forensics)
# Open: http://localhost:6006/
# View Integrity Report
open outputs/performance_curve.html
cat outputs/experiment_results.csv| File | Purpose | Security Context |
|---|---|---|
TASK_SUMMARY.md |
Complete project report | |
run_experiment.py |
Main script | The Audit Engine |
outputs/performance_curve.html |
Visualization | Integrity Report |
outputs/experiment_results.csv |
All results with AI estimates | |
notebook/experiment_v1.ipynb |
Jupyter notebook version | |
.env.example |
Config Template | Secret Management |
requirements.txt |
Python dependencies | |
SECURITY.md |
Security Policy | Best Practices |
This audit reveals a critical security insight: AI models are confident, but not always correct.
Our "Integrity Gap" analysis shows that while the model's reasoning often sounds plausible, the actual output can deviate significantly from reality (up to $98K in our test). For a financial application, this variance is a high-risk vulnerability.
Key Takeaways for Security Professionals:
- Trust but Verify: Never deploy an LLM without an observability layer (like Arize Phoenix) to audit its actual behavior.
- The Black Box Risk: Without tracing, you are blind to "silent failures" where the model hallucinates convincingly.
- Continuous Auditing: Security is not a one-time check. Automated pipelines like this one are essential to detect model drift and new failure modes over time.
Final Verdict: The model is a powerful tool, but it requires a "human-in-the-loop" or strict guardrails for high-stakes decision-making.
Audit Status: ✅ COMPLETE - All telemetry captured and integrity verified.
