Backend for graph-centric fraud analysis and dynamic visualization payloads.
- Generates a large synthetic IEEE-CIS-style dataset using real-column distribution profiles.
- Assumes clean entity IDs (
uid_clean) for initial implementation. - Trains
XGBoostwith asymmetric focal objective. - Builds UID aggregate features and graph-structural embeddings.
- Produces API-ready artifacts for:
- force-directed graph nodes/links
- hop-based expansion
- star/ring pattern endpoints
- transaction explainability payloads (waterfall-ready contributions)
- embedding-space payloads
- entity timelines
docs/foundational_flow_review.md walks the pipeline stage by stage (real data, synthetic
generator, cleaning, embeddings, objective, UID blend, evaluation, graph queries, dashboard
panels) with the math and economic reasoning behind each step and what each step actually
produces. Every number in it is reproduced by:
python scripts/diagnose_foundations.py --n-transactions 120000 --out outputs/diagnostics.jsondocs/generator_v2_verification.md describes the v2 generator, the causal graph features and the
verification results (ceiling PR-AUC 0.72; graph features lift a transaction-only model from 0.67 to 0.70).
Build v2 artifacts with:
python scripts/build_backend_api_artifacts.py --generator v2 --objective logistic \
--causal-graph-features --calibrate --n-transactions 120000 --output-dir outputs/backend_api
python scripts/verify_generator_v2.py --out docs/generator_v2_verification.jsonsrc/fraud_graphs/synthetic.py: v1 marginal-profile generator.src/fraud_graphs/synthetic_v2.py: v2 actor-level generator (households, rings, ATO, first-party fraud).src/fraud_graphs/graph_features.py: causal neighbour aggregates (leak-free graph features).src/fraud_graphs/features.py: UID aggregations + UID label blending.src/fraud_graphs/graph_embeddings.py: sparse graph embedding features.src/fraud_graphs/modeling.py: XGBoost training + asymmetric focal objective + explanation helper.src/fraud_graphs/backend_builder.py: builds full backend artifacts.src/fraud_graphs/graph_module.py: graph store and neighborhood/pattern queries.src/fraud_graphs/api_server.py: FastAPI app with all endpoints.
conda run -n gpt2-pytorch python scripts/build_backend_api_artifacts.py \
--real-transaction-path data/train_transaction.csv \
--real-identity-path data/train_identity.csv \
--n-transactions 250000 \
--output-dir outputs/backend_apiArtifacts are written under outputs/backend_api/artifacts.
conda run -n gpt2-pytorch python scripts/run_api.py \
--artifacts-dir outputs/backend_api/artifacts \
--host 0.0.0.0 \
--port 8000Alternative:
conda run -n gpt2-pytorch uvicorn fraud_graphs.asgi:app --host 0.0.0.0 --port 8000GET /healthGET /api/v1/metaGET /api/v1/graph/overviewGET /api/v1/graph/neighborhood?node_id=tx:2000001&hops=2GET /api/v1/graph/transaction/{transaction_id}GET /api/v1/graph/patterns/starsGET /api/v1/graph/patterns/ringsGET /api/v1/transactions/{transaction_id}GET /api/v1/transactions/{transaction_id}/explainGET /api/v1/embedding-space?sample_size=2500GET /api/v1/timeline/entity?entity_type=DeviceInfo&entity_value=iOSGET /api/v1/search?q=2000
Most endpoints accept:
dataset=main(default)dataset=demo(small, idealized motif dataset for graph UX validation)
conda run -n gpt2-pytorch python scripts/smoke_test_api.py \
--artifacts-dir outputs/backend_api/artifactsThe frontend/ directory contains a single-page Next.js investigation dashboard:
- force-directed graph explorer
- transaction explainability panel
- UID timeline view
- embedding scatter snapshot
- star/ring pattern lists
cd frontend
cp .env.example .env.local
npm install
npm run devOpen http://localhost:3000.
Generate precomputed JSON snapshots locally:
conda run -n gpt2-pytorch python scripts/export_frontend_snapshots.py \
--artifacts-dir outputs/backend_api/artifacts \
--out-dir frontend/public/snapshotsThen configure frontend env for static serving only:
NEXT_PUBLIC_DATA_SOURCE=staticNEXT_PUBLIC_FRAUD_API_BASE_URLis ignored in static mode
In this mode, the frontend reads only files under frontend/public/snapshots/ and does zero backend/model computation in cloud runtime.
- Commit
frontend/public/snapshots/*.json. - In Vercel, set project root to
frontend/. - Set env:
NEXT_PUBLIC_DATA_SOURCE=static
- Deploy.
Because the frontend is exported static (output: "export"), Vercel serves files only. No model training/feature engineering/inference runs in cloud runtime.
NEXT_PUBLIC_FRAUD_API_BASE_URLdefault:http://127.0.0.1:8000NEXT_PUBLIC_DATA_SOURCE:api= live API callsstatic= precomputed snapshot files only (recommended for Vercel free tier)
Backend CORS now allows local frontend origins by default:
http://localhost:3000http://127.0.0.1:3000
Override with:
FRAUD_API_CORS_ORIGINS=http://localhost:3000,http://your-host