๐ SF Bay Area, CA | ๐งญ Principal PM | ๐ Building full-cycle AI products
I turn human judgment into measurable product quality. After a decade shipping 0โ1 and at-scale AI/ML products, I'm now building full-cycle AI tools end to end โ from the underlying architecture to the practical apps people actually use.
I'm a Principal Product Manager in Walmart's Data & Ads Global Product Org, where I run the human-in-the-loop evaluation program behind AI-driven products โ designing eval frameworks, scorecards, and annotator operations that make models measurably better.
My background blends Computer Science, Cognitive Science, and an MBA (Cornell University) โ technical fluency paired with consulting-grade strategy. I care equally about the low-level architecture and the real-world application.
What I work on
- ๐งช LLM & AI evaluation โ eval harnesses, scorecards, quality benchmarks, and human-in-the-loop pipelines
- ๐งญ Annotator operations at scale โ onboarding, calibration, and quality governance across global vendor teams
- ๐ฏ Recommendation & personalization โ ranking, relevance, and retention-driving ML products
- ๐ ๏ธ Full-cycle product building โ 0โ1 prototypes through scaled launches, hands-on as both planner and builder
- ๐ Quantitative impact โ every initiative tied to revenue, retention, quality, or efficiency
Twenty-nine repos is a pile, not a portfolio. Here is the same material sorted by what each piece is actually evidence of.
| Project | What it demonstrates |
|---|---|
| LLM Eval Scorecard ยท live | Side-by-side human eval with the parts most scorecards skip: randomized pane order with a position-bias test, bootstrap confidence intervals on the score gap, Krippendorff's ฮฑ and weighted ฮบ across raters, and a sample-size readout. Answers can you act on this yet?, not just who won? |
| AI Trackers | Three scheduled trackers whose entire logic lives in natural-language SKILL.md files rather than code โ an agent reads them, gathers live data, and delivers a bilingual digest. A bet on prompts-as-programs. |
| Project | What it demonstrates |
|---|---|
| Ad Creative Optimizer ยท live | Predicts creative fatigue across Google, Meta and TikTok, then auto-rotates. Decisioning and pacing logic made inspectable. |
| Offsite Ads Demo | Off-site campaigns on Amazon DSP and Google โ retail media mechanics end to end. |
| FarmVend ยท live | Vendor management for farmers markets: inventory prediction, payments, dynamic pricing. A whole two-sided marketplace at small scale. |
Each of these was built for a specific role conversation. The README in each explains the problem it picks, the product positions it defends, and what it deliberately is not.
| Project | The argument |
|---|---|
| Glean Delivery Intelligence | Proactive beats search โ but only if every pushed insight can show the weight and measured precision of the signals beneath it. |
| Glean Compass | Priority drift, detected. Precision over recall on purpose: two flags, not forty, and dismissal is training data rather than a delete button. |
| Glean EI ยท full-stack | The same thesis with a real Express backend โ confidence genuinely recomputes server-side, agents stream over SSE. |
| NerdWallet Money Next Steps | Sequencing beats ranking: clear the 24% APR balance before the rewards card. Plus a full-stack build on Next.js + SQLite with a live funnel. |
| Lyft Verticals ยท live | Airport, scheduled rides and teens as one system rather than three features. |
| Stitch Fix Household ยท live | Extending personal styling to a household โ cross-category browse and shared profiles. |
| Midi Health ยท live | Five specific conversion changes to a telehealth funnel, each tied to the anxiety it removes. |
| Project | What it is |
|---|---|
| Sotto ๐ | Privacy-first on-device voice dictation. Speak softly, never type. |
| Notch ๆฅ่ฟน ๐ | Local-first macOS work journal that traces your day into AI summaries. |
| Meeting Scribe | Menu-bar app that notices a Zoom/Meet/Teams call has started and handles the notes. |
| LinkedIn Schedule Send | The button LinkedIn should have shipped. |
| Project | What it is |
|---|---|
| Adaptation Radar | Ranks books by screen-adaptation potential from four live public APIs, with every weight a slider so the model is something you argue with. A scheduled harvest records history, so the board shows the slope, not just the level. |
| StreamScope | Unifies viewing history across Netflix, HBO Max, Hulu, Disney+ and Prime. |
Creative work โ animatics from Romance of the Three Kingdoms, and a song
| Project | |
|---|---|
| Red Cliffs ่ตคๅฃ | The battle that split the empire three ways |
| Guandu ๅฎๆธก | Cao Cao outnumbered ten to one |
| Three Visits ไธ้กง่ ๅปฌ | Liu Bei at the thatched cottage, three times |
| Six Sorties ๅ ญๅบ็ฅๅฑฑ | Zhuge Liang's northern expeditions |
| ๅคฉไธๅๅ | Lyrics, arrangement, and three locally generated vocal takes |
- Walmart โ Principal PM, Data & Ads ยท lifted model response quality 42% and helpfulness 38% across 10,000+ enterprise users; scaled eval ops to 600+ annotators across 4 vendor partners
- iHerb โ Lead PM ยท improved recommendation precision 33%, cut churn 22%, raised inter-annotator agreement from 71% โ 94%
- Wish โ Senior PM ยท boosted engagement 38% and grew active users 147% on the personalized recommendation engine
- Amazon โ Senior PM ยท lifted forecast accuracy 24% and cut stockouts 18% across millions of SKUs
- Building a full-cycle AI product โ owning architecture and application end to end
- Open to remote Product Manager roles โ AI/ML, evaluation, personalization, and platform products
- Sharing what I learn โ practical patterns for human-in-the-loop evaluation and AI product development
"Ship measurable quality." โ I build evaluation systems and AI products that turn human judgment into outcomes you can prove.
