Skip to content
View chloe4ai's full-sized avatar

Block or report chloe4ai

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please donโ€™t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this userโ€™s behavior. Learn more about reporting abuse.

Report abuse
chloe4ai/README.md

Hi, I'm Chloe ๐Ÿ‘‹

๐Ÿ“ SF Bay Area, CA | ๐Ÿงญ Principal PM | ๐Ÿš€ Building full-cycle AI products

Product Management LLM Evaluation RAG Recommendation Systems Python TypeScript Claude

I turn human judgment into measurable product quality. After a decade shipping 0โ†’1 and at-scale AI/ML products, I'm now building full-cycle AI tools end to end โ€” from the underlying architecture to the practical apps people actually use.

About Me

I'm a Principal Product Manager in Walmart's Data & Ads Global Product Org, where I run the human-in-the-loop evaluation program behind AI-driven products โ€” designing eval frameworks, scorecards, and annotator operations that make models measurably better.

My background blends Computer Science, Cognitive Science, and an MBA (Cornell University) โ€” technical fluency paired with consulting-grade strategy. I care equally about the low-level architecture and the real-world application.

What I work on

  • ๐Ÿงช LLM & AI evaluation โ€” eval harnesses, scorecards, quality benchmarks, and human-in-the-loop pipelines
  • ๐Ÿงญ Annotator operations at scale โ€” onboarding, calibration, and quality governance across global vendor teams
  • ๐ŸŽฏ Recommendation & personalization โ€” ranking, relevance, and retention-driving ML products
  • ๐Ÿ› ๏ธ Full-cycle product building โ€” 0โ†’1 prototypes through scaled launches, hands-on as both planner and builder
  • ๐Ÿ“Š Quantitative impact โ€” every initiative tied to revenue, retention, quality, or efficiency

๐Ÿ—‚๏ธ Selected Work

Twenty-nine repos is a pile, not a portfolio. Here is the same material sorted by what each piece is actually evidence of.

AI evaluation & data quality โ€” the thesis

Project What it demonstrates
LLM Eval Scorecard ยท live Side-by-side human eval with the parts most scorecards skip: randomized pane order with a position-bias test, bootstrap confidence intervals on the score gap, Krippendorff's ฮฑ and weighted ฮบ across raters, and a sample-size readout. Answers can you act on this yet?, not just who won?
AI Trackers Three scheduled trackers whose entire logic lives in natural-language SKILL.md files rather than code โ€” an agent reads them, gathers live data, and delivers a bilingual digest. A bet on prompts-as-programs.

Ad tech & marketplace โ€” the day job, built out

Project What it demonstrates
Ad Creative Optimizer ยท live Predicts creative fatigue across Google, Meta and TikTok, then auto-rotates. Decisioning and pacing logic made inspectable.
Offsite Ads Demo Off-site campaigns on Amazon DSP and Google โ€” retail media mechanics end to end.
FarmVend ยท live Vendor management for farmers markets: inventory prediction, payments, dynamic pricing. A whole two-sided marketplace at small scale.

Product prototypes โ€” built to argue a position, not to look pretty

Each of these was built for a specific role conversation. The README in each explains the problem it picks, the product positions it defends, and what it deliberately is not.

Project The argument
Glean Delivery Intelligence Proactive beats search โ€” but only if every pushed insight can show the weight and measured precision of the signals beneath it.
Glean Compass Priority drift, detected. Precision over recall on purpose: two flags, not forty, and dismissal is training data rather than a delete button.
Glean EI ยท full-stack The same thesis with a real Express backend โ€” confidence genuinely recomputes server-side, agents stream over SSE.
NerdWallet Money Next Steps Sequencing beats ranking: clear the 24% APR balance before the rewards card. Plus a full-stack build on Next.js + SQLite with a live funnel.
Lyft Verticals ยท live Airport, scheduled rides and teens as one system rather than three features.
Stitch Fix Household ยท live Extending personal styling to a household โ€” cross-category browse and shared profiles.
Midi Health ยท live Five specific conversion changes to a telehealth funnel, each tied to the anxiety it removes.

Tools I built because I wanted them

Project What it is
Sotto ๐ŸŽ™ Privacy-first on-device voice dictation. Speak softly, never type.
Notch ๆ—ฅ่ฟน ๐Ÿ“” Local-first macOS work journal that traces your day into AI summaries.
Meeting Scribe Menu-bar app that notices a Zoom/Meet/Teams call has started and handles the notes.
LinkedIn Schedule Send The button LinkedIn should have shipped.

Data-driven scouting

Project What it is
Adaptation Radar Ranks books by screen-adaptation potential from four live public APIs, with every weight a slider so the model is something you argue with. A scheduled harvest records history, so the board shows the slope, not just the level.
StreamScope Unifies viewing history across Netflix, HBO Max, Hulu, Disney+ and Prime.
Creative work โ€” animatics from Romance of the Three Kingdoms, and a song
Project
Red Cliffs ่ตคๅฃ The battle that split the empire three ways
Guandu ๅฎ˜ๆธก Cao Cao outnumbered ten to one
Three Visits ไธ‰้กง่Œ…ๅปฌ Liu Bei at the thatched cottage, three times
Six Sorties ๅ…ญๅ‡บ็ฅๅฑฑ Zhuge Liang's northern expeditions
ๅคฉไธๅ†ๅ€Ÿ Lyrics, arrangement, and three locally generated vocal takes

๐Ÿ“ˆ Track Record

  • Walmart โ€” Principal PM, Data & Ads ยท lifted model response quality 42% and helpfulness 38% across 10,000+ enterprise users; scaled eval ops to 600+ annotators across 4 vendor partners
  • iHerb โ€” Lead PM ยท improved recommendation precision 33%, cut churn 22%, raised inter-annotator agreement from 71% โ†’ 94%
  • Wish โ€” Senior PM ยท boosted engagement 38% and grew active users 147% on the personalized recommendation engine
  • Amazon โ€” Senior PM ยท lifted forecast accuracy 24% and cut stockouts 18% across millions of SKUs

๐Ÿ“Š GitHub Activity

GitHub Contribution Graph

๐ŸŒฑ What I'm Focused On

  • Building a full-cycle AI product โ€” owning architecture and application end to end
  • Open to remote Product Manager roles โ€” AI/ML, evaluation, personalization, and platform products
  • Sharing what I learn โ€” practical patterns for human-in-the-loop evaluation and AI product development

๐Ÿค Connect

LinkedIn Email GitHub


"Ship measurable quality." โ€” I build evaluation systems and AI products that turn human judgment into outcomes you can prove.

Pinned Loading

  1. streamscope streamscope Public

    Unifies viewing history across Netflix, HBO Max, Hulu, Disney+, and Prime Video into a single cross-platform recommendation dashboard.

    JavaScript

  2. farmers-market-app farmers-market-app Public

    Mochi Market โ€” a vendor app for farmers markets with inventory prediction, dynamic pricing, and sales reporting.

    TypeScript

  3. ai-trackers ai-trackers Public

    Three prompt-driven scheduled AI trackers (Thinker / Feature / Trend) that deliver bilingual EN+ไธญๆ–‡ digests as Google Calendar reminders.

  4. notch notch Public

    Notchๆ—ฅ่ฟน is a local-first macOS work journal that automatically traces your day โ€” apps, windows, and screenshots โ€” and turns it into AI summaries you can ask questions about.

    JavaScript

  5. sotto sotto Public

    ๐ŸŽ™ Speak softly, never type. Privacy-first on-device AI voice dictation for macOS.

    Python

  6. linkedin-schedule-send linkedin-schedule-send Public

    Chrome extension that adds a Schedule Send button to LinkedIn's message composer. Write on the weekend, send on a weekday.

    JavaScript