An AI-powered customer-support agent built for SpotifyCares using the Customer Support on Twitter dataset.
The system is designed to:
- Classify incoming customer messages into support intents.
- Retrieve relevant historical SpotifyCares support interactions.
- Generate a response grounded in historical support behavior.
- Decide whether a request should be auto-handled or escalated to a human, with an explicit reason.
- Evaluate the system using a manually reviewed Golden Evaluation Set, automated metrics, an LLM-as-judge, and human agreement checks.
Customer-support messages on Twitter are often:
- Short
- Informal
- Noisy
- Ambiguous
- Missing important context
The goal of this project is to build a practical support-agent pipeline that can understand an incoming Spotify customer message, identify its intent, find relevant historical support responses, and safely determine whether the request can be automatically handled.
The system prioritizes safe routing and evidence-based responses rather than attempting to automatically answer every customer message.
The overall pipeline is:
Customer Message
|
v
Intent Classification
|
+------------+------------+
| |
v v
Historical Retrieval Confidence / Risk Checks
| |
+------------+------------+
|
v
Auto-handle / Escalate
|
+---------+---------+
| |
v v
Draft Response Human Escalation
For auto-handled requests, the response generator uses retrieved historical SpotifyCares interactions as grounding evidence.
SpotifyCares was selected because the dataset contains a large number of customer-support interactions for the brand, allowing the system to learn recurring support patterns.
The processed SpotifyCares data contains approximately:
| Artifact | Count |
|---|---|
| SpotifyCares brand/support tweets | 43,265 |
| Inbound customer messages | 41,585 |
| Direct customer → SpotifyCares response links | 39,188 |
| Historical customer-response pairs | 43,092 |
| Cleaned retrieval examples | 41,050 |
| Weakly labelled classifier examples | 19,137 |
The original full dataset is intentionally not included in this repository because of its size.
Only the processed artifacts required to run and evaluate the project are included.
The project uses:
Customer Support on Twitter
Kaggle dataset:
thoughtvector/customer-support-on-twitter
The original dataset contains approximately 3 million tweets across multiple brands.
The SpotifyCares subset was extracted and processed from the original dataset.
Customer/support relationships were reconstructed using the response relationships provided in the dataset.
The preprocessing pipeline:
- Identifies SpotifyCares support tweets.
- Identifies inbound customer messages.
- Connects customer messages to their corresponding SpotifyCares responses.
- Removes duplicates.
- Removes Twitter usernames and internal routing markers.
- Normalizes text.
- Builds a historical customer → support-response corpus.
The final taxonomy contains 13 operational intents.
| Intent | Description |
|---|---|
account_login |
Login, password, account access and account-management problems |
premium_subscription |
Premium subscriptions, upgrades, cancellations and subscription status |
billing_payment |
Charges, payment methods, billing and payment issues |
family_plan |
Spotify Family plan questions |
student_plan |
Student-plan eligibility and offers |
technical_app_issue |
App crashes, bugs, compatibility and technical problems |
playback_features |
Playback, shuffle, queue, offline playback and related features |
content_availability |
Missing songs, albums, artists or regional availability |
playlist_library |
Playlists, saved music and library-related issues |
ads_free_experience |
Advertisements and ad-free listening |
integrations_devices |
TVs, watches, speakers, phones and third-party/device integrations |
feature_request_general |
Requests or suggestions for product features |
unclear |
Insufficient information or ambiguous requests |
The unclear category is particularly important for safe routing because many real support messages do not provide enough information to confidently determine an intent.
A 200-example Golden Evaluation Set was created and manually reviewed.
Each example contains:
- Tweet ID
- Customer message
- Intent
- Expected action
- Label notes
The Golden Set is kept separate from classifier training.
It is used to evaluate:
- Intent classification
- Auto-handle vs escalation decisions
- Retrieval behavior
- End-to-end agent behavior
This separation prevents the evaluation examples from simply becoming training examples.
The main classifier uses:
TF-IDF
+
Logistic Regression
+
Balanced Class Weights
Historical support messages were weakly labelled using a deterministic keyword/rule-based labelling process.
The resulting training data contains:
19,137 examples across 12 trainable intents.
The unclear intent is handled primarily through confidence-based routing rather than being treated as a normal supervised training class.
A classifier confidence threshold of:
0.45
is used.
Predictions below this confidence level are routed to:
unclear
This provides a safety mechanism for uncertain messages.
The response-generation system is grounded in historical SpotifyCares support interactions.
The retrieval corpus is cleaned by:
- Grouping customer messages with corresponding support responses
- Removing duplicate customer examples
- Removing Twitter usernames
- Removing internal routing markers
- Normalizing text
The baseline retriever uses:
TF-IDF
+
Unigrams + Bigrams
+
Cosine Similarity
+
Top-K Retrieval
The system retrieves the most similar historical customer-support interactions.
The main demonstration uses the top historical examples as grounding evidence for response generation.
A leave-out retrieval evaluation is used so that a Golden Set example is not evaluated against itself.
This is important because directly retrieving the same example from the corpus would create data leakage and artificially inflate the retrieval score.
| Metric | Mean Similarity |
|---|---|
| Top-1 | 0.4250 |
| Top-3 | 0.3804 |
| Top-5 | 0.3557 |
These are retrieval-similarity diagnostics, not classification accuracy.
For automatically handled requests, the system uses retrieved historical support interactions as grounding evidence.
The local model used for response generation is:
qwen2.5:7b-instruct
through Ollama.
Using a local model avoids requiring a paid external LLM API during evaluation.
The response generator is instructed to:
- Use historical evidence
- Stay relevant to the customer request
- Avoid unsupported claims
- Avoid inventing policies
- Provide useful support guidance
- Ask for clarification when necessary
The system intentionally does not automatically answer every request.
A request is escalated when one or more safety/routing conditions are triggered.
- Predicted intent is
unclear - Classifier confidence is below the operational threshold
- Historical retrieval evidence is weak
- The issue involves higher-risk account situations
- The issue involves billing/payment situations
Higher-risk categories such as account access and billing/payment are handled conservatively.
The system returns:
Intent
Confidence
Action
Reason
Retrieved Evidence
Draft Response
Example:
Intent: billing_payment
Confidence: 0.82
Action: ESCALATE
Reason:
Billing-related requests are treated conservatively because
incorrect automated guidance could affect a customer's payment
or subscription status.
Two baseline approaches were implemented.
A deterministic keyword-based classifier provides an interpretable baseline.
| Metric | Rule Baseline |
|---|---|
| Accuracy | 57.00% |
| Macro-F1 | 0.5806 |
A TF-IDF historical-response retriever is evaluated independently of the classifier.
| Metric | Mean Similarity |
|---|---|
| Top-1 | 0.4250 |
| Top-3 | 0.3804 |
| Top-5 | 0.3557 |
Retrieval similarity is reported separately rather than incorrectly comparing it directly with classification accuracy.
The final classifier was evaluated on the untouched 200-example Golden Set.
| Metric | Result |
|---|---|
| Intent Accuracy | 57.50% |
| Macro-F1 | 0.5859 |
| Action Accuracy | 65.00% |
| Evaluation Examples | 200 |
The classifier performs well on some intents but struggles with semantically overlapping categories such as:
- Content availability
- Feature requests
- Family-plan questions
- Short technical complaints
- Subscription vs billing questions
A separate local Qwen model was used as an LLM judge on a 30-example response-quality audit.
The judge evaluated:
- Relevance
- Grounding
- Helpfulness
- Overall quality
| Dimension | Mean Score / 5 |
|---|---|
| Relevance | 2.00 |
| Grounding | 2.03 |
| Helpfulness | 1.87 |
| Overall | 1.97 |
The relatively low response-quality score is an important finding.
It shows that although the classification and routing pipeline is functional, the retrieval and response-generation layers still require improvement before being trusted for broad production use.
The LLM judge should be treated as an evaluation signal rather than absolute ground truth because the audit is based on a 30-example sample and uses a local model.
A second-pass human consistency check was performed on 30 examples using the same labelling rubric.
| Measure | Agreement | Cohen's Kappa |
|---|---|---|
| Intent | 83.33% | 0.8120 |
| Action | 90.00% | 0.8000 |
This provides evidence that the evaluation rubric is reasonably consistent.
However, this is a second-pass consistency check rather than a fully independent blinded inter-annotator study because the same rubric and assisted labelling workflow were used.
This is the largest classification problem.
- 26 Golden Set examples
- 23 errors
- 18 direct predictions as
unclear - Error rate: 88.5%
Many messages about missing songs, albums, artists, or regional availability are short and overlap with general content questions.
- 20 examples
- 17 errors
- 12 direct predictions as
unclear - Error rate: 85.0%
Feature requests can be phrased in many different ways and can resemble technical complaints.
- 8 examples
- 7 errors
- Error rate: 87.5%
Family-plan messages frequently contain generic subscription terminology, making the boundary between Family and general Premium questions difficult.
- 16 examples
- 8 errors
- Error rate: 50.0%
Short technical complaints often lack:
- Device information
- Operating system
- Application version
- Error messages
- Reproduction steps
- 16 examples
- 7 errors
- Error rate: 43.8%
Subscription, upgrade, offer, and payment language frequently overlaps with billing/payment issues.
The headline intent accuracy is:
57.50%
However, this number should not be interpreted as:
"57.5% of customer-support problems can be solved correctly in production."
There are several reasons.
The evaluation contains only 200 examples.
Some intents are represented more heavily than others.
The model performs much better on some categories than others and struggles heavily with categories such as:
- Content availability
- Feature requests
- Family plan
Macro-F1 provides a more balanced view:
Macro-F1 = 0.5859
An earlier end-to-end evaluator produced a retrieval score of 1.0 because Golden Set examples were present in the retrieval corpus.
That was identified as data leakage and is not reported as a valid retrieval result.
The final retrieval diagnostics use leave-out evaluation instead.
The major non-obvious decisions are documented separately in:
report/decision_log.md
Important decisions include:
- Selecting SpotifyCares as the target brand.
- Reconstructing customer/support conversations using response IDs.
- Cleaning duplicate historical interactions.
- Creating a compact operational intent taxonomy.
- Including
unclearas a safety-oriented routing class. - Keeping the Golden Set separate from training.
- Using TF-IDF + Logistic Regression for an interpretable classifier.
- Using historical responses as grounding evidence.
- Adding confidence and retrieval thresholds.
- Conservatively escalating account and billing issues.
- Using a local Qwen model for response generation.
- Using leave-out retrieval evaluation to prevent leakage.
Review the largest confusion groups and create clearer operational boundaries for:
- Content availability
- Feature requests
- Family-plan questions
Create a larger, high-quality manually labelled training set instead of relying mainly on keyword-based weak supervision.
Replace TF-IDF-only retrieval with:
Semantic Embeddings
+
Candidate Retrieval
+
Reranking
This should improve retrieval for short messages that use different wording from historical examples.
The response generator should:
- Explicitly ground answers in retrieved evidence
- Avoid unsupported claims
- Ask targeted clarification questions when evidence is insufficient
- Produce safer escalation messages
- Avoid copying irrelevant historical responses
Future evaluation should include:
- A larger Golden Set
- A genuinely independent second annotator
- Larger response-quality evaluation
- Human review alongside the LLM judge
- More robust retrieval metrics
- Per-intent precision, recall and F1
- Confidence calibration
Hiver-SDE-Assignment/
│
├── app.py
├── evaluate.py
├── requirements.txt
├── README.md
├── .gitignore
│
├── data/
│ ├── spotifycares_conversations.csv
│ ├── golden_set.csv
│ └── ...
│
├── src/
│ ├── classifier.py
│ ├── retrieval.py
│ ├── response_generator.py
│ ├── routing.py
│ └── preprocessing.py
│
├── models/
│ └── ...
│
├── evaluation/
│ ├── ...
│
├── report/
│ ├── decision_log.md
│ └── ...
│
└── outputs/
└── ...
The exact file list may vary slightly depending on the repository version.
Recommended environment:
Python 3.11+
Install Python dependencies:
pip install -r requirements.txtThe project uses Ollama for local response generation.
Install Ollama separately and make sure the required model is available:
ollama pull qwen2.5:7b-instructThe repository is designed so that the headline evaluation can be reproduced using the provided processed artifacts rather than downloading and processing the full ~3M-row dataset.
git clone <YOUR_GITHUB_REPOSITORY_URL>
cd Hiver-SDE-AssignmentWindows:
python -m venv venv
venv\Scripts\activateLinux/macOS:
python3 -m venv venv
source venv/bin/activatepip install -r requirements.txtpython evaluate.pyThe evaluation should produce the main metrics on the 200-example Golden Set.
Expected headline results:
Intent Accuracy : 57.50%
Macro-F1 : 0.5859
Action Accuracy : 65.00%
streamlit run app.pyOpen the local Streamlit URL shown in the terminal.
For an incoming message such as:
@SpotifyCares I can't log into my account anymore. Please help!
the system produces a structured decision similar to:
Intent:
account_login
Confidence:
<model confidence>
Action:
ESCALATE
Reason:
Account-access requests are handled conservatively because
incorrect automated instructions may affect account security.
Retrieved Historical Examples:
<top relevant SpotifyCares interactions>
Draft Response:
<grounded support response>
The exact output depends on the classifier prediction, retrieved historical evidence, and routing thresholds.
This project is a research/assignment prototype rather than a production customer-support system.
Important limitations include:
- The classifier is trained partly using weak labels.
- The Golden Set contains only 200 examples.
- Several intent categories have substantial semantic overlap.
- Retrieval is based on lexical TF-IDF similarity.
- Response quality remains relatively low according to the LLM-as-judge audit.
- The LLM judge is itself imperfect.
- Human agreement was checked on only a 30-example sample.
- The original dataset is noisy and contains incomplete conversations.
- The system should not automatically perform sensitive account or payment actions.
The safest deployment strategy would therefore be confidence-aware automation with conservative escalation.
The project demonstrates an end-to-end support-agent architecture consisting of:
Data Processing
↓
Intent Taxonomy
↓
Intent Classification
↓
Historical Response Retrieval
↓
Confidence / Risk Routing
↓
Auto-handle or Escalate
↓
Grounded Response Generation
↓
Evaluation
The main results are:
| Component | Result |
|---|---|
| Golden Evaluation Set | 200 examples |
| Intent Accuracy | 57.50% |
| Intent Macro-F1 | 0.5859 |
| Action Accuracy | 65.00% |
| Rule Baseline Accuracy | 57.00% |
| Rule Baseline Macro-F1 | 0.5806 |
| Retrieval Top-1 Similarity | 0.4250 |
| Retrieval Top-3 Similarity | 0.3804 |
| Retrieval Top-5 Similarity | 0.3557 |
| LLM Judge Overall | 1.97 / 5 |
| Human Intent Agreement | 83.33% |
| Human Action Agreement | 90.00% |
The key conclusion is that classification and routing are functional but not yet production-ready, while the response-generation and retrieval components are the most important areas for further improvement.
| Hiver Requirement | Status |
|---|---|
| Runnable repository | ✅ |
| README with reproduction instructions | ✅ |
| Brand-specific AI support agent | ✅ |
| Intent classification | ✅ |
| Historical response grounding | ✅ |
| Auto-handle / escalation decision | ✅ |
| 150–250 example Golden Set | ✅ 200 examples |
| Automated evaluation | ✅ |
| LLM-as-judge | ✅ |
| Human agreement check | ✅ |
| Two baselines | ✅ |
| Failure analysis | ✅ Top 5 |
| Misleading headline number section | ✅ |
| One-week improvement plan | ✅ |
| Decision log | ✅ |
This project focuses not only on building an AI support agent, but also on measuring whether the system is trustworthy.
The evaluation shows that a simple and interpretable architecture can provide a functional baseline for customer-support automation, while also exposing important weaknesses in intent ambiguity, retrieval quality, and response generation.
The most important next step is not simply increasing the headline accuracy, but improving per-intent reliability, retrieval grounding, response quality, and safe escalation behavior.