Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Hiver SDE Intern Assignment — SpotifyCares AI Support Agent

An AI-powered customer-support agent built for SpotifyCares using the Customer Support on Twitter dataset.

The system is designed to:

  1. Classify incoming customer messages into support intents.
  2. Retrieve relevant historical SpotifyCares support interactions.
  3. Generate a response grounded in historical support behavior.
  4. Decide whether a request should be auto-handled or escalated to a human, with an explicit reason.
  5. Evaluate the system using a manually reviewed Golden Evaluation Set, automated metrics, an LLM-as-judge, and human agreement checks.

1. Problem Statement

Customer-support messages on Twitter are often:

  • Short
  • Informal
  • Noisy
  • Ambiguous
  • Missing important context

The goal of this project is to build a practical support-agent pipeline that can understand an incoming Spotify customer message, identify its intent, find relevant historical support responses, and safely determine whether the request can be automatically handled.

The system prioritizes safe routing and evidence-based responses rather than attempting to automatically answer every customer message.


2. System Overview

The overall pipeline is:

                    Customer Message
                           |
                           v
                  Intent Classification
                           |
              +------------+------------+
              |                         |
              v                         v
     Historical Retrieval       Confidence / Risk Checks
              |                         |
              +------------+------------+
                           |
                           v
                 Auto-handle / Escalate
                           |
                 +---------+---------+
                 |                   |
                 v                   v
          Draft Response       Human Escalation

For auto-handled requests, the response generator uses retrieved historical SpotifyCares interactions as grounding evidence.


3. Selected Brand

SpotifyCares

SpotifyCares was selected because the dataset contains a large number of customer-support interactions for the brand, allowing the system to learn recurring support patterns.

The processed SpotifyCares data contains approximately:

Artifact Count
SpotifyCares brand/support tweets 43,265
Inbound customer messages 41,585
Direct customer → SpotifyCares response links 39,188
Historical customer-response pairs 43,092
Cleaned retrieval examples 41,050
Weakly labelled classifier examples 19,137

The original full dataset is intentionally not included in this repository because of its size.

Only the processed artifacts required to run and evaluate the project are included.


4. Dataset

The project uses:

Customer Support on Twitter

Kaggle dataset:

thoughtvector/customer-support-on-twitter

The original dataset contains approximately 3 million tweets across multiple brands.

The SpotifyCares subset was extracted and processed from the original dataset.

Conversation Reconstruction

Customer/support relationships were reconstructed using the response relationships provided in the dataset.

The preprocessing pipeline:

  1. Identifies SpotifyCares support tweets.
  2. Identifies inbound customer messages.
  3. Connects customer messages to their corresponding SpotifyCares responses.
  4. Removes duplicates.
  5. Removes Twitter usernames and internal routing markers.
  6. Normalizes text.
  7. Builds a historical customer → support-response corpus.

5. Intent Taxonomy

The final taxonomy contains 13 operational intents.

Intent Description
account_login Login, password, account access and account-management problems
premium_subscription Premium subscriptions, upgrades, cancellations and subscription status
billing_payment Charges, payment methods, billing and payment issues
family_plan Spotify Family plan questions
student_plan Student-plan eligibility and offers
technical_app_issue App crashes, bugs, compatibility and technical problems
playback_features Playback, shuffle, queue, offline playback and related features
content_availability Missing songs, albums, artists or regional availability
playlist_library Playlists, saved music and library-related issues
ads_free_experience Advertisements and ad-free listening
integrations_devices TVs, watches, speakers, phones and third-party/device integrations
feature_request_general Requests or suggestions for product features
unclear Insufficient information or ambiguous requests

The unclear category is particularly important for safe routing because many real support messages do not provide enough information to confidently determine an intent.


6. Golden Evaluation Set

A 200-example Golden Evaluation Set was created and manually reviewed.

Each example contains:

  • Tweet ID
  • Customer message
  • Intent
  • Expected action
  • Label notes

The Golden Set is kept separate from classifier training.

It is used to evaluate:

  • Intent classification
  • Auto-handle vs escalation decisions
  • Retrieval behavior
  • End-to-end agent behavior

This separation prevents the evaluation examples from simply becoming training examples.


7. Intent Classification

The main classifier uses:

TF-IDF
   +
Logistic Regression
   +
Balanced Class Weights

Historical support messages were weakly labelled using a deterministic keyword/rule-based labelling process.

The resulting training data contains:

19,137 examples across 12 trainable intents.

The unclear intent is handled primarily through confidence-based routing rather than being treated as a normal supervised training class.

Confidence Threshold

A classifier confidence threshold of:

0.45

is used.

Predictions below this confidence level are routed to:

unclear

This provides a safety mechanism for uncertain messages.


8. Historical Response Retrieval

The response-generation system is grounded in historical SpotifyCares support interactions.

The retrieval corpus is cleaned by:

  • Grouping customer messages with corresponding support responses
  • Removing duplicate customer examples
  • Removing Twitter usernames
  • Removing internal routing markers
  • Normalizing text

Retrieval Method

The baseline retriever uses:

TF-IDF
+
Unigrams + Bigrams
+
Cosine Similarity
+
Top-K Retrieval

The system retrieves the most similar historical customer-support interactions.

The main demonstration uses the top historical examples as grounding evidence for response generation.

Leave-Out Evaluation

A leave-out retrieval evaluation is used so that a Golden Set example is not evaluated against itself.

This is important because directly retrieving the same example from the corpus would create data leakage and artificially inflate the retrieval score.

Retrieval Diagnostics

Metric Mean Similarity
Top-1 0.4250
Top-3 0.3804
Top-5 0.3557

These are retrieval-similarity diagnostics, not classification accuracy.


9. Response Generation

For automatically handled requests, the system uses retrieved historical support interactions as grounding evidence.

The local model used for response generation is:

qwen2.5:7b-instruct

through Ollama.

Using a local model avoids requiring a paid external LLM API during evaluation.

The response generator is instructed to:

  • Use historical evidence
  • Stay relevant to the customer request
  • Avoid unsupported claims
  • Avoid inventing policies
  • Provide useful support guidance
  • Ask for clarification when necessary

10. Auto-Handle vs Escalation

The system intentionally does not automatically answer every request.

A request is escalated when one or more safety/routing conditions are triggered.

Escalation Conditions

  • Predicted intent is unclear
  • Classifier confidence is below the operational threshold
  • Historical retrieval evidence is weak
  • The issue involves higher-risk account situations
  • The issue involves billing/payment situations

Higher-risk categories such as account access and billing/payment are handled conservatively.

Decision Output

The system returns:

Intent
Confidence
Action
Reason
Retrieved Evidence
Draft Response

Example:

Intent: billing_payment
Confidence: 0.82

Action: ESCALATE

Reason:
Billing-related requests are treated conservatively because
incorrect automated guidance could affect a customer's payment
or subscription status.

11. Baselines

Two baseline approaches were implemented.

Baseline 1 — Rule-Based Intent Classifier

A deterministic keyword-based classifier provides an interpretable baseline.

Results

Metric Rule Baseline
Accuracy 57.00%
Macro-F1 0.5806

Baseline 2 — Retrieval-Only Baseline

A TF-IDF historical-response retriever is evaluated independently of the classifier.

Leave-Out Retrieval Diagnostics

Metric Mean Similarity
Top-1 0.4250
Top-3 0.3804
Top-5 0.3557

Retrieval similarity is reported separately rather than incorrectly comparing it directly with classification accuracy.


12. Main Evaluation Results

The final classifier was evaluated on the untouched 200-example Golden Set.

Metric Result
Intent Accuracy 57.50%
Macro-F1 0.5859
Action Accuracy 65.00%
Evaluation Examples 200

The classifier performs well on some intents but struggles with semantically overlapping categories such as:

  • Content availability
  • Feature requests
  • Family-plan questions
  • Short technical complaints
  • Subscription vs billing questions

13. LLM-as-Judge Evaluation

A separate local Qwen model was used as an LLM judge on a 30-example response-quality audit.

The judge evaluated:

  • Relevance
  • Grounding
  • Helpfulness
  • Overall quality

Results

Dimension Mean Score / 5
Relevance 2.00
Grounding 2.03
Helpfulness 1.87
Overall 1.97

The relatively low response-quality score is an important finding.

It shows that although the classification and routing pipeline is functional, the retrieval and response-generation layers still require improvement before being trusted for broad production use.

The LLM judge should be treated as an evaluation signal rather than absolute ground truth because the audit is based on a 30-example sample and uses a local model.


14. Human Agreement Check

A second-pass human consistency check was performed on 30 examples using the same labelling rubric.

Measure Agreement Cohen's Kappa
Intent 83.33% 0.8120
Action 90.00% 0.8000

This provides evidence that the evaluation rubric is reasonably consistent.

However, this is a second-pass consistency check rather than a fully independent blinded inter-annotator study because the same rubric and assisted labelling workflow were used.


15. Top Failure Modes

1. Content Availability → unclear

This is the largest classification problem.

  • 26 Golden Set examples
  • 23 errors
  • 18 direct predictions as unclear
  • Error rate: 88.5%

Many messages about missing songs, albums, artists, or regional availability are short and overlap with general content questions.


2. Feature Requests → unclear

  • 20 examples
  • 17 errors
  • 12 direct predictions as unclear
  • Error rate: 85.0%

Feature requests can be phrased in many different ways and can resemble technical complaints.


3. Family Plan → premium_subscription / unclear

  • 8 examples
  • 7 errors
  • Error rate: 87.5%

Family-plan messages frequently contain generic subscription terminology, making the boundary between Family and general Premium questions difficult.


4. Technical Issues → unclear or Wrong Intent

  • 16 examples
  • 8 errors
  • Error rate: 50.0%

Short technical complaints often lack:

  • Device information
  • Operating system
  • Application version
  • Error messages
  • Reproduction steps

5. Premium Subscription → Billing / unclear

  • 16 examples
  • 7 errors
  • Error rate: 43.8%

Subscription, upgrade, offer, and payment language frequently overlaps with billing/payment issues.


16. What Is Misleading About My Headline Number?

The headline intent accuracy is:

57.50%

However, this number should not be interpreted as:

"57.5% of customer-support problems can be solved correctly in production."

There are several reasons.

1. The Golden Set is relatively small

The evaluation contains only 200 examples.

2. The class distribution is not perfectly balanced

Some intents are represented more heavily than others.

3. Performance varies significantly by intent

The model performs much better on some categories than others and struggles heavily with categories such as:

  • Content availability
  • Feature requests
  • Family plan

4. Accuracy hides class-level weaknesses

Macro-F1 provides a more balanced view:

Macro-F1 = 0.5859

5. Retrieval can leak if evaluated incorrectly

An earlier end-to-end evaluator produced a retrieval score of 1.0 because Golden Set examples were present in the retrieval corpus.

That was identified as data leakage and is not reported as a valid retrieval result.

The final retrieval diagnostics use leave-out evaluation instead.


17. Key Design Decisions

The major non-obvious decisions are documented separately in:

report/decision_log.md

Important decisions include:

  1. Selecting SpotifyCares as the target brand.
  2. Reconstructing customer/support conversations using response IDs.
  3. Cleaning duplicate historical interactions.
  4. Creating a compact operational intent taxonomy.
  5. Including unclear as a safety-oriented routing class.
  6. Keeping the Golden Set separate from training.
  7. Using TF-IDF + Logistic Regression for an interpretable classifier.
  8. Using historical responses as grounding evidence.
  9. Adding confidence and retrieval thresholds.
  10. Conservatively escalating account and billing issues.
  11. Using a local Qwen model for response generation.
  12. Using leave-out retrieval evaluation to prevent leakage.

18. What I Would Improve With One More Week

1. Improve the Intent Taxonomy

Review the largest confusion groups and create clearer operational boundaries for:

  • Content availability
  • Feature requests
  • Family-plan questions

2. Replace Weak Labels With Better Supervision

Create a larger, high-quality manually labelled training set instead of relying mainly on keyword-based weak supervision.


3. Improve Retrieval

Replace TF-IDF-only retrieval with:

Semantic Embeddings
        +
Candidate Retrieval
        +
Reranking

This should improve retrieval for short messages that use different wording from historical examples.


4. Improve Response Generation

The response generator should:

  • Explicitly ground answers in retrieved evidence
  • Avoid unsupported claims
  • Ask targeted clarification questions when evidence is insufficient
  • Produce safer escalation messages
  • Avoid copying irrelevant historical responses

5. Strengthen Evaluation

Future evaluation should include:

  • A larger Golden Set
  • A genuinely independent second annotator
  • Larger response-quality evaluation
  • Human review alongside the LLM judge
  • More robust retrieval metrics
  • Per-intent precision, recall and F1
  • Confidence calibration

19. Project Structure

Hiver-SDE-Assignment/
│
├── app.py
├── evaluate.py
├── requirements.txt
├── README.md
├── .gitignore
│
├── data/
│   ├── spotifycares_conversations.csv
│   ├── golden_set.csv
│   └── ...
│
├── src/
│   ├── classifier.py
│   ├── retrieval.py
│   ├── response_generator.py
│   ├── routing.py
│   └── preprocessing.py
│
├── models/
│   └── ...
│
├── evaluation/
│   ├── ...
│
├── report/
│   ├── decision_log.md
│   └── ...
│
└── outputs/
    └── ...

The exact file list may vary slightly depending on the repository version.


20. Requirements

Recommended environment:

Python 3.11+

Install Python dependencies:

pip install -r requirements.txt

The project uses Ollama for local response generation.

Install Ollama separately and make sure the required model is available:

ollama pull qwen2.5:7b-instruct

21. Reproducing the Results

The repository is designed so that the headline evaluation can be reproduced using the provided processed artifacts rather than downloading and processing the full ~3M-row dataset.

Step 1 — Clone the repository

git clone <YOUR_GITHUB_REPOSITORY_URL>
cd Hiver-SDE-Assignment

Step 2 — Create a virtual environment

Windows:

python -m venv venv
venv\Scripts\activate

Linux/macOS:

python3 -m venv venv
source venv/bin/activate

Step 3 — Install dependencies

pip install -r requirements.txt

Step 4 — Run evaluation

python evaluate.py

The evaluation should produce the main metrics on the 200-example Golden Set.

Expected headline results:

Intent Accuracy : 57.50%
Macro-F1        : 0.5859
Action Accuracy : 65.00%

Step 5 — Run the demo

streamlit run app.py

Open the local Streamlit URL shown in the terminal.


22. Example System Output

For an incoming message such as:

@SpotifyCares I can't log into my account anymore. Please help!

the system produces a structured decision similar to:

Intent:
account_login

Confidence:
<model confidence>

Action:
ESCALATE

Reason:
Account-access requests are handled conservatively because
incorrect automated instructions may affect account security.

Retrieved Historical Examples:
<top relevant SpotifyCares interactions>

Draft Response:
<grounded support response>

The exact output depends on the classifier prediction, retrieved historical evidence, and routing thresholds.


23. Safety and Limitations

This project is a research/assignment prototype rather than a production customer-support system.

Important limitations include:

  • The classifier is trained partly using weak labels.
  • The Golden Set contains only 200 examples.
  • Several intent categories have substantial semantic overlap.
  • Retrieval is based on lexical TF-IDF similarity.
  • Response quality remains relatively low according to the LLM-as-judge audit.
  • The LLM judge is itself imperfect.
  • Human agreement was checked on only a 30-example sample.
  • The original dataset is noisy and contains incomplete conversations.
  • The system should not automatically perform sensitive account or payment actions.

The safest deployment strategy would therefore be confidence-aware automation with conservative escalation.


24. Main Takeaways

The project demonstrates an end-to-end support-agent architecture consisting of:

Data Processing
      ↓
Intent Taxonomy
      ↓
Intent Classification
      ↓
Historical Response Retrieval
      ↓
Confidence / Risk Routing
      ↓
Auto-handle or Escalate
      ↓
Grounded Response Generation
      ↓
Evaluation

The main results are:

Component Result
Golden Evaluation Set 200 examples
Intent Accuracy 57.50%
Intent Macro-F1 0.5859
Action Accuracy 65.00%
Rule Baseline Accuracy 57.00%
Rule Baseline Macro-F1 0.5806
Retrieval Top-1 Similarity 0.4250
Retrieval Top-3 Similarity 0.3804
Retrieval Top-5 Similarity 0.3557
LLM Judge Overall 1.97 / 5
Human Intent Agreement 83.33%
Human Action Agreement 90.00%

The key conclusion is that classification and routing are functional but not yet production-ready, while the response-generation and retrieval components are the most important areas for further improvement.


25. Assignment Deliverables Checklist

Hiver Requirement Status
Runnable repository ✅
README with reproduction instructions ✅
Brand-specific AI support agent ✅
Intent classification ✅
Historical response grounding ✅
Auto-handle / escalation decision ✅
150–250 example Golden Set ✅ 200 examples
Automated evaluation ✅
LLM-as-judge ✅
Human agreement check ✅
Two baselines ✅
Failure analysis ✅ Top 5
Misleading headline number section ✅
One-week improvement plan ✅
Decision log ✅

26. Conclusion

This project focuses not only on building an AI support agent, but also on measuring whether the system is trustworthy.

The evaluation shows that a simple and interpretable architecture can provide a functional baseline for customer-support automation, while also exposing important weaknesses in intent ambiguity, retrieval quality, and response generation.

The most important next step is not simply increasing the headline accuracy, but improving per-intent reliability, retrieval grounding, response quality, and safe escalation behavior.

About

AI-powered Spotify customer support agent that understands user issues, searches relevant information, and provides helpful support responses.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages