AI systems • Reliable agent evaluation • Full-stack product engineering
Building practical software that is measurable, maintainable, and ready to ship.
| Project | What it demonstrates |
|---|---|
| ToolReliBench | Reliability evaluation for long-horizon, tool-using LLM agents |
| ToolReliBench1 | Extended benchmark implementation with metrics, synthetic traces, and reproducibility docs |
| College AI Assistant | Django-backed college support portal with authentication, student services, and AI-assistant UI |
| Portfolio overview | Curated index of projects, skills, and engineering focus |
The account also contains complete working builds of VidStream (Next.js + Firebase video platform) and the College AI Assistant backend (Django + REST + Celery). They are currently private while deployment and configuration review is completed.
- AI and evaluation: TF-IDF, cosine similarity, reliability metrics, failure taxonomies, experiment pipelines
- Backend: Python, Django, Django REST Framework, Flask, Celery, PostgreSQL, Redis
- Frontend and mobile: Next.js, React, Flutter, Tailwind CSS
- Engineering practice: Docker, testing, reproducible experiments, security-conscious configuration, clear documentation
I enjoy turning ambitious product ideas and research questions into working systems. My projects combine applied AI, full-stack development, and careful engineering fundamentals: explicit metrics, reproducible setup, and documentation that helps another developer run the project.
- GitHub: @sarvesh2003dev
- Email: sarveshtamse08@gmail.com