01
GenAI Regression Evaluation
Problem
A production GenAI / RAG stack was shipping changes with no reliable way to know whether retrieval quality had regressed for real customers.
Approach
Built a golden dataset grounded in real customer utterances and automated daily regression runs. Led embedding-model evaluation (BGE-M3, Stella 1.5B) for the retrieval path, and shipped per-utterance explainability — MRR, mAP, Hit-Rate on every query — so non-ML stakeholders can debug a bad score without filing a ticket.
Impact
4M+
Datapoints evaluated / yr
Daily
Automated regression runs