Se connecterObtenir l’app

Offres d’emploi › Annonce

Data Scientist - Evaluations, Chanakya

Sarvam AI · Bengaluru

Temps pleinSur siteAnglais

Le poste

Tiré de l’annonce de l’employeur · Sarvam AI · publiée le 17 avril 2026

About Sarvam Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC. About the Role This is an intellectually demanding role. The Data Scientist anchors the evaluations function for this vertical — designing, building, and maintaining evaluation frameworks that measure model and system quality in operational context. You are not running standard benchmarks. You are building domain-specific eval harnesses for high-stakes use cases where a wrong answer carries real consequences. You will work closely with the MLOps Engineer, the PM, and the deployment team. The evaluations you design are the mechanism by which the team determines whether what we've built is good enough to deploy — and whether it stays good after deployment. What You'll Do • Design and build evaluation frameworks for Sarvam's AI outputs across domain-specific requirements: document comprehension, command summarisation, geospatial reasoning, enterprise workflow automation, and others as they emerge • Define quality metrics in collaboration with domain experts and clients; translate operational requirements into measurable, defensible signals • Run structured evaluation cycles pre- and post-deployment; build dashboards that surface model quality in production • Identify failure modes, edge cases, and distribution shifts — with the bias of someone looking for what's wrong, not confirming what's right • Collaborate with the MLOps Engineer to operationalise eval pipelines — automated, triggered by deployment events, versioned, and reproducible • Build and manage domain-specific datasets for fine-tuning, evaluation, and benchmarking — including human annotation workflows where needed • Publish internal findings and quality reports that feed the product and engineering roadmap What We're Looking For • 3–6 years in data science, ML research, or applied AI; at least 2 years working with LLMs in production contexts • Strong statistics and probability fundamentals — you understand what makes an evaluation valid and what makes it misleading • Experience designing evaluation frameworks from scratch: custom metrics, inter-rater reliability, red-teaming methodologies • Python proficiency; comfort with pandas, NumPy, HuggingFace datasets, RAGAS, EleutherAI Eval Harness, LangSmith, or equivalent • Experience with prompt engineering, model fine-tuning, or RLHF in applied settings • Ability to work with unstructured domain data: PDFs, doctrine documents, transcripts, and field reports Bonus Points • Prior work in high-stakes domains (healthcare, legal, defence, finance) where output quality

Voyez votre score de compatibilité pour chaque poste

BabZituna note chaque offre selon votre profil sur six critères réels et vous montre POURQUOI elle a obtenu ce score, avec un audit d’équité (lire l’audit public des biais).

Obtenir l’app → ✓ 100 % gratuit pour les chercheurs d’emploi
Comment une offre a été notée Exemple
Compétences96Expérience90Localisation84Mode de travail74Type de contrat61SalairePas de données

Chiffres d’exemple, pas un vrai candidat. Chaque critère est noté sur 100 à partir de votre propre profil, et celui que nous ne pouvons pas mesurer le dit au lieu de deviner.