로그인앱 받기

채용 공고 › 공고

Data Scientist - Evaluations, Chanakya

Sarvam AI · Bengaluru

풀타임현장 근무영어

직무 소개

기업 공고 원문에서 · Sarvam AI · 2026년 4월 17일 게시

About Sarvam Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC. About the Role This is an intellectually demanding role. The Data Scientist anchors the evaluations function for this vertical — designing, building, and maintaining evaluation frameworks that measure model and system quality in operational context. You are not running standard benchmarks. You are building domain-specific eval harnesses for high-stakes use cases where a wrong answer carries real consequences. You will work closely with the MLOps Engineer, the PM, and the deployment team. The evaluations you design are the mechanism by which the team determines whether what we've built is good enough to deploy — and whether it stays good after deployment. What You'll Do • Design and build evaluation frameworks for Sarvam's AI outputs across domain-specific requirements: document comprehension, command summarisation, geospatial reasoning, enterprise workflow automation, and others as they emerge • Define quality metrics in collaboration with domain experts and clients; translate operational requirements into measurable, defensible signals • Run structured evaluation cycles pre- and post-deployment; build dashboards that surface model quality in production • Identify failure modes, edge cases, and distribution shifts — with the bias of someone looking for what's wrong, not confirming what's right • Collaborate with the MLOps Engineer to operationalise eval pipelines — automated, triggered by deployment events, versioned, and reproducible • Build and manage domain-specific datasets for fine-tuning, evaluation, and benchmarking — including human annotation workflows where needed • Publish internal findings and quality reports that feed the product and engineering roadmap What We're Looking For • 3–6 years in data science, ML research, or applied AI; at least 2 years working with LLMs in production contexts • Strong statistics and probability fundamentals — you understand what makes an evaluation valid and what makes it misleading • Experience designing evaluation frameworks from scratch: custom metrics, inter-rater reliability, red-teaming methodologies • Python proficiency; comfort with pandas, NumPy, HuggingFace datasets, RAGAS, EleutherAI Eval Harness, LangSmith, or equivalent • Experience with prompt engineering, model fine-tuning, or RLHF in applied settings • Ability to work with unstructured domain data: PDFs, doctrine documents, transcripts, and field reports Bonus Points • Prior work in high-stakes domains (healthcare, legal, defence, finance) where output quality

모든 공고에서 나의 매치 점수를 확인하세요

BabZituna는 모든 공고를 내 프로필과 비교해 여섯 가지 실제 기준으로 점수를 매기고, 왜 그 점수가 나왔는지 보여 줍니다. 공정성 감사도 거쳤습니다(공개 편향 감사 읽기).

앱 받기 → ✓ 구직자는 100% 무료
한 공고의 점수 산출 예시
기술96경력90근무지84근무 방식74고용 형태61급여데이터 없음

예시 수치이며 실제 지원자가 아닙니다. 각 기준은 내 프로필을 바탕으로 100점 만점으로 채점하며, 측정할 수 없는 기준은 추측하지 않고 그렇다고 표시합니다.