로그인앱 받기

채용 공고 › 공고

Infrastructure SRE - HPC

Sarvam AI · Bengaluru

풀타임현장 근무영어

직무 소개

기업 공고 원문에서 · Sarvam AI · 2026년 6월 26일 게시

About Sarvam Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC. About the Role Sarvam runs a large, multi-vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs and must run uninterrupted for weeks, and inference services that must hold a flat p99 under production load. Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve. This is not a Kubernetes administration role. We assume Kubernetes fluency as a baseline. The difficulty lies above and below it - in parallel filesystems under heavy checkpoint load, in RDMA fabrics that degrade quietly, in NCCL hangs whose root cause may be the network or the kernel, in driver and firmware drift across heterogeneous hardware, and in distributed training failures that masquerade as infrastructure faults. We are hiring a team of specialists rather than a set of identical generalists. This posting covers five areas of focus. We expect candidates to bring genuine depth in one and working fluency across the others, because on a shared fleet a storage problem often first appears as a training hang, and the engineer on call must route an incident correctly before anyone can resolve it. When you apply, please indicate the area of focus that best matches your experience. Strong generalists are welcome; we will place you where your depth is most useful. What You’ll Do Operate the GPU fleet end to end across training and serving - provisioning, observability, capacity, and fleet health. Hold a meaningful on-call rotation, write runbooks that hold up under pressure, and drive postmortems that produce durable fixes. Build the internal tooling the team relies on, rather than operating off-the-shelf systems alone. Partner with ML and platform teams to keep large runs alive and serving latency predictable. What We're Looking For 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.* Demonstrated on-call ownership of infrastructure that mattered, with a track record of postmortems that led to real change. Proficiency in Python or Go, used to build and maintain internal tooling. Working fluency across all five areas of focus below - enough to recognize, triage, and route a problem outside your specialty, even if the fix belongs to a teammate. * For the Storage and Fabric areas of focus, we will weigh deep domain expertise against the GPU-cluster requirement; exceptional specialists with less direct GPU-fleet time

언급된 기술

Infrastructure

모든 공고에서 나의 매치 점수를 확인하세요

BabZituna는 모든 공고를 내 프로필과 비교해 여섯 가지 실제 기준으로 점수를 매기고, 왜 그 점수가 나왔는지 보여 줍니다. 공정성 감사도 거쳤습니다(공개 편향 감사 읽기).

앱 받기 → ✓ 구직자는 100% 무료
한 공고의 점수 산출 예시
기술96경력90근무지84근무 방식74고용 형태61급여데이터 없음

예시 수치이며 실제 지원자가 아닙니다. 각 기준은 내 프로필을 바탕으로 100점 만점으로 채점하며, 측정할 수 없는 기준은 추측하지 않고 그렇다고 표시합니다.