MasukUnduh aplikasi

Lowongan › Iklan

Platform Engineer - AI Infrastructure

Sarvam AI · Bengaluru

Penuh waktuDi lokasiInggris

Tentang posisi ini

Dari iklan pemberi kerja · Sarvam AI · diterbitkan 29 Juni 2026

Platform Engineer - AI Infrastructure About Sarvam Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC. About the Role Sarvam runs a large, multi-vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs, and inference services that must hold a flat p99 under production load. This role builds the platform that sits on top of that fleet - the scheduling, scaling, multi-tenancy, serving, and self-service layers that let our ML teams use thousands of GPUs without a human in the loop for every job. This is the build side of infrastructure, not the operate side. Our SREs keep the fleet reliable and carry the pager; you build the systems that make the fleet usable - and that make it break less in the first place. The work is heavy software engineering: you will design and ship control-plane services, scheduler integrations, autoscaling controllers, the inference-serving platform, RBAC and quota systems, observability and cost tooling, and the CLI and APIs ML engineers touch every day. You will treat the platform as a product and its internal users as customers. You should be a strong systems engineer who is fluent in Kubernetes at the controller and internals level, comfortable in Go or Python, and literate in the GPU-specific constraints - MIG and GPU sharing, gang scheduling, topology-aware placement, RDMA - that make this different from a generic cloud platform. What You’ll Do This is the surface area of the platform. You won't own all of it at once - you'll take a capability and build it end to end - but you should be excited by most of it. The serving platform. The control plane that turns a model artifact into a scalable, multi-tenant endpoint: intelligent routing and load balancing across replicas, rollout machinery (canary, blue-green, rollback), traffic splitting, model-version integration etc. The scaling and elasticity layer. Autoscaling for both training (elastic and gang scaling, scale-to-fit) and serving (queue-depth and utilization-driven), capacity pooling across clusters, burst handling, preemption and reclaim, and efficient bin-packing across GPUs etc. Scheduling and orchestration. The scheduler layer itself - Kueue, Volcano, Slurm-on-Kubernetes, or custom controllers - with gang scheduling, priority and preemption, queue fairness, quota enforcement, and topology-aware placement that keeps a job's ranks on the same fabric island. Multi-tenancy, RBAC, and isolation. The tenant model and namespacing, RBAC and policy, quota and fair-share enforcement, M

Keahlian yang disebutkan

Infrastructure

Lihat skor kecocokan Anda untuk setiap posisi

BabZituna menilai setiap lowongan terhadap profil Anda dalam enam dimensi nyata dan menunjukkan MENGAPA skornya demikian, diaudit untuk keadilan (baca audit bias publik).

Unduh aplikasi → ✓ Gratis 100% untuk pencari kerja
Bagaimana satu lowongan dinilai Contoh
Keahlian96Pengalaman90Lokasi84Gaya kerja74Jenis pekerjaan61GajiTidak ada data

Angka contoh, bukan kandidat sungguhan. Setiap dimensi dinilai dari 100 berdasarkan profil Anda sendiri, dan dimensi yang tidak bisa kami ukur kami nyatakan begitu, bukan kami tebak.