登录获取应用

职位 › Seattle › 职位详情 今日新增

Staff Technical Program Manager – Infrastructure Maintenance & Change Management

CoreWeave · Bellevue, WA

实习现场办公英语

职位介绍

摘自雇主发布的职位信息 · CoreWeave · 发布于2026年10月6日

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

About the Role

CoreWeave is building the world's largest AI Cloud platform, and our fleet is growing at extraordinary speed. We're seeking a Senior Technical Program Manager to own fleet-wide reliability and operations programs that keep pace with that growth. This role partners closely with Compute, Networking, Data Center, and Operations teams to drive improvements in fleet operating efficiency — maximizing the percentage of fleet capacity that is healthy, available, and sellable — alongside reliability and stability goals.

You'll act as the central owner for fleet reliability and operations programs: aligning teams on priorities, defining success metrics, driving execution, and ensuring improvements hold at scale — not just land once and regress as the fleet scales.

The Technical Program Manager will

  • Establish and own fleet reliability metrics and dashboards (e.g., failure rates, MTTR, incident trends, capacity availability and utilization rates), with visibility into how these trend as the fleet expands
  • Drive alignment on fleet reliability OKRs across engineering and operations teams
  • Identify systemic reliability gaps across hardware, firmware, software, networking, storage, and operational processes
  • Lead complex, cross-functional programs to improve fleet delivery, readiness, and operational stability that scale with fleet growth, avoiding solutions that only work at current size
  • Define program plans, milestones, dependencies, risks, and success criteria for reliability initiatives
  • Proactively manage cross-team dependencies and unblock execution across multiple engineering organizations
  • Track progress against goals, surface risks early, and communicate status clearly to stakeholders and leadership
  • Participate in major incident reviews and root cause analysis, ensuring foll

提及的技能

Engineering OperationsProduction EngineeringNot Applicable

每个职位,都有你的匹配分

BabZituna按六个真实维度,将每个职位与你的资料对照评分,并告诉你为什么得出这个分数,公平性经过审计(阅读公开的偏见审计)。

获取应用 → ✓ 求职者100%免费
一个职位如何评分 示例
技能96经验90地点84工作方式74工作类型61薪资无数据

示例数据,并非真实候选人。每个维度根据你本人的资料按100分制评分;无法衡量的维度会如实标明,而不是猜测。