Đăng nhậpTải ứng dụng

Việc làm › Tin đăng

Senior Site Reliability Engineer (SRE & AI Platform Operations)

Helloprint · Rotterdam, Zuid-Holland, Netherlands

Làm tại chỗTiếng Anh

Về vị trí này

Từ tin đăng của nhà tuyển dụng · Helloprint · đăng ngày 2 tháng 9, 2026

HelloPrint is mid-transformation. The entire platform is being rebuilt from the ground up: new frontend, new pricing engine, new content engine, new product engine. Everything agent-ready. We ship in a week what used to take a year. What we do not yet have is someone who owns the reliability of all of it: the SLOs, the cost, the AI runtime, and the guardrails that let product engineers move fast without breaking things. That is this role. Core Objective: Take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure across high-velocity microservices and cloud workloads. What you will do: SLOs & Error Budgets: Define, track, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error-budget policies across core customer journeys and critical services (including checkout, payments, catalog pipelines, and routing engines). Distributed Observability & Telemetry: Expand Telemetry, Sentry + Google Cloud Monitoring distributed tracing, and automated diagnostic tooling across microservices, Laravel Horizon queue workers on Redis, and Google Cloud infrastructure. Deployment Safety & Progressive Delivery: Partner with product teams to evolve CI/CD pipelines with canary traffic shifting, automated SLO-driven rollbacks, and automated health gates in GitHub Actions. AI Platform Operations & FinOps: Architect, monitor, and scale the runtime infrastructure supporting AI agents, semantic pipelines, and background automation. Take full ownership of runtime cost control, model and token budget tracking, latency profiles, rate limits, queue backpressure, and provider availability. Cloud Infrastructure Cost & Budget Optimization: Partner with engineering leadership to drive continuous FinOps practices across Google Cloud workloads, actively identifying resource inefficiencies, optimizing compute/storage footprint, and enforcing infrastructure budget guardrails. Incident Management & Post-Mortems: Lead on-call incident response and blameless post-mortems, systematically turning root causes into automated tests, synthetic checks, and architectural guardrails. Resilience, Capacity & DR: Drive capacity forecasting, dependency isolation, automated load testing, and disaster recovery validations against strict RTO/RPO targets. Infrastructure as Code & Security: Own declarative infrastructure workflows using Terraform and Google Cloud Run, ensuring strict IAM least privilege, Secret Manager, and deterministic environments. Toil Elimination & DevEx: Build pragmatic internal tooling, runbooks, and self-service deployment primitives that eliminate firefighting and enable product engineers to ship reliably. What we are looking for: Production SRE & Platform Background: Proven experience operating and scaling high-traffic distributed production systems where uptime, low latency, and safe release cycles are paramount. Deep

Kỹ năng được nhắc đến

In OfficeRotterdam

Xem điểm phù hợp của bạn cho mọi vị trí

BabZituna chấm điểm mọi công việc so với hồ sơ của bạn trên sáu tiêu chí thực tế và cho bạn thấy TẠI SAO lại có điểm đó, đã được kiểm định về tính công bằng (đọc báo cáo kiểm định thiên lệch công khai).

Tải ứng dụng → ✓ Miễn phí 100% cho người tìm việc
Một công việc được chấm điểm thế nào Ví dụ
Kỹ năng96Kinh nghiệm90Địa điểm84Hình thức làm việc74Loại công việc61Mức lươngKhông có dữ liệu

Số liệu minh họa, không phải ứng viên thật. Mỗi tiêu chí được chấm trên thang 100 từ chính hồ sơ của bạn, và tiêu chí nào chúng tôi không đo được sẽ ghi rõ thay vì phỏng đoán.