직무 소개
기업 공고 원문에서 · Temporal · 2026년 9월 28일 게시
Role Summary We’re hiring a Staff Software Engineer to join the Replication Foundations team within Temporal’s Cloud Global Services (CGS) organization. Replication Foundations owns and evolves Temporal’s core replication stack in Temporal OSS—the distributed systems backbone behind key Temporal Cloud capabilities such as High Availability namespaces, cross-cluster and cross-region failover, and migration products that enable customers to move workloads between self-hosted Temporal and Temporal Cloud. The team also builds foundational scalability and reliability mechanisms that support Temporal at scale. In this role, you’ll help set the technical direction for Temporal’s distributed replication systems. You’ll lead complex, correctness-critical initiatives spanning architecture, design, implementation, rollout, and operations. You’ll work across teams to evolve reliable and scalable replication capabilities that support both the open source project and Temporal Cloud. What You'll Do Set the technical direction and evolve the architecture of Temporal’s OSS replication stack, from problem definition through rollout and operational support. Lead the design and implementation of replication protocols that power: High Availability namespaces Cross-cluster and cross-region replication Migration between Temporal clusters, including cloud-to-self-hosted and cloud-to-cloud scenarios Drive scalability and reliability initiatives such as: Multi-cell namespaces Enabling a namespace to span multiple clusters Improving load distribution and handling hot spots Define and communicate system-level guarantees, including consistency models, ordering, idempotency, failure recovery, performance, and operational behavior. Identify architectural risks and opportunities, and shape the technical roadmap for replication capabilities that support current and future cloud products. Partner with Cloud Enablement, CGS, Product, and other engineering teams to align OSS replication foundations with customer and product needs. Lead design reviews, raise the quality of implementation and testing practices, mentor engineers, and provide technical guidance across the organization. Lead or contribute to debugging complex production issues, incident response, and follow-up improvements related to replication and core system behavior. What You'll Bring A track record of designing and delivering complex production distributed systems, including systems where correctness, availability, and performance are critical. Deep understanding of distributed systems fundamentals such as replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery. Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces. Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks. Proficiency writing production-quality
















