Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects

Other Jobs To Apply

No other job posts for this day.

<div>We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long-term technical strategy, organizational structure, and operational roadmap for global, mission-critical infrastructure platforms on the Services Special Projects team.<br/><br/>This position requires a rare blend of deep technical domain expertise-spanning distributed systems, Kubernetes, and AI workload orchestration-and proven organizational leadership managing large, globally distributed engineering teams.<br/><br/><b>Description</b><br/><br/>In this role, you will be responsible for defining and building infrastructure strategy that balances continuous innovation with high reliability, performance, and cost efficiency. You will lead a growing, multi-tiered team of engineers who are responsible for foundational platforms that power large-scale consumer and enterprise worklo.<br/><br/>Beyond operational delivery, you will establish standards for operational excellence, Site Reliability Engineering (SRE), and capacity planning. You will be a key strategic partner, translating complex business imperatives into scalable platform designs while cultivating a strong engineering culture focused on automation, technical ownership, accountability, and continuous improvement.<br/><br/><b>Minimum Qualifications</b><br/><br/>MS Degree in Computer Science or related degree and 12+ years of experience of progressive engineering leadership experience building, scaling, and operating mission-critical infrastructure platforms and global services.<br/><br/>Management & Leadership Scope: 6+ years managing multi-layered engineering organizations (manager-of-managers) with a proven track record of hiring, developing, and retaining top-tier technical talent across global sites.<br/><br/>Cloud & Distributed Compute Expertise: Demonstrated hands-on and architectural mastery of cloud-native infrastructure, Kubernetes platform engineering, and hybrid cloud operations (AWS, GCP, private data centers).<br/><br/>Accelerated Computing & AI Infrastructure: Direct operational and architectural experience running large-scale systems for AI/ML training and inference worklo, including utilization optimization, scheduling, and high-performance storage/networking.<br/><br/>SRE & Production Operations: Deep background in Site Reliability Engineering (SRE) principles, telemetry, observability frameworks, disaster recovery, and managing 24/7 high-availability infrastructure at scale.<br/><br/>Technical Communication: Exceptional ability to seamlessly bridge executive strategy and low-level technical trade-offs-communicating vision to executive stakeholders while driving detailed technical discussions with principal engineers.<br/><br/><b>Preferred Qualifications</b><br/><br/>Large-Scale Enterprise Provenance: Experience leading core infrastructure or foundational platform SRE for a global, tier-1 technology organization operating at massive scale.<br/><br/>Multi-Engine Database & Data Infrastructure: Familiarity overseeing diverse open-source and proprietary storage/data ecosystems (e.g., Cassandra, FoundationDB, Kafka, Redis, PostgreSQL).<br/><br/>Financial & Capacity Governance: Proven competency managing large-scale infrastructure investments, capital expenditures, operational budgets, capacity forecasting, and cloud optimization strategies.</div>

Share Share
Apply Now →