All jobs

VP of Engineering

Gabriel Talent

NewOpen
San Francisco, California, United States Hybrid Engineering full-time Lead / Managerial12+ yrs experience Posted today

About the Company

An open-access AI cloud platform offering affordable, on-demand GPUs, serverless inference, and decentralized compute for developers, researchers, and AI builders worldwide. Roughly 30 people today, at Series A stage.

About the Role

This is the company's first dedicated engineering executive hire, building the Infrastructure, Platform, and SRE functions from the ground up. You'll own infrastructure strategy, organizational growth, and executive-level decision making, but this is explicitly not a step-back-and-manage role. You're expected to stay more than 40% hands-on: personally contributing to architecture reviews, debugging critical production issues, and partnering directly with engineers on implementation.

What You'll Own

  • The design and evolution of the AI cloud platform architecture: GPU orchestration, compute scheduling, networking, storage, and distributed systems
  • Building and scaling large GPU clusters supporting customer workloads, including provisioning, scheduling, utilization optimization, and capacity management
  • Personal participation in architecture reviews, system design, and key technical initiatives
  • Acting as the technical escalation point for complex infrastructure challenges
  • Establishing best practices for Kubernetes, observability, CI/CD, security, and operational excellence
  • Building the SRE and Platform Engineering functions from scratch: SLOs, SLIs, incident response, capacity planning
  • Recruiting and developing the Infrastructure, Platform, and SRE teams
  • Partnering with executive leadership on company strategy and infrastructure investment

What We're Looking For

  • 12+ years building and operating large-scale infrastructure systems, including leadership of infrastructure organizations while remaining deeply hands-on technically
  • Experience building or operating a cloud platform at scale, ideally GPU-native infrastructure supporting AI training and inference workloads, background from a GPU cloud or AI infrastructure company (similar to CoreWeave, Lambda, Modal, Together AI, RunPod, or Crusoe) is the strongest signal here
  • Expert-level Kubernetes knowledge and experience designing multi-region cloud infrastructure
  • Deep expertise in Linux, networking, distributed systems, and storage architecture
  • A proven track record scaling infrastructure in high-growth startup environments, not just maintaining systems at large, established companies. A background that's exclusively enterprise or big tech is a real mismatch here unless paired with meaningful startup infrastructure experience
  • Strong grounding in Infrastructure-as-Code, automation, observability, monitoring, and reliability engineering
  • Experience building highly available production systems with real SLOs and incident response processes
  • Bonus: GPU scheduling experience (Slurm, Kubernetes GPU operators, Ray, or distributed training systems), managing thousands of GPUs in production, or bare-metal provisioning and lifecycle management

Logistics

San Francisco, hybrid. This is an executive-level role with competitive compensation and meaningful equity, Visa: US citizens or green card holders are preferred; open to H-1B transfers.

✨ Not ready to apply? Build a free, ATS-friendly résumé tailored to this role.Open résumé builder →

More open roles at Gabriel Talent

VP of Engineering in San Francisco, California, United States at Gabriel Talent