VP of Engineering
Gabriel Talent
About the Company
An open-access AI cloud platform offering affordable, on-demand GPUs, serverless inference, and decentralized compute for developers, researchers, and AI builders worldwide. Roughly 30 people today, at Series A stage.
About the Role
This is the company's first dedicated engineering executive hire, building the Infrastructure, Platform, and SRE functions from the ground up. You'll own infrastructure strategy, organizational growth, and executive-level decision making, but this is explicitly not a step-back-and-manage role. You're expected to stay more than 40% hands-on: personally contributing to architecture reviews, debugging critical production issues, and partnering directly with engineers on implementation.
What You'll Own
- The design and evolution of the AI cloud platform architecture: GPU orchestration, compute scheduling, networking, storage, and distributed systems
- Building and scaling large GPU clusters supporting customer workloads, including provisioning, scheduling, utilization optimization, and capacity management
- Personal participation in architecture reviews, system design, and key technical initiatives
- Acting as the technical escalation point for complex infrastructure challenges
- Establishing best practices for Kubernetes, observability, CI/CD, security, and operational excellence
- Building the SRE and Platform Engineering functions from scratch: SLOs, SLIs, incident response, capacity planning
- Recruiting and developing the Infrastructure, Platform, and SRE teams
- Partnering with executive leadership on company strategy and infrastructure investment
What We're Looking For
- 12+ years building and operating large-scale infrastructure systems, including leadership of infrastructure organizations while remaining deeply hands-on technically
- Experience building or operating a cloud platform at scale, ideally GPU-native infrastructure supporting AI training and inference workloads, background from a GPU cloud or AI infrastructure company (similar to CoreWeave, Lambda, Modal, Together AI, RunPod, or Crusoe) is the strongest signal here
- Expert-level Kubernetes knowledge and experience designing multi-region cloud infrastructure
- Deep expertise in Linux, networking, distributed systems, and storage architecture
- A proven track record scaling infrastructure in high-growth startup environments, not just maintaining systems at large, established companies. A background that's exclusively enterprise or big tech is a real mismatch here unless paired with meaningful startup infrastructure experience
- Strong grounding in Infrastructure-as-Code, automation, observability, monitoring, and reliability engineering
- Experience building highly available production systems with real SLOs and incident response processes
- Bonus: GPU scheduling experience (Slurm, Kubernetes GPU operators, Ray, or distributed training systems), managing thousands of GPUs in production, or bare-metal provisioning and lifecycle management
Logistics
San Francisco, hybrid. This is an executive-level role with competitive compensation and meaningful equity, Visa: US citizens or green card holders are preferred; open to H-1B transfers.
