Skip to main content

Technical Program Manager - AI Infrastructure / GPU Clusters

Job Description

About the Company

\n

\n

GMI Cloud is building next- AI infrastructure designed for large-scale GPU training and inference workloads. Our platform supports high-density GPU clusters deployed in modern data centers across multiple regions.

\n

\n

\n

About the Role

\n

\n

We are looking for a Technical Program Manager (TPM) to drive the deployment and delivery of GPU cluster infrastructure. This role will work at the intersection of AI hardware platforms, high-performance networking, and data center infrastructure, coordinating across solution architects, engineering teams, vendors, and contractors to deliver production-ready AI clusters.

\n

\n

\n

Responsibilities

\n

\n

GPU Cluster Deployment

\n

    \n
  • Lead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.
  • \n

  • Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.
  • \n

  • Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.
  • \n

\n

\n

Infrastructure Architecture Collaboration

\n

    \n
  • Work closely with Infrastructure Solution Architects (SA) to define:
  • \n

  • GPU server platform selection
  • \n

  • Network architecture for distributed GPU clusters
  • \n

  • Storage integration and cluster infrastructure design
  • \n

  • Support development of the cluster Bill of Materials (BOM) including compute, networking, storage, and supporting infrastructure components.
  • \n

  • Ensure architecture decisions align with data center constraints such as power density, cooling capacity, and rack layout.
  • \n

\n

\n

System Integration

\n

    \n
  • Drive system integration for large-scale GPU clusters, including:
  • \n

  • Rack elevation planning
  • \n

  • GPU server deployment and configuration
  • \n

  • High-speed network topology implementation
  • \n

  • Power and cooling readiness
  • \n

  • Ensure deployments align with vendor reference architectures and validated cluster designs.
  • \n

\n

\n

Contractor & Field Deployment Management

\n

    \n
  • Work closely with General Contractors (GC) and system integrators to manage on-site infrastructure implementation.
  • \n

  • Lead contractor onboarding, including SOW development, scope definition, and delivery milestone alignment.
  • \n

  • Coordinate and oversee field deployment activities such as:
  • \n

  • Structured cabling installation
  • \n

  • Rack installation and equipment mounting
  • \n

  • Network and power connectivity preparation
  • \n

  • Hardware staging and deployment logistics
  • \n

\n

\n

Cluster Validation & Performance Testing

\n

    \n
  • Coordinate cluster bring-up and validation activities including:
  • \n

  • Single-node GPU validation
  • \n

  • Multi-node cluster deployment
  • \n

  • GPU interconnect validation (P2P, RDMA)
  • \n

  • Drive cluster benchmarking, stress testing, and performance verification before production release.
  • \n

\n

\n

Operational Readiness

\n

    \n
  • Ensure deployed GPU clusters are fully ready for production workloads by driving:
  • \n

  • Hardware and network validation
  • \n

  • Monitoring and telemetry integration
  • \n

  • Operational documentation and runbooks
  • \n

  • Handover to operations teams
  • \n

\n

\n

\n

Qualifications

\n

\n

    \n
  • 5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure delivery
  • \n

  • Experience with GPU cluster deployments or high-performance computing environments
  • \n

  • Familiarity with GPU server architecture and distributed computing infrastructure
  • \n

  • Experience working with Infrastructure Solution Architects to define system architecture and hardware BOM
  • \n

  • Experience managing data center hardware deployments and system integration
  • \n

  • Ability to coordinate multi-vendor infrastructure projects across regions
  • \n

\n

\n

\n

Required Skills

\n

    \n
  • Experience deploying large-scale AI infrastructure or GPU clusters
  • \n

  • Familiarity with:
  • \n

  • InfiniBand / RoCE / high-speed Ethernet networking
  • \n

  • GPU interconnect validation (P2P / RDMA)
  • \n

  • Rack elevation and high-density rack deployment
  • \n

  • Experience with cluster validation and performance benchmarking
  • \n

  • Background as Systems Engineer, HPC Engineer, or Infrastructure Architect
  • \n

  • Experience working in AI infrastructure, cloud infrastructure, or hyperscale data centers
  • \n

\n

\n

\n

Skills

\n

    \n
  • Experience deploying liquid-cooled GPU clusters or high-power racks
  • \n

  • Experience working with NVIDIA AI infrastructure platforms
  • \n

  • Familiarity with AI training environments and distributed workloads
  • \n