# Cluster Engineer

- Company: [Stninc](<https://jobstar.asia/company/stninc>)
- Location: Remote
- Remote: Yes
- Team: AI
- Employment type: Full Time
- Posted: August 6, 2026

## Job description

**Position Summary**  
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization.  
The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.  
  
**Responsibilities**

* Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
* Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
* Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
* Build and support production AI infrastructure running hundreds to thousands of GPUs.
* Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
* Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
* Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
* Configure and tune distributed AI software stacks including:

  + PyTorch
  + NCCL
  + CUDA
  + UCX
  + MPI
  + Slurm
  + Pyxis/Enroot
* Optimize GPU scheduling and resource allocation for both training and inference environments.
* Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
* Identify performance regressions and troubleshoot distributed training issues at scale.
* Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
* Work closely with ML engineers to improve training scalability and inference efficiency.
* Create automation to deploy, validate, benchmark, and monitor GPU clusters.
* Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.

  **Required Qualifications**
* 7+ years designing or operating large-scale Linux infrastructure.
* 5+ years supporting production GPU clusters for AI or HPC workloads.
* Demonstrated experience building multi-node GPU training environments from the ground up.
* Deep expertise with distributed PyTorch training.
* Extensive experience troubleshooting and optimizing NCCL communications.
* Strong understanding of distributed AI communication patterns, including:

  + AllReduce
  + ReduceScatter
  + AllGather
  + Broadcast
  + Point-to-point communications
* Experience benchmarking distributed training using tools such as:

  + nccl-tests
  + NVIDIA DCGM
  + Nsight Systems
  + MLPerf (preferred)
* Strong understanding of GPU memory management, including:

  + KV Cache
  + Activation checkpointing
  + Tensor Parallelism
  + Pipeline Parallelism
  + Data Parallelism
* Experience optimizing LLM inference throughput, including:

  + Tokens/sec optimization
  + Batch sizing
  + Continuous batching
  + KV cache tuning
  + Memory bandwidth optimization
* Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
* Expert-level Linux systems administration skills.
* Experience with Slurm workload manager.
* Experience using Pyxis and Enroot for containerized GPU workloads.
* Strong scripting skills using Python and Bash.

**Technical Expertise**  
**AI Frameworks**

* PyTorch
* CUDA
* NCCL
* Triton (preferred)
* TensorRT-LLM (preferred)

**Cluster Scheduling**

* Slurm
* Pyxis
* Enroot

**GPU Networking**  
**Strong understanding of:**

* InfiniBand
* RoCE v2
* RDMA
* GPUDirect RDMA
* GPUDirect Storage
* UCX
* MPI
* Network topology optimization
* Congestion control
* QoS
* ECN/PFC
* High-speed Ethernet (200/400/800 GbE)

**Storage**  
**Experience designing or tuning storage for AI workloads, including:**

* Parallel file systems
* Distributed storage
* Object storage
* NVMe
* Checkpoint optimization
* Dataset staging
* GPUDirect Storage
* Storage bandwidth optimization
* Metadata performance

**Performance Engineering**  
**Experience with:**

* NCCL benchmarking
* Multi-node scaling analysis
* GPU utilization optimization
* Communication/computation overlap
* NUMA optimization
* CPU affinity
* PCIe topology
* GPU topology (NVLink/NVSwitch)
* Memory bandwidth analysis
* End-to-end performance profiling

**Preferred Qualifications**

* Experience deploying AI workloads on Kubernetes.
* Experience with NVIDIA GPU Operator.
* Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).
* Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
* Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
* Familiarity with MLPerf benchmarking.
* Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
* Experience automating infrastructure using Ansible, Terraform, or similar tools.
* Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.

## Apply

[Apply on Stninc](<https://jobs.ashbyhq.com/stninc/e99fcaf9-584e-48a9-b558-286515147ce5>)

Canonical job page: <https://jobstar.asia/job/cluster-engineer-stninc-ae0b9521478538fe>
