# Senior Site Reliability Engineer

- Company: [Lumaai](<https://jobstar.asia/company/lumaai>)
- Location: Redwood City, CA
- Remote: Yes
- Team: Product & Engineering
- Employment type: Full Time
- Salary: $168K – $252K
- Posted: July 24, 2026

## Job description

*Team: Infra Reliability · SF Bay Area / Remote (US)*

You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference clusters reliable and fast, and you help redesign them for the next level of scale.

This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it.

**What You'll Own**

* Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
* Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
* Tune Linux performance deeply, at the OS and kernel level.
* Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
* Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
* Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.

**First 90 Days**

*One way the first 90 could unfold.*

* **Days 1–30 — Immerse & Diagnose:** Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
* **Days 30–60 — Ship & Validate:** Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
* **Days 60–90 — Scale & Systemize:** Contribute to the next-gen re-architecture and harden security and compliance practices.

**What You Bring**

* 5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
* Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
* Working experience with Terraform, Airflow, and Ray.
* Strong experience with AWS or OCI.
* Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
* Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
* Comfort in a less-structured, fast-paced environment.

**Nice to Have**

* Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
* Experience managing large-scale GPU clusters for AI/ML training or inference.
* Familiarity with Kubernetes or orchestration frameworks like Ray.
* Deep expertise in data pipelines and infrastructure.

*About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.*

## Apply

[Apply on Lumaai](<https://jobs.ashbyhq.com/lumaai/75cdb0eb-3cfe-4808-b15e-938e0fbf3bee>)

Canonical job page: <https://jobstar.asia/job/senior-site-reliability-engineer-lumaai-redwood-city-802917a9a5e75ca1>
