# Software Engineer, Inference

- Company: [Lumaai](<https://jobstar.asia/company/lumaai>)
- Location: Redwood City, CA
- Remote: Yes
- Team: Research & AI
- Employment type: Full Time
- Posted: July 24, 2026

## Job description

You'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs.

This is large-scale inference systems work: scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. It fits a strong systems engineer comfortable with model serving and Kubernetes at scale. If you want pure modeling rather than the systems that run models, this is firmly the systems side.

**What You'll Own**

* Ship new model architectures by integrating them into the inference engine.
* Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
* Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
* Automate, test, and maintain inference services for maximum uptime and reliability.
* Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
* Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.

**First 90 Days**

*One way the first 90 could unfold.*

* **Days 1–30 — Immerse & Diagnose:** Learn the inference stack, the fleets, and where reliability or utilization break.
* **Days 30–60 — Ship & Validate:** Integrate a model or ship tooling/scheduling that improves uptime or GPU utilization.
* **Days 60–90 — Scale & Systemize:** Harden deployment pipelines and scheduling across clusters and providers.

**What You Bring**

* Strong Python and system-architecture skills.
* Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.
* Experience with queues, scheduling, traffic control, and fleet management at scale.
* Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.
* Familiarity with Redis and S3-compatible storage.

**Nice to Have**

* Modern networking stacks including RDMA (RoCE, InfiniBand, NVLink).
* High-performance large-scale ML systems (100+ GPUs).
* CUDA, and FFmpeg or multimedia processing.

*About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.*

## Apply

[Apply on Lumaai](<https://jobs.ashbyhq.com/lumaai/c4b9ff5f-40b2-40d4-9f6f-8b7aad9b8860>)

Canonical job page: <https://jobstar.asia/job/software-engineer-inference-lumaai-redwood-city-10c555f89c7e0a2a>
