# ML Research Engineer \(Distributed Training\)

- Company: [Metamorphic](<https://jobstar.asia/company/metamorphic>)
- Location: Palo Alto
- Team: Technical Staff
- Employment type: Full Time
- Salary: $200K – $280K
- Posted: March 26, 2026

## Job description

# About Metamorphic

Metamorphic is developing new approaches to intelligence by combining machine learning with large-scale experimental neuroscience, informed by the principles that make the brain efficient, flexible, and robust. We are building foundation models trained on rich, continuous neural data — a high-resolution model of the brain at a scale never before possible.

Our founding team spans machine learning, neuroscience, and neurotechnology, with prior work including the [MICrONS project](https://www.nature.com/immersive/d42859-025-00001-w/index.html), [Neuropixels](http://neuropixels.org/), and the [Enigma project](https://www.enigmaproject.ai/), as well as foundational scientific contributions in learning, neural computation, and generative modeling. Our work sits at the frontier of AI research, and we believe the highest-impact discoveries will come from researchers and engineers working as a single, tightly collaborative team.

The name Metamorphic reflects our belief that the next advances in intelligence will come from a change in form, beyond scale — from artificial to natural intelligence.

# About the Role

We are hiring Research Engineers to join our growing AI research team. You will work on building and scaling the distributed systems that enable training Metamorphic’s state-of-the-art foundation models across thousands of GPU’s. This is a high-impact, technically deep role working at the frontier of ML research and engineering. You will design and optimize our distributed training framework, implement advanced parallelism strategies, build fault-tolerant infrastructure, and provide the tooling researchers need to run large-scale experiments quickly and reproducibly. You'll have substantial autonomy to shape foundational technical decisions on a small, high-impact team.

You'll thrive in this role if you:

* Have significant software engineering experience and can move quickly without sacrificing rigor
* Are able to balance research goals with practical engineering constraints
* Are happy to take on tasks outside your job description to support the team
* Enjoy pair programming and deeply collaborative work
* Are eager to learn more about machine learning research in a novel scientific domain
* Are enthusiastic to work at an organization that functions as a single, cohesive team pursuing large-scale AI research
* Have ambitious goals for AI progress and are excited to create the best outcomes over the long term

We offer:

* The chance to work on one of the most scientifically consequential AI projects being pursued today
* A small, world-class team where your contributions directly shape the science and the company
* Competitive compensation and benefits, along with visa sponsorship
* Strong mentorship and career development

# **Salary Range**

$200,000 - $280,000 USD

Based on experience. We additionally offer a competitive equity package and comprehensive benefits, as well as visa sponsorship for international candidates.

# **Minimum Qualifications**

* Bachelor's degree or equivalent experience in Computer Science, Machine Learning, or a related field
* Strong software engineering skills with a proven track record of building complex systems
* Hands-on experience building and debugging distributed training infrastructure (PyTorch FSDP, DeepSpeed ZeRO, Megatron, TorchTitan, or similar) and optimizing advanced parallelism strategies
* Strong understanding of GPU architecture and performance: memory hierarchy, tensor core utilization, bandwidth vs compute limitations
* Strong understanding of the NVIDIA ecosystem: CUDA, NCCL, NVLink/NVSwitch topologies, mixed-precision training (MXFP8/NVFP4), and profiling tools
* Deep familiarity with PyTorch internals, including torch.distributed, autograd, memory management, and torch.compile
* Experience with cloud/HPC environments and job orchestration across hundreds of GPUs

# **Nice to Have**

* Experience building fault-tolerant training pipelines, including checkpointing, automatic recovery, and infrastructure for reproducible experimentation
* Experience with the latest in mixture-of-experts architectures, diffusion model training, or multimodal models
* Experience with inference serving frameworks (vLLM, TensorRT-LLM) or building custom inference solutions

We encourage you to apply even if you do not believe you meet every single qualification. If you don't see a role that fits, we encourage you to submit a general application and tell us how you'd like to contribute to our mission.

## Apply

[Apply on Metamorphic](<https://jobs.ashbyhq.com/metamorphic/20d7f6a3-d768-40d5-9d40-84fee852e866>)

Canonical job page: <https://jobstar.asia/job/ml-research-engineer-distributed-training-metamorphic-palo-alto-33001ba4722abdf4>
