# Reinforcement Learning Engineer \($400k - $800k salary\)

- Company: [Batoncorporation](<https://jobstar.asia/company/batoncorporation>)
- Location: New York · London
- Team: Engineering
- Employment type: Full Time
- Salary: $400K – $800K
- Posted: May 13, 2026

## Job description

# **Who We Are**

Baton Corporation is the development company that builds and operates the entire technology stack behind [pump.fun](http://pump.fun), the largest memecoin launchpad in production today. The systems are low latency, high throughput, live under constant load, and break if you get them wrong.

# **What You’ll Do**

As our Reinforcement Learning Engineer, you will own a production trading system that directly deploys real capital. This is not a research role - it’s about building learning systems that are robust, measurable, and safe under real-world constraints.

* Own and ship an RL-driven trading agent using real capital to increase trading volume and user participation in a memecoin ecosystem
* Design reward functions and policies aligned with product goals while enforcing strict downside risk constraints
* Build evaluation and validation frameworks (simulation, offline analysis) to minimize reliance on live sequential testing
* Safely transition an existing heuristic-based production system toward learning-based approaches
* Take end-to-end ownership and technical leadership as the sole RL expert, from data and modeling through deployment, monitoring, and safeguards

# Who You Are:

* You have previously put an autonomous learning system into production that directly controlled capital, pricing, traffic, or resources and can explain what broke and how they fixed it
* Have personally designed and enforced hard risk limits (capital caps, loss bounds, circuit breakers) in a live system, not just talked about “risk-aware objectives.
* Have built a policy evaluation loop from scratch (simulators, replay, counterfactuals, shadow deployments) before trusting live rollout.
* Can make and defend uncomfortable tradeoffs (e.g. heuristic > RL, bandit > deep RL) based on empirical results instead of ideology
* Have operated as the single owner of a complex ML system in a small team, with no safety net of research orgs, infra teams, or “ML platforms.”

# What it's like to work here

* We work in person
* Hours can be long and unconventional
* The pace is intense
* Expectations are high, and impact is immediate
* Working at Baton is not for everyone

# Why Join Us?

* Unmatched ownership and autonomy
* Exposure to systems operating at the edge of crypto scale
* The ability to ship fast and see real-world impact immediately

If you’re motivated by responsibility, speed, and building products used by massive audiences, you’ll feel at home here.

## Apply

[Apply on Batoncorporation](<https://jobs.ashbyhq.com/batoncorporation/729d473f-e114-4294-af3c-ed6c05fcba0f>)

Canonical job page: <https://jobstar.asia/job/reinforcement-learning-engineer-400k-800k-salary-batoncorporation-new-york-32eba18fde238ef7>
