# Research Engineer

- Company: [Kog](<https://jobstar.asia/company/kog>)
- Location: Paris, France
- Remote: Yes
- Team: Engineering
- Employment type: Full Time
- Posted: June 9, 2026

## Job description

**About Kog**   
  
Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).  
  
We co-design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter-GPU communication overlaps with computation rather than blocking it.  
  
We pre-trained a 2B-parameter DTP model on 6T tokens on 256 H100 GPUs.  
  
We are a team of 11 people, including 10 engineers and 5 PhDs.  
  
Test it at [playground.kog.ai](http://playground.kog.ai). Read the technical details on the [Kog Labs blog](https://blog.kog.ai).  
  
**What you will work on**  
  
You will imagine, design, and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open-weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality.

* Design new model architecture variants, including routing strategies, attention mechanisms, and MoE structure, with execution constraints as a first-order design input.
* Extend the Laneformer thesis by exploring inference-aware architectural variants such as DTP, Ladder Residual, and PT-Transformer, and finding what compounds at scale.
* Own the post-training pipeline across fine-tuning, evaluation methodology, and adaptation of existing open-weight models toward architecture variants optimized for inference speed.
* Scale the stack to large MoE models such as DeepSeek v4 and Qwen 3, working through routing, expert parallelism, and communication patterns at inference time.
* Write up findings as research papers, submit them to top venues, and present them at conferences.
* Contribute to building AI agents that will perform architecture research and training experiments autonomously, starting from the research foundations we are building now.

**What we look for**

* You have designed or changed model architecture, where the structure itself was the object of the work. Showing that work, a paper, a repository, or a thesis, is a requirement to move forward.
* You reason about model design and hardware together, tracing how communication structure and layer dependencies shape inference behavior, with fluency in Transformers and MoE deep enough to weigh trade-offs.
* Stronger signals include inference-aware architectural variants such as DTP, Ladder Residual, or PT-Transformer, and post-training methods such as fine-tuning, preference optimization, or quantization, including at research scale.
* A top engineering school or a PhD with concrete architecture work counts, even without industry experience.

**What we offer**

* Direct access to AMD and NVIDIA datacenter GPUs from day one
* A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions
* Problems that sit on the critical path of model execution speed and that directly influence what the system can become
* A remote-friendly working model, with one mandatory week per month in our Paris office. Travel and accommodation covered by the company.
* Compensation aligned with top AI research profiles, including equity

## Apply

[Apply on Kog](<https://jobs.ashbyhq.com/kog/364b3702-a7a9-431d-8e70-f8fe3f4d2f01>)

Canonical job page: <https://jobstar.asia/job/research-engineer-kog-paris-ab93f096a81d34cb>
