RemNavi

Remote position

ML Infrastructure Engineer

at Npv

Apply on Npv →

New remote Data jobs that show the pay — weekly, free, unsubscribe anytime.

May require office time ● Posted yesterday Paris Data

About this role

We're looking for an ML Infrastructure Engineer to join White Circle, an AI Safety company building the policy enforcement and optimization layer for AI systems. Backed by $11M from senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, and DeepMind, White Circle processes 100M+ API calls monthly and runs its own LLMs in production.

You will

  • Build scalable RL and post-training pipelines, including smoke tuning runs for quality testing and ablations.

  • Design data control systems for rollouts, replay, filtering, evaluation, and policy updates.

  • Tune training and inference end-to-end for throughput: networking, memory, scheduling, data loading, storage, checkpointing, I/O.

  • Build infrastructure for model iteration (experiment runs, artifacts, evals, dashboards, reproducibility, cost visibility) and inference infrastructure for post-training and eval loops.

  • Build agentic development environments: coding-agent harnesses, tool integrations, runtime sandboxes, multi-agent orchestration.

Requirements

  • Hands-on experience designing and running distributed RL/post-training systems at scale (rollouts, replay buffers, reward signals, policy updates, eval loops).

  • Strong Python (concurrency, async, multiprocessing, performance optimization) and PyTorch or JAX.

  • Debugging distributed GPU workloads across CUDA, drivers, containers, NCCL, networking, storage, and checkpointing.

  • Profiling across the stack (py-spy, PyTorch profiler, Nsight, perf, tracing).

  • Inference stacks: vLLM, SGLang, TensorRT-LLM, Dynamo, or custom serving.

  • Ability to connect system metrics to model behavior and learning dynamics.

  • Relocation to Paris (hybrid) required.

Bonus

  • Public builder footprint: open-source contributions to RL, distributed ML, inference, eval, or agent infra; active technical presence on X.

  • Experience at high-bar AI infra/research teams (xAI, Qwen, ByteDance, Prime Intellect, or similar).

  • Ownership of custom training frameworks, trainers, schedulers, or data loaders.

  • GPU clusters on Kubernetes, Slurm, Ray; NCCL, RDMA, InfiniBand, RoCE, or EFA.

  • Rust, C++, CUDA, or Go; serious use of agentic coding tools (Claude Code, Codex, or similar).

We offer

  • Competitive salary + equity.

  • Hybrid work from Paris with relocation package.

  • Top-tier medical insurance in France and flexible time off.

  • L&D budget, all hardware and tools you need, plus covered AI agent and IDE subscriptions.

  • Team off-sites twice a year.

Find Jobs in France on Arbeitnow

Real Remote Score 32/100 · see the full breakdown

Real Remote Score

32/100

Weak

Comp
0/25
Location
4/25
Source
5/15
Clarity
3/15
Freshness
20/20
Why this score?
  • Compensation · No salary disclosed 0/25
  • Location · Specific city or narrow scope 4/25
  • Source · Generic aggregator 5/15
  • Role clarity · Neither seniority nor stack in title 3/15
  • Freshness · Posted yesterday 20/20

How the Real Remote Score is calculated → · Score appeals & corrections

Hybrid Transparency Score 15/100 · see the full breakdown

Hybrid Transparency Score

15/100

Weak

Days
0/30
Location
0/30
Schedule
0/15
Relocation
15/15
Source
0/10

This role is hybrid: it expects some in-office presence. HTS grades how clearly the employer discloses the hybrid terms. How the Hybrid Transparency Score works →

Posted via Arbeitnow. Applications are handled by Npv. RemNavi earns no commission.

Apply on Npv →

New remote Data jobs that show the pay

Weekly, free, unsubscribe anytime. We link straight to the employer — we never take your application.

Compare this role