Jobgether
Jobgether

Research Engineer (Reinforcement Learning)

engineeringfull-timeRomania
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Romania. Join a small, senior engineering team building the next generation of voice- and text-driven AI agents. You’ll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions. Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production deployment. You’ll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior. The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model performance. You’ll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued. Your contributions will directly shape AI systems operating at significant production scale.

Accountabilities

    • Build training environments, verifiers, and supporting infrastructure for post-training models.
    • Own the synthetic data pipeline from data generation through quality assurance and validation.
    • Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements.
    • Design and maintain evaluations that models must pass before production releases.
    • Select and adapt suitable open-weight foundation models for specific agent and product requirements.
    • Develop trained behaviors that perform consistently across both voice and text-based agents.
    • Deploy trained models to production and continuously improve them based on real-world usage and feedback.
    • Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations.
    • Requirements:

      • Strong Python engineering skills and the ability to build reliable, production-quality systems.
      • Demonstrated experience taking a machine learning model from raw data through experimentation and into production.
      • A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage.
      • The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms.
      • Practical experience working with GPUs and a realistic understanding of their capabilities and limitations.
      • Strong judgment around when model training is the right solution—and when a simpler approach is preferable.
      • Ability to collaborate effectively within a remote, distributed, and highly autonomous team.
      • Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable.
      • Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus.
      • Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous.
      • Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable.
      • Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus.
      • Benefits:

        • Opportunity to make a significant impact on a fast-growing developer platform and help shape its future.
        • Collaboration with a small, highly experienced team that values technical excellence, creativity, and ownership.
        • Competitive salary and equity package.
        • Health, dental, and vision benefits.
        • Flexible vacation policy.
        • Remote-friendly working environment with flexibility and autonomy.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now