Jobgether
Jobgether

Applied ML Engineer

engineeringfull-timeGermany
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities

    • Reproduce and evaluate machine learning research methods using open-weight and API-accessible models.
    • Design evaluation datasets, probes, scoring approaches, baselines, calibration tests, and experiment harnesses.
    • Work directly with model weights, logits, hidden states, activations, model APIs, and inference infrastructure when required.
    • Build and extend evaluation infrastructure covering experiment runners, judges, persistence, orchestration, reporting, and reproducibility.
    • Turn research workflows into intuitive product experiences, including experiment configuration, execution, traces, comparisons, reports, and review workflows.
    • Investigate how verification methods behave when models are modified through fine-tuning, merging, quantization, distillation, safety removal, or deliberate evasion.
    • Design controlled experiments that distinguish meaningful signals from artifacts, confounders, and misleading correlations.
    • Produce clear technical reports that separate measured evidence from interpretation and hypotheses.
    • Deliver production-quality systems with APIs, asynchronous jobs, databases, observability, testing, deployment, and documentation.
    • Contribute across research, experimentation, engineering, and product as priorities evolve.
    • During the first six months, reproduce and document at least one published model-provenance or verification method, including its capabilities, assumptions, and limitations.
    • Build a repeatable model-verification runner with versioned inputs, artifacts, metrics, and reports, and make at least one verification workflow accessible through the product interface.
    • Run controlled experiments across base, fine-tuned, merged, quantized, and known distilled models, improving understanding of when verification methods succeed, fail, and why.
    • Requirements:

      • Strong Python engineering skills, with hands-on experience using PyTorch and Hugging Face Transformers.
      • Solid understanding of machine learning evaluation, including dataset design, baselines, metrics, calibration, false positives and negatives, statistical uncertainty, and reproducibility.
      • Ability to read ML research papers critically and implement methods from first principles rather than relying entirely on existing packages.
      • Professional software engineering experience beyond notebooks, including APIs, asynchronous jobs, databases, logging, testing, deployment, and documentation.
      • Familiarity with open-weight models and a practical understanding of how modern LLM inference systems operate.
      • Ability to work across backend and frontend boundaries, with sufficient React/TypeScript knowledge to help make complex experiments and results understandable to users.
      • Strong experimental and analytical judgment, particularly around distinguishing what evidence demonstrates from what it merely suggests.
      • High ownership and initiative, with the ability to identify problems, propose solutions, and drive projects forward independently.
      • Comfort working in a fast-moving startup environment where priorities can change quickly and engineers may operate across multiple functions.
      • Experience with model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluation, interpretability, or related areas is a plus.
      • Experience with activation and representation analysis, probing, model hooks, logits, hidden states, or other model-internals techniques is advantageous.
      • Familiarity with evaluation and inference infrastructure such as DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, or comparable technologies is beneficial.
      • Experience with Next.js, React, TypeScript, data visualization, or experiment dashboards is a plus.
      • Experience running and serving open-weight models on GPUs, including reasoning about latency, throughput, memory, precision, and cost trade-offs, is valuable.
      • Experience designing adversarial evaluations or testing systems against deliberate attempts to evade detection is an advantage.
      • A strong commitment to producing production-quality code, tests, tooling, and documentation that other engineers can confidently operate and extend.
      • Benefits:

        • Opportunity to work on applied machine learning at the intersection of research, experimentation, engineering, and product.
        • End-to-end ownership across model evaluation, model internals, infrastructure, backend systems, and user-facing experiences.
        • Exposure to modern open-weight models, LLM inference systems, and emerging ML verification techniques.
        • A role with significant technical autonomy and the opportunity to shape both experiments and production systems.
        • Fast-moving startup environment with evolving priorities and cross-functional collaboration.
        • Opportunity to translate cutting-edge research into practical, measurable, and user-accessible products.
        • The opportunity to build systems and evaluation methodologies designed to produce evidence that users can understand and trust.
        • Additional compensation, flexibility, healthcare, and other benefits may be provided according to the partner company's employment package and local arrangements.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now