Jobgether
Jobgether

Staff ML Software Engineer (L6) — Platform Systems, AIMS Engineering

engineeringfull-timeUS
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Design, build, and operate platform subsystems for observability, evaluation, tooling, and operational automation supporting next-generation ML architectures.
    • Prove new platform capabilities against current production operations, including anomaly detection, root-cause analysis, and automated operational workflows.
    • Build observability infrastructure that provides deep visibility into model behavior, training pipeline health, serving latency, and data quality, enabling teams to detect and diagnose issues before they become incidents.
    • Identify and drive infrastructure cost optimization across ML training and serving, developing increasingly automated frameworks and tools that make compute efficiency a core engineering priority.
    • Architect reliability improvements across the AI/ML stack, reducing operational toil, improving on-call experiences, and establishing strong standards for production excellence.
    • Contribute to the target architecture and migration strategy for a modernized ML platform, coordinating with partner engineering teams and managing technical dependencies.
    • Evaluate emerging infrastructure approaches, model paradigms, and platform capabilities, translating relevant developments into a forward-looking technical roadmap.
    • Develop reusable frameworks and common platform patterns that can be adopted across engineering teams and ML workloads.
    • Lead complex technical programs across organizational boundaries, building alignment and consensus while influencing teams without direct authority.
    • Continuously improve critical systems by balancing long-term platform investments with immediate operational priorities.
    • Requirements:

      • Significant experience designing, building, and operating production AI/ML systems at scale, including ML training pipelines and model-serving or online-inference environments.
      • Hands-on experience building infrastructure for advanced agentic or complex model architectures, such as memory, trace, evaluation, replay, orchestration, or routing systems.
      • Strong software engineering fundamentals with deep expertise in Python and working proficiency in at least one JVM language, such as Java or Scala.
      • Proven experience improving the reliability, scalability, operational maturity, and cost efficiency of AI/ML infrastructure.
      • Strong background building observability and monitoring systems for ML workloads, with an understanding of visibility requirements across training, serving, and data pipelines.
      • Deep distributed-systems expertise, including large-scale batch processing and real-time serving infrastructure.
      • Experience collaborating with multiple engineering and partner teams to establish technical direction, manage dependencies, and deliver complex programs.
      • Excellent technical judgment and the ability to identify reusable patterns, determine where investment is valuable, and make pragmatic decisions about scope and priorities.
      • Ability to operate effectively in ambiguous environments, independently defining problems, developing approaches, and adapting as new information emerges.
      • Preferred experience with LLM evaluation, trace, replay, observability, or debugging tools and frameworks for complex model systems.
      • Familiarity with modern ML infrastructure such as feature stores, model-serving platforms, and experimentation frameworks.
      • Experience migrating production AI/ML systems between technology generations and improving their architecture over time.
      • Experience in personalization domains such as recommendation systems, search, or discovery is a plus.
      • Benefits:

        • Annual salary range of $600,000–$1,066,000, with compensation determined based on factors such as role, location, background, skills, and experience.
        • Annual compensation structure consisting of salary and stock options, with employees able to determine their preferred mix each year.
        • Comprehensive health plans and mental health support.
        • 401(k) retirement plan with employer matching.
        • Stock option program.
        • Health Savings Accounts and Flexible Spending Accounts.
        • Disability programs and life and serious injury benefits.
        • Family-forming benefits.
        • Paid leave programs.
        • Flexible time off for full-time salaried employees.
        • Fully remote position within the United States.
        • Opportunity to work on highly scalable AI/ML systems with significant technical and organizational impact.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Staff ML Software Engineer (L6) — Platform Systems, AIMS Engineering at Jobgether — Remote