Jobgether
Staff ML Software Engineer (L6) — Platform Systems, AIMS Engineering
engineeringfull-timeUS
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities:
- Design, build, and operate platform subsystems for observability, evaluation, tooling, and operational automation supporting next-generation ML architectures.
- Prove new platform capabilities against current production operations, including anomaly detection, root-cause analysis, and automated operational workflows.
- Build observability infrastructure that provides deep visibility into model behavior, training pipeline health, serving latency, and data quality, enabling teams to detect and diagnose issues before they become incidents.
- Identify and drive infrastructure cost optimization across ML training and serving, developing increasingly automated frameworks and tools that make compute efficiency a core engineering priority.
- Architect reliability improvements across the AI/ML stack, reducing operational toil, improving on-call experiences, and establishing strong standards for production excellence.
- Contribute to the target architecture and migration strategy for a modernized ML platform, coordinating with partner engineering teams and managing technical dependencies.
- Evaluate emerging infrastructure approaches, model paradigms, and platform capabilities, translating relevant developments into a forward-looking technical roadmap.
- Develop reusable frameworks and common platform patterns that can be adopted across engineering teams and ML workloads.
- Lead complex technical programs across organizational boundaries, building alignment and consensus while influencing teams without direct authority.
- Continuously improve critical systems by balancing long-term platform investments with immediate operational priorities.
- Significant experience designing, building, and operating production AI/ML systems at scale, including ML training pipelines and model-serving or online-inference environments.
- Hands-on experience building infrastructure for advanced agentic or complex model architectures, such as memory, trace, evaluation, replay, orchestration, or routing systems.
- Strong software engineering fundamentals with deep expertise in Python and working proficiency in at least one JVM language, such as Java or Scala.
- Proven experience improving the reliability, scalability, operational maturity, and cost efficiency of AI/ML infrastructure.
- Strong background building observability and monitoring systems for ML workloads, with an understanding of visibility requirements across training, serving, and data pipelines.
- Deep distributed-systems expertise, including large-scale batch processing and real-time serving infrastructure.
- Experience collaborating with multiple engineering and partner teams to establish technical direction, manage dependencies, and deliver complex programs.
- Excellent technical judgment and the ability to identify reusable patterns, determine where investment is valuable, and make pragmatic decisions about scope and priorities.
- Ability to operate effectively in ambiguous environments, independently defining problems, developing approaches, and adapting as new information emerges.
- Preferred experience with LLM evaluation, trace, replay, observability, or debugging tools and frameworks for complex model systems.
- Familiarity with modern ML infrastructure such as feature stores, model-serving platforms, and experimentation frameworks.
- Experience migrating production AI/ML systems between technology generations and improving their architecture over time.
- Experience in personalization domains such as recommendation systems, search, or discovery is a plus.
- Annual salary range of $600,000–$1,066,000, with compensation determined based on factors such as role, location, background, skills, and experience.
- Annual compensation structure consisting of salary and stock options, with employees able to determine their preferred mix each year.
- Comprehensive health plans and mental health support.
- 401(k) retirement plan with employer matching.
- Stock option program.
- Health Savings Accounts and Flexible Spending Accounts.
- Disability programs and life and serious injury benefits.
- Family-forming benefits.
- Paid leave programs.
- Flexible time off for full-time salaried employees.
- Fully remote position within the United States.
- Opportunity to work on highly scalable AI/ML systems with significant technical and organizational impact.
Requirements:
Benefits:
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply