Jobgether
Jobgether

Staff Applied AI Engineer, Product & Agent Performance

engineeringfull-timeUS
SALARY
$175k – $200k/yr
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities

    • Design, implement, and continuously improve agent behavior across live, long-horizon, multi-turn, and multi-agent workflows.
    • Architect retrieval and context strategies that deliver the right source data to models in the right structure while keeping agents grounded in reliable information.
    • Design memory and state-management approaches for multi-turn and multi-agent experiences, determining what information should be retained, summarized, or discarded.
    • Develop prompt and context templates using few-shot examples, structured formats, reasoning scaffolding, and other techniques to create consistent agent behavior.
    • Improve agent performance through experimentation with prompting, tool-use strategies, retrieval, and context construction rather than relying on assumptions.
    • Build production-representative evaluation suites and regression checks to measure accuracy, reliability, regressions, failure modes, edge cases, latency, and cost.
    • Create evaluation rubrics, quality heuristics, and performance thresholds that account for the severity and business or safety impact of failures, not simply their frequency.
    • Design and validate escalation mechanisms that route uncertain or high-risk cases to human review while maintaining safe and consistent behavior.
    • Establish cost-aware approaches to AI performance, balancing accuracy, reliability, latency, context efficiency, and tool-call usage.
    • Baseline existing behavior, conduct comparative evaluations, and assess model or system changes to make evidence-based go/no-go recommendations before customer release.
    • Maintain product-level AI documentation, including model cards, intended-use guidance, limitations, known failure modes, and performance information.
    • Partner closely with Product and Engineering to ensure agentic systems are not only capable but also steerable, trustworthy, transparent, and scalable.
    • Translate production failures and performance evidence into clear diagnoses, experiments, fixes, and actionable recommendations for cross-functional teams.
    • Requirements

      • 8+ years of production software engineering experience, including at least 3 years of hands-on ownership of ML, LLM, or agentic systems in production.
      • Professional experience working with AI systems in healthcare, finance, or another regulated environment where reliability, safety, and transparency are important.
      • Demonstrated ability to diagnose agent failures and determine whether improvements should come from instructions, retrieval, context, memory, tool use, or other system components.
      • Strong understanding of how to evaluate AI failures based on severity, risk, and cost rather than frequency alone.
      • Hands-on experience designing and implementing RAG architectures and production-grounded evaluation frameworks.
      • Experience developing fallback mechanisms, human-in-the-loop workflows, escalation logic, or comparable safety mechanisms for automated systems.
      • Practical familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build, test, and evaluate AI systems in production environments.
      • Strong software engineering foundations and the ability to move comfortably between technical implementation, experimentation, evaluation, and product-level decision-making.
      • Evidence-driven judgment and confidence to challenge launch decisions when performance or safety standards are not met.
      • A strong builder mentality, with the ability to move quickly from identifying a production problem to designing an experiment, validating a solution, and implementing a fix.
      • Equivalent practical experience demonstrating staff-level technical depth may be considered in place of specific educational credentials.
      • Experience applying AI to healthcare data or workflows, particularly where calibrated uncertainty and transparency affect clinicians, care teams, or patients, is highly valued.
      • Experience with long-horizon, multi-turn, or multi-agent systems and product-level AI documentation such as model cards is a plus.
      • Benefits

        • Competitive compensation: $175,000–$200,000 per year.
        • Meaningful technical ownership: Define how AI agent performance, safety, reliability, and production readiness are measured.
        • High-impact scope: Influence prompts, retrieval, context, memory, evaluations, escalation patterns, and AI product decisions at scale.
        • Mission-driven work: Contribute to technology designed to improve healthcare delivery and patient outcomes.
        • Remote-friendly culture: Flexible working environment designed to support distributed collaboration.
        • Professional development: Employee-driven programs and initiatives supporting personal and career growth.
        • Collaborative community: Work alongside a diverse, talented, energized, and purpose-driven team.
        • Cross-functional exposure: Partner closely with Product and Engineering while translating production evidence into improvements across AI-powered workflows.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now