Jobgether
Jobgether

AI Engineer – Trust & Explainability (AI Platform)

engineeringfull-timeUS
SALARY
$104k – $178k/yr
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities

    • Build end-to-end tracing across AI platform components, including gateways, orchestration, memory, tools, model calls, agent handoffs, parallel branches, and retries.

    • Develop correlation capabilities that connect actions across multiple agents into coherent, readable workflow traces.

    • Build developer-facing trace and debugging experiences that allow engineers to understand complete agent interactions.

    • Develop explanation capabilities that transform raw trace information into human-readable accounts of what an agent did, what information it relied on, and why it followed a particular path.

    • Create platform primitives for customer-facing trust, including explanation records, confidence and provenance metadata, and summaries of information used by agents.

    • Partner with product engineering teams to integrate trust and explainability capabilities into production AI applications and iterate based on feedback.

    • Evaluate and integrate open-source observability, tracing, and evaluation frameworks, extending them when existing capabilities do not meet platform requirements.

    • Build missing trust and explainability capabilities and contribute useful fixes or extensions to open-source projects where appropriate.

    • Develop and maintain evaluation tooling covering golden datasets, test runners, scoring pipelines, regression reporting, and model or prompt comparisons.

    • Create processes for generating and versioning evaluation datasets using de-identified traffic and synthetic scenarios.

    • Build automated quality checks for model, prompt, and tool changes so regressions are identified before production release.

    • Develop adversarial and red-team testing for prompt injection, jailbreaks, tool misuse, and data-exfiltration risks.

    • Build automated tenant-isolation tests to ensure agents cannot access another customer's data through memory, retrieval, tools, or model context.

    • Collaborate with Security Operations to incorporate threat models and emerging attack patterns into platform testing.

    • Participate in design discussions and code reviews, support onboarding of junior AI engineers, and contribute documentation that reduces tribal knowledge.

    • Requirements

      • 3+ years of professional software engineering experience delivering production features independently.

      • Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.

      • Strong foundations in algorithms, data structures, software design, and modern engineering practices.

      • Production-level proficiency with Python; TypeScript experience is a plus.

      • Hands-on experience building applications that integrate LLMs, such as LLM APIs, agent frameworks, RAG pipelines, or comparable systems, either professionally or through substantial personal or open-source work.

      • Active daily use of AI-assisted software development tools and an interest in applying them effectively to engineering workflows.

      • Experience with Git, Docker, automated testing, and modern scripting or development tooling.

      • Hands-on experience in at least one area such as LLM observability and tracing, LLM evaluation and testing, agent frameworks and multi-agent orchestration, or application security testing, with an interest in developing broader expertise.

      • Experience with distributed tracing or observability tooling in a production environment.

      • Strong automated testing practices, particularly for systems where outputs can vary between runs.

      • Experience running workloads on AWS or Azure, including foundational knowledge of identity and access management, networking, and secrets management.

      • Strong analytical and problem-solving skills, with the ability to work independently while knowing when to seek guidance on complex system design.

      • Clear written and verbal communication skills and the ability to collaborate effectively across engineering, product, architecture, and security teams.

      • Preferred experience with OpenTelemetry, GenAI semantic conventions, OpenLLMetry, or comparable tracing standards.

      • Familiarity with LLM observability and evaluation platforms such as Langfuse, Arize Phoenix, LangSmith, Braintrust, promptfoo, DeepEval, or equivalent tools.

      • Experience contributing to open-source AI observability, evaluation, or agent-framework projects is a plus.

      • Experience with AWS Bedrock, Azure OpenAI, or other cloud-managed model services is beneficial.

      • Experience building multi-tenant SaaS platforms where tenant isolation is a critical security requirement is preferred.

      • Experience in financial services, fintech, or another regulated environment where explainability influenced technical decisions is advantageous.

      • Experience building developer-facing debugging, monitoring, or visualization tools is a plus.

      • Benefits

        • Annual compensation range of $104,148–$177,600, with final compensation determined by experience, qualifications, and other job-related factors.

        • Fully remote work within the United States.

        • Medical, dental, vision, life, and disability insurance coverage.

        • Flexible paid time off.

        • Paid company holidays.

        • 401(k) plan with company matching.

        • Opportunity to work on AI platform infrastructure spanning agent tracing, explainability, evaluation, observability, and security.

        • Exposure to emerging AI engineering practices and open-source observability and evaluation frameworks.

        • Collaborative environment with opportunities for technical growth alongside experienced engineers and architects.

        • Role supporting production AI systems in a regulated financial-services environment.

        • Background, credit, and drug screening may be required as part of the hiring process.

✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now