Jobgether
Jobgether

Software Engineer | AI Training Data & Evals Lab

engineeringfull-timeUS
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
beginner
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities:

    • Build and maintain evaluation harnesses that measure the performance of AI models and agents on real-world tasks.
    • Improve evaluation reliability, coverage, and signal quality through better rubrics, task design support, and scoring approaches.
    • Develop tools that enable researchers and operators to run experiments efficiently without repeatedly rebuilding the same workflows.
    • Build and maintain APIs and backend services supporting human-in-the-loop workflows, task routing, and quality-control processes.
    • Improve data pipelines that transform expert work into structured training and evaluation datasets.
    • Strengthen system observability, scalability, and operational reliability through effective logging, metrics, monitoring, and debugging capabilities.
    • Write clear, maintainable production code and actively participate in code reviews, architecture discussions, and technical design decisions.
    • Document technical decisions and system behavior clearly so that other engineers and collaborators can build upon and operate the systems effectively.
    • Take ownership of core systems from development through production operation, with an expectation of delivering meaningful improvements within the first 30–90 days.
    • Requirements:

      • Strong software engineering fundamentals with professional experience in Node.js and TypeScript.
      • Strong coding ability in Python and/or Go.
      • Demonstrated experience building, deploying, and owning production systems, including APIs, backend services, and data pipelines.
      • Solid understanding of distributed systems, scalability, reliability, and engineering trade-offs.
      • Experience working with AWS or GCP and modern infrastructure technologies such as containers and Kubernetes.
      • Proven track record of shipping and maintaining production systems that other people depend on, rather than working exclusively on prototypes.
      • Strong written communication skills and the ability to collaborate effectively in an asynchronous, distributed environment.
      • Comfortable taking ownership of ambiguous technical problems, making sound engineering decisions, and following projects through to production.
      • Experience with evaluation frameworks, experimentation platforms, or machine-learning tooling is a plus.
      • Experience with data pipelines, workflow orchestration, or internal platforms for research and operations teams is a plus.
      • Experience working in early-stage environments or high-ownership B2B SaaS and platform teams is a plus.
      • Benefits:

        • Full-time, fully remote position with a LATAM focus and meaningful overlap with U.S. time zones.
        • Compensation of $7,000–$10,000 USD per month, based on experience.
        • Significant ownership and opportunities to grow into larger systems, deeper technical leadership, and projects central to the organization’s growth.
        • Lean, async-first working environment focused on clear writing, sound judgment, and strong follow-through.
        • Opportunity to work on research-adjacent engineering challenges at the frontier of AI while building practical production platforms.
        • Direct impact on the training data and evaluation systems used by leading AI labs.
        • Structured hiring process including a practical take-home assignment, team review, technical screen, real-world work trial, and final offer stage.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now