Jobgether
Jobgether

Member of Technical Staff | Inference Platform

engineeringfull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities:

    • Evolve and operate a Kubernetes-based online and batch inference runtime supporting production machine learning workloads.

    • Run large-scale batch inference through ephemeral jobs, implementing multi-dimensional admission control across CPU, memory, and GPU resources.

    • Build and extend Kubernetes controllers and custom resources to support reliable model execution and scheduling.

    • Optimize model inference engines and feature-processing pipelines using efficient, vectorized, and columnar operations.

    • Develop efficient mechanisms for serving graphs and data from Lance-based storage.

    • Own the execution of training, post-training, and fine-tuning jobs across both cloud infrastructure and customer-hosted Kubernetes environments.

    • Design and improve autoscaling strategies, GPU serving, inference performance, and infrastructure cost efficiency.

    • Implement comprehensive telemetry and monitoring for model execution, enabling performance, reliability, and cost optimization.

    • Solve complex infrastructure challenges including deterministic job sizing, checkpointing, recovery of batch workloads, difficult input files, and automated profiling of newly accepted models.

    • Design robust approaches to per-customer encryption, workload isolation, and secure execution across shared and customer environments.

    • Improve serving availability, online inference latency, batch throughput, GPU utilization, and cost per prediction or training job.

    • Ensure training and batch workloads complete reliably and on schedule without requiring manual intervention or repeated retries.

    • Write production-quality code, participate in rigorous code reviews, and take operational ownership of the systems you build.

    • Requirements:

      • Professional experience operating model-serving infrastructure or large-scale batch compute workloads on Kubernetes.

      • Strong experience building Kubernetes controllers, operators, or comparable Kubernetes-native infrastructure.

      • Strong software engineering skills with production-quality Python and experience developing reliable, maintainable systems.

      • Demonstrated ability to profile and optimize data-intensive Python pipelines and identify performance bottlenecks.

      • Understanding of distributed systems, workload scheduling, resource allocation, and production infrastructure.

      • Experience with performance optimization across compute, memory, storage, and GPU resources.

      • Strong cost-awareness and the ability to treat infrastructure efficiency as an important product requirement.

      • Experience operating production systems and willingness to take responsibility for the reliability and performance of systems you build.

      • Strong engineering judgment, problem-solving skills, and ability to work effectively on open-ended infrastructure challenges.

      • Familiarity with ML inference workloads and the operational requirements of serving models at scale.

      • Experience with Ray, Ray Serve, or KubeRay in production is a strong advantage.

      • Knowledge of Kueue or other batch scheduling and admission-control technologies is beneficial.

      • Experience with GPU serving and performance optimization is a plus.

      • Familiarity with Arrow, Parquet, Lance, or other columnar data formats is advantageous.

      • Experience shipping and operating software on customer-hosted Kubernetes environments is valuable.

      • Experience with GCP or AWS and platforms such as GKE or EKS is a plus.

      • Experience working in financial services or other regulated environments is beneficial.

      • Benefits:

        • Full-time, fully remote position based in Brazil.

        • Opportunity to own critical production infrastructure powering both real-time and large-scale batch machine learning workloads.

        • Broad technical scope across Kubernetes, ML inference, distributed systems, GPU infrastructure, data processing, and cloud platforms.

        • Hands-on opportunity to build and evolve Kubernetes controllers, scheduling systems, autoscaling infrastructure, and model-serving platforms.

        • Direct impact on measurable engineering and business outcomes, including availability, latency, throughput, GPU utilization, and cost per prediction.

        • Opportunity to solve complex infrastructure challenges involving reliability, recovery, security, encryption, isolation, and multi-tenant execution.

        • Exposure to cloud and customer-hosted environments, including production Kubernetes deployments.

        • Engineering culture centered on ownership, production quality, operational responsibility, and measurable outcomes.

        • Opportunity to work on infrastructure where compute efficiency is treated as a core product capability.

✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now
Member of Technical Staff | Inference Platform at Jobgether — Remote