Jobgether
Jobgether

Senior Inference Engineer

engineeringfull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities

    • Build and deploy production-grade LLM inference systems across one or multiple GPU machines, owning the complete pipeline from customer query to served response.
    • Design, implement, and operate model-serving infrastructure using technologies such as vLLM, SGLang, and TensorRT-LLM.
    • Optimize inference workloads for scale, balancing latency, throughput, reliability, and infrastructure costs.
    • Apply techniques such as quantization, batching, caching, and intelligent request routing to improve inference performance.
    • Develop robust production infrastructure using Python or Golang, with an emphasis on maintainable, scalable engineering rather than configuration-only work.
    • Establish the initial inference platform in close collaboration with the CTO and take ownership of its evolution as the organization scales.
    • Define and execute the technical roadmap for inference infrastructure, identifying opportunities to improve performance, reliability, and developer or customer experience.
    • Partner with Product to translate evolving customer and market requirements into practical inference-platform capabilities.
    • Evaluate emerging inference technologies and approaches and determine where they can create meaningful improvements.
    • Collaborate with engineering stakeholders to establish reliable operational practices for production GPU infrastructure.
    • Requirements:

      • Significant professional experience building and operating production software or infrastructure systems, with strong hands-on engineering capabilities.
      • Demonstrated experience deploying and serving large language models in production, ideally using vLLM, SGLang, TensorRT-LLM, or comparable inference frameworks.
      • Practical expertise optimizing inference workloads through quantization, batching, caching, routing, or similar techniques.
      • Strong programming skills in Python or Golang, with a track record of writing and maintaining production-quality code.
      • Strong understanding of production inference architectures, including the journey from user request through model execution to a reliable served response.
      • Excellent problem-solving skills and the ability to independently investigate complex performance, scalability, and reliability challenges.
      • Strong communication skills, with the ability to explain sophisticated technical concepts clearly to engineers, product stakeholders, and other audiences.
      • Comfortable taking significant ownership, working with ambiguity, and making pragmatic technical decisions in a fast-moving environment.
      • Familiarity with Docker and Kubernetes is a plus.
      • Hands-on experience with generative AI technologies such as PyTorch or Transformers is advantageous.
      • Knowledge of the GPU software stack, including CUDA, NCCL, drivers, and related libraries, is beneficial.
      • Understanding of model architectures and fine-tuning techniques is a plus.
      • Experience with NVIDIA Dynamo is an additional advantage.
      • Benefits:

        • Competitive compensation package including equity.
        • Health, dental, vision, and life insurance, with coverage for eligible dependents where available.
        • Benefits adapted to the country of employment.
        • Flexible working schedule focused on outcomes rather than fixed working hours.
        • High degree of workplace flexibility, supporting remote work and changing personal circumstances.
        • Remote-first environment with a globally distributed team.
        • Significant ownership over the architecture, implementation, and long-term roadmap of the inference platform.
        • Direct collaboration with senior technical leadership and Product teams.
        • Opportunity to work on production-grade GPU and AI infrastructure at scale.
        • Exposure to modern LLM serving, inference optimization, Kubernetes, cloud infrastructure, and open-source technologies.
        • A culture built around ownership, action, continuous improvement, open-source collaboration, and technically ambitious engineering.
        • Opportunity to help define emerging standards for AI infrastructure and production inference.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now