Jobgether
Lead ML/AI Platform Engineer
engineeringfull-timeIreland
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more
About the role
Accountabilities
- Own the ML/AI platform: Lead the architecture and operation of training infrastructure, model serving, inference pipelines, model registries, feature pipelines, and production integrations.
- Set technical direction: Partner with the Data Platform Architect to define ML/AI architecture, engineering standards, tooling strategies, and build-versus-buy decisions.
- Build production ML systems: Oversee feature engineering, model training, tuning, registration, deployment, monitoring, and hosted inference using AWS managed services and open-source tooling.
- Drive GenAI and LLM engineering: Develop practical strategies for RAG, prompt engineering, evaluation, fine-tuning, model serving, and agentic workflows while balancing cost, latency, quality, and safety.
- Develop agentic capabilities: Evaluate and implement emerging agent technologies and workflow patterns, including MCP, stateless and stateful architectures, and appropriate guardrails for regulated environments.
- Own the serving layer: Design scalable ML services and API contracts that integrate cleanly with Java microservices, making informed decisions around performance, latency, throughput, and reliability.
- Enable research-to-production: Partner with Data Science to productionize models and create the infrastructure, pipelines, and deployment processes required to move experimentation into reliable production.
- Establish MLOps practices: Drive model monitoring, drift detection, reproducibility, experiment tracking, model registries, and cost observability in collaboration with DevOps.
- Shape the ML roadmap: Work with product, data, and engineering leadership to identify high-impact AI/ML opportunities and translate them into actionable technical roadmaps.
- Provide technical leadership: Mentor senior engineers, guide complex architectural decisions, and represent the AI/ML function in cross-functional discussions.
- Communicate technical strategy: Translate complex architectural trade-offs into clear design documents and executive-level recommendations for both technical and non-technical stakeholders.
- Senior engineering experience: 8+ years in software or ML engineering, including at least 5 years delivering production machine learning systems and owning complex, ambiguous problems end to end.
- Technical leadership: Demonstrated experience shaping ML strategy, mentoring senior engineers, and serving as a trusted decision-maker for challenging architectural problems.
- AWS ML expertise: Hands-on experience with Amazon SageMaker for training, tuning, hosted endpoints, and model registry, as well as Amazon Bedrock and AgentCore for GenAI and agentic applications.
- Open-source ML tooling: Practical experience with JupyterLab, Spark, MLflow, and related machine learning development and experimentation tools.
- AWS platform knowledge: Deep familiarity with services including S3, Athena, Redshift, Glue, Step Functions, and Lambda, combined with strong SQL skills for analytical and ML workloads.
- GenAI/LLM expertise: Production experience with RAG, prompt engineering, evaluation, and the practical trade-offs between quality, cost, latency, and safety.
- Vector search knowledge: Experience with vector databases such as pgvector or Pinecone, together with good judgment about when vector retrieval is appropriate versus alternative approaches.
- Programming skills: Deep Python expertise and strong knowledge of core ML libraries such as scikit-learn, pandas, NumPy, PyTorch and/or TensorFlow, and XGBoost or LightGBM.
- Java familiarity: Sufficient working knowledge of Java to review service code, define API contracts, and troubleshoot integrations with backend microservices.
- MLOps expertise: Strong understanding of model monitoring, drift detection, reproducibility, experiment tracking, model registries, deployment patterns, and cost observability.
- Cross-functional collaboration: Comfortable partnering with Data Science, DevOps, backend engineering, product, and leadership while maintaining clear ownership across shared technical boundaries.
- Communication: Excellent written and verbal communication skills, with the ability to produce both detailed technical designs and concise executive-level recommendations.
- Contracting requirements: Ability to work independently as an independent contractor through your own entity or an approved contracting arrangement, with reliable overlap with US Pacific business hours.
- Nice to have: Experience with ClickHouse or similar analytical databases, LLM fine-tuning techniques such as LoRA or QLoRA, streaming and real-time inference, Kafka or Kinesis, infrastructure-as-code, large-scale ML systems, or open-source ML contributions.
- Equity compensation package.
- Flexible Time Off (FTO) to take time away when needed to rest and recharge.
- Medical, dental, and vision coverage, with 100% of employee premiums covered where applicable.
- Disability and life insurance.
- Learning and career development opportunities within a growing technology environment.
- Remote-work setup reimbursement.
- Monthly phone and internet stipend.
- Team-building events, cultural activities, and company-wide gatherings.
- Paid time off for volunteering and community service.
- Half-day Fridays.
- 401(k) matching contribution.
- Opportunity to work on AI/ML infrastructure supporting products used by more than 1,600 financial institutions.
- Fully remote work environment as part of a globally distributed team.
Requirements:
Benefits:
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply