Spotify
Multilingual AI Quality Specialist
datafull-timeStockholm
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
What You'll Do:
- Define quality frameworks, evaluation rubrics, thresholds, and methodologies for multilingual AI experiences.
- Design and execute structured evaluations for AI-generated, AI-translated, AI-curated, and recommendation-driven experiences.
- Lead multilingual dataset curation, annotation, enrichment, and ground-truth creation to support AI model development and evaluation.
- Analyze evaluation results, identify quality gaps, and provide actionable recommendations to improve multilingual AI quality.
- Support LLM-as-a-judge workflows, evaluator calibration, and human-AI agreement studies.
- Partner closely with Product, Engineering, Data Science, Research, Localization, vendors, and market experts to improve AI quality signals and inform launch decisions.
- Document best practices and help define quality standards across languages, markets, and AI use cases.
- Contribute to building scalable evaluation capabilities that support the next generation of AI-powered experiences across Spotify.
Who You Are:
- You have experience in multilingual quality evaluation, localization, data curation, annotation, AI evaluation, or related fields, including text-to-text and text-to-speech experiences.
- You understand language quality, cultural relevance, content quality, and user experience across multiple languages and markets.
- You have experience designing or conducting structured evaluations using quality rubrics, audits, annotation projects, or review methodologies.
- You are comfortable using qualitative and quantitative data to identify trends, measure quality, and make recommendations.
- You are familiar with large language models (LLMs), generative AI evaluation, human-in-the-loop workflows, or LLM-as-a-judge methodologies.
- You enjoy working through ambiguity and turning complex quality challenges into practical evaluation strategies.
- You communicate effectively and thrive in highly cross-functional environments, collaborating with technical and non-technical partners alike.
- Experience with recommendation systems, personalization, search, ranking, machine translation, generative AI, dataset creation, annotation operations, evaluator calibration, prompt testing, model evaluation, SQL, Python, dashboards, or annotation platforms is a plus.
Where You'll Be:
- This role is based in London or Stockholm.
- We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply