Research Data Scientist
About the role
Scope of the Role:
We are looking for a highly skilled Research Data Scientist – GenAI/LLM to join our AI/LLM Delivery Unit and work on research-driven AI/ML initiatives involving Generative AI, Large Language Models (LLMs), NLP, multimodal AI, model evaluation, and AI data.
The role combines strong research and analytical capabilities with hands-on AI/ML expertise, requiring the candidate to design experiments, develop evaluation methodologies, analyze complex datasets, build research prototypes, and translate research findings into practical AI/ML solutions.
The ideal candidate will have a strong research orientation, excellent statistical and analytical skills, and the ability to work collaboratively with researchers, data scientists, AI/ML engineers, domain experts, and client-facing teams.
What You’ll Own:
AI/ML & Generative AI Research:
- Conduct independent and collaborative research in Generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation, and AI data.
- Formulate research questions and translate complex AI/ML problems into structured research methodologies and experiments.
- Design, execute, and analyze experiments to evaluate and improve AI/ML models and solutions.
- Build analytical models, prototypes, and research pipelines using Python and relevant ML frameworks.
- Stay current with emerging research, methodologies, papers, and developments in GenAI, LLMs, NLP, multimodal models, and AI evaluation.
LLM & Model Evaluation:
- Develop and implement LLM evaluation frameworks, benchmarks, datasets, and evaluation criteria.
- Evaluate models for accuracy, robustness, bias, hallucination, reasoning, relevance, response quality, and other performance dimensions.
- Conduct model benchmarking, error analysis, comparative analysis, and performance evaluation.
- Work on areas such as RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and LLM optimization, as applicable.
- Identify model and data gaps and recommend improvements to enhance model performance and reliability.
Data Science & Statistical Research:
- Collect, clean, analyze, and interpret large and complex structured and unstructured datasets.
- Perform EDA, statistical analysis, hypothesis testing, significance testing, correlation analysis, sampling, and error analysis.
- Develop data-driven insights and identify patterns, trends, and relationships relevant to AI/ML research.
- Apply appropriate statistical and quantitative methodologies to validate research findings
AI Data & Dataset Development:
- Develop and evaluate datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks, and evaluation criteria for AI/ML models.
- Analyze data quality and identify issues affecting model performance.
- Collaborate with annotation, data engineering, and AI/ML teams to improve AI training and evaluation data.
- Translate data and research findings into actionable recommendations for improving AI system
Research & Innovation:
- Contribute to research papers, technical reports, whitepapers, patents, benchmarks, internal publications, and other research outputs, where applicable.
- Identify opportunities to apply emerging research and technologies to real-world AI and data challenges.
- Explore new methodologies, models, datasets, and evaluation approaches to improve AI capabilities.
- Contribute to capability building and innovation within the AI/LLM practice
Collaboration & Stakeholder Engagement:
- Work closely with researchers, data scientists, AI/ML engineers, data/annotation teams, domain experts, and delivery teams.
- Present research findings, analytical insights, and technical recommendations to senior technical stakeholders.
- Translate complex research and technical concepts into clear, actionable recommendations.
- Where required, participate in client-facing technical discussions and presentations and help translate business requirements into AI/ML solutions.
You’ll Thrive in This Role If You Have:
- Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or a related discipline.
- Bachelor’s/Master’s degree from IITs, NITs, or other premier engineering/research institutions is strongly preferred.
- 4–7 years of hands-on research experience in AI/ML, Data Science, NLP, Generative AI, LLMs, or related areas.
- Strong demonstrated research experience with the ability to independently formulate research questions, design experiments, analyze results, and communicate findings.
- Demonstrated research track record through research publications, patents, conference presentations, open-source contributions, or significant AI/ML research projects.
- Candidates with publications in reputed conferences/journals and a strong academic/research profile will be preferred
- Strong proficiency in Python and SQL.
- Strong hands-on experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch/TensorFlow.
- Strong understanding of:
- Machine learning algorithms
- Statistics and experimentation
- Data analysis and feature engineering
- Model evaluation and performance metrics
- Hypothesis testing and statistical inference
- Hands-on exposure to LLMs, NLP, Generative AI, and multimodal AI.
- Experience with one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF/DPO, embeddings, or model benchmarking.
- Experience working with large-scale structured and unstructured datasets.
- Familiarity with Git and cloud platforms such as AWS, Azure, or GCP is desirable.
The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.