Jobgether
Senior Data Scientist
datafull-timeUS
SALARY
$140k – $170k/yr
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities
- Lead AI, ML, and NLP initiatives end to end, covering problem framing, solution design, model development, validation, deployment, monitoring, and ongoing production maintenance.
- Develop LLM-powered applications that automate complex, high-volume workflows while balancing accuracy, response time, throughput, and inference costs through model selection, caching, prompt engineering, and related techniques.
- Design evaluation datasets and automated evaluation pipelines to measure model quality, error rates, and performance across clinical and financial content, identifying regressions caused by changes in models, prompts, or data.
- Build and maintain production-grade data foundations, including ETL/ELT pipelines and feature datasets sourced from PostgreSQL transactional systems and Redshift data warehouses.
- Establish strong standards for data quality, reliability, documentation, lineage, and maintainability across analytical and machine learning data workflows.
- Collaborate with clinicians, pharmacists, and other domain experts to establish ground truth, assess edge cases, and validate model behavior against real-world clinical workflows.
- Ensure AI and ML solutions meet appropriate healthcare standards and regulatory requirements, including HIPAA, while monitoring for bias, safety issues, and performance drift.
- Serve as technical lead on complex, ambiguous projects by defining approaches, establishing best practices, setting technical direction, and influencing broader engineering and data strategy.
- Coach and mentor data scientists through technical pairing, design reviews, code reviews, and constructive feedback while contributing to technical hiring and team development.
- Communicate complex analytical results and technical concepts effectively to technical and non-technical audiences, including executive leadership, through presentations, reports, and visualizations.
- Partner closely with software architects and engineering teams to ensure models and data products integrate effectively into production systems and meet platform requirements.
- Master’s or Ph.D. in Computer Science, Data Science, Statistics, Mathematics, or a related quantitative discipline.
- 5+ years of professional experience in data science, machine learning, and NLP, with a demonstrated track record of delivering models into production environments.
- Advanced proficiency in Python and SQL.
- Hands-on experience developing and orchestrating ETL/ELT pipelines using technologies such as dbt, Airflow, and AWS data services including Glue, DMS, Lambda, and S3.
- Strong understanding of incremental data loads, idempotency, data quality testing, documentation, and data lineage.
- Strong software engineering fundamentals, including modular and tested production-quality Python, Git-based workflows, code reviews, Docker, and CI/CD.
- Experience productionizing and monitoring machine learning models in AWS environments, including model registries, ML CI/CD, versioning, and post-deployment monitoring using tools such as SageMaker and MLflow.
- Demonstrated experience building and evaluating LLM-powered applications, including solutions using models such as Claude through Amazon Bedrock, prompt design, retrieval and RAG pipelines, and systematic evaluation of output quality.
- Proven ability to independently scope ambiguous technical problems and deliver solutions from conception through production.
- Demonstrated project leadership and experience coaching or mentoring other data scientists.
- Excellent analytical, problem-solving, communication, and collaboration skills, with the ability to operate effectively in a fast-paced environment.
- Healthcare or insurance industry experience is preferred.
- Familiarity with healthcare data standards such as FHIR, HL7v2, and X12 and clinical vocabularies including RxNorm, NDC, ICD-10, and SNOMED is a plus.
- Experience modeling analytical data for Amazon Redshift or comparable columnar MPP platforms, extracting from PostgreSQL, and writing performant SQL against large datasets is preferred.
- Ability to analyze query plans and optimize slow SQL queries independently.
- Experience deploying containerized applications and models using Docker and cloud-managed infrastructure such as SageMaker endpoints, ECS, or Lambda; Kubernetes experience is a plus.
- Comfortable working primarily at a desk and traveling approximately 10% for client sites, conferences, or internal meetings.
- Competitive base salary of $140,000–$170,000, with exact compensation determined by skills and experience.
- Eligibility for a discretionary performance-based bonus.
- Medical, dental, and vision insurance.
- 401(k) eligibility after three months, with a 50% company match on the first 5% of salary contributed and a three-year vesting schedule.
- Health Savings Account for eligible HDHP participants, with company contributions of up to $500 for individual coverage and $1,000 for family coverage annually.
- 100% company-paid short- and long-term disability, AD&D, and group life insurance.
- 18 days of accrued PTO during the first three years, increasing thereafter.
- 7 paid holidays.
- Employee Assistance Program.
- Up to $1,500 annually in continuing education funding for eligible programs after one year of service.
- Voluntary benefits including FSA, hospital indemnity, accident, and critical illness insurance.
- Remote-first work environment designed to provide flexibility and support collaboration across locations.
- Opportunity to work on complex AI and ML challenges with significant impact on healthcare workflows and patient outcomes.
Requirements
Benefits
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply