← Back to jobsApply for this position
Eclipse
Data Scientist (AI Data & LLM Specialist)
datafull-timeRemote
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
crypto
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Qualifications
- Proven experience as a Data Scientist or Machine Learning Engineer with a focus on data quality and preparation.
- Strong understanding of data labeling methodologies and hands-on experience with data annotation platforms and workflows.
- Demonstrated experience preparing datasets for training and fine-tuning Large Language Models (LLMs), including knowledge of techniques like tokenization, embeddings, and NER.
- Proficiency in Python and common data science libraries (e.g., Pandas, NumPy, Scikit-learn, spaCy, Hugging Face).
- Experience using APIs/SDKs to automate data annotation and active learning loops.
- Excellent communication skills, with an ability to create clear documentation for technical and non-technical audiences.
Responsibilities
- Develop Data Labeling Strategies: Design and document a formal data annotation strategy, including clear, scalable, and efficient guidelines for labeling our data. Define and enforce quality metrics, including inter-annotator agreement.
- Optimize for LLM Consumption: Research, define, and prototype the optimal data formats, structures, and pre-processing steps required for fine-tuning and training LLMs on our datasets.
- Data Quality Analysis: Establish automated processes and metrics to analyze the quality of both raw and labeled data, providing feedback to improve our data collection and labeling workflows.
- Collaborate with Engineering: Work closely with the engineering team to guide the implementation of data processing pipelines and ensure the data infrastructure meets the needs of ML applications.
Nice-to-Haves
- Experience with audio data processing and relevant libraries.
- Familiarity with data annotation platforms and tools.
- Knowledge of modern MLOps principles and practices.
- Experience with large language model data curation and Reinforcement Learning from Human Feedback (RLHF) pipelines.
Join the Eclipse Team
- Opportunity. We believe blockchains should be fast AND highly usable. You’ll do high-impact work to enhance Ethereum’s scalability, shaping the future of crypto.
- Flexibility. We collaborate synchronously and asynchronously, across weekly all-hands meetings, Slack messaging, and quarterly in-person meetups.
- Team. Our founding team has experience launching and scaling blue-chip projects such as dYdX, Uniswap, and zkSync. We’re backed by leading funds and leaders including Polychain, Tribe, Placeholder, DBA, Mustafa Al-Bassam, Tarun Chitra, Meltem Demirors, and others.
- Culture. As an early member of our team, you’ll have a unique opportunity to help shape our culture. We value intellectual honesty, bias towards action, and believe every member plays a key role in achieving our ambitious goals.
- Compensation. You’ll receive a competitive compensation package.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $14.99/mo. Cancel anytime.
Join waitlist