Jobgether
Data Engineer (Databricks) | Specialist
datafull-timeBrazil
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more
About the role
Accountabilities:
- Build and maintain scalable, reliable data pipelines using modern distributed processing technologies to support high-quality data ingestion, transformation, and delivery.
- Organize and manage data tables using Delta Lake and Unity Catalog, ensuring effective data management and governance.
- Design and implement an automated, native MLOps pipeline on Databricks and Google Cloud Platform (GCP) covering the complete machine learning model lifecycle.
- Develop processes for data preparation, feature engineering, model training, validation, registration, deployment, serving, monitoring, and automated retraining.
- Participate in technical discovery activities, including inventorying existing machine learning models and assessing their migration requirements.
- Develop and validate a standardized MLOps pipeline template through a pilot implementation, followed by progressive migration of models in waves based on business and technical criticality.
- Collaborate with engineering and data teams to ensure solutions are scalable, maintainable, reliable, and aligned with technical standards.
- Work within an agile delivery model, actively participating in sprints, refinement sessions, reviews, retrospectives, and other team rituals.
- Continuously identify opportunities to improve data pipeline performance, automation, reliability, and operational efficiency.
- Demonstrated professional experience working with Databricks and modern data engineering environments.
- Strong hands-on experience with PySpark and Apache Spark for distributed data processing.
- Experience developing and orchestrating workflows using Apache Airflow.
- Practical experience with Google BigQuery and cloud-based data platforms.
- Experience integrating and using MLflow for machine learning lifecycle management.
- Knowledge of AWS Glue and its application within data integration and processing workflows.
- Experience working with both SQL and NoSQL databases, including technologies such as PostgreSQL, MongoDB, and Cassandra.
- Ability to design and maintain scalable data pipelines with a strong focus on data quality, performance, automation, and reliability.
- Experience working in Agile/Scrum environments, including sprint planning, refinement, reviews, and retrospectives.
- Strong analytical and problem-solving abilities, with the capacity to work independently and collaboratively on complex technical challenges.
- Nice to have: Knowledge of Apache Kafka for event streaming and real-time data architectures.
- Nice to have: Experience with dbt for data transformation and analytics engineering.
- Opportunity to work with modern data engineering, AI, cloud, and MLOps technologies.
- Exposure to large-scale Databricks and GCP-based data environments.
- Opportunity to contribute to end-to-end machine learning lifecycle automation and reusable technical solutions.
- Collaborative and agile working environment.
- Continuous learning and professional development opportunities.
- Exposure to emerging trends in Artificial Intelligence, Generative AI, and advanced technology.
- Opportunity to work on technically challenging projects with meaningful business impact.
- Career growth within a technology-focused and innovation-driven environment.
- Compensation and benefits package aligned with the role and local market
Requirements
Benefits
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply