Jobgether
Databricks Engineer - Senior/ Lead
engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities:
- Design, develop, and maintain scalable data pipelines and data engineering solutions using Azure Databricks.
- Build robust ETL/ELT workflows using Python, PySpark, Spark SQL, and Apache Spark.
- Develop complex SQL queries and transformations to process and prepare large datasets for analytics and machine learning.
- Optimize Spark jobs and distributed processing workloads for performance, scalability, reliability, and cloud cost efficiency.
- Work with structured and semi-structured data formats including Parquet, Delta, JSON, and CSV.
- Design, build, and maintain Delta Lake tables, leveraging capabilities such as ACID transactions, time travel, and schema evolution.
- Integrate Databricks workloads with Azure Data Lake Storage Gen2 and other Azure data services.
- Apply distributed computing principles to efficiently process large-scale datasets.
- Implement data quality, validation, monitoring, and error-handling processes across data pipelines.
- Collaborate with data scientists, analysts, and business stakeholders to understand requirements and deliver reliable data solutions for analytics and ML use cases.
- Apply Azure security, access control, governance, and RBAC best practices across data engineering environments.
- Use Git-based version control and follow collaborative software development practices.
- Contribute to CI/CD pipelines and automated deployment processes using tools such as Azure DevOps or GitHub Actions.
- Support data modeling initiatives and help ensure data structures are optimized for downstream analytics and reporting.
- Contribute to streaming data solutions using technologies such as Spark Structured Streaming, Azure Event Hubs, or Kafka.
- Identify opportunities to improve pipeline architecture, engineering standards, performance, and operational efficiency.
- Provide technical guidance and contribute to engineering best practices appropriate for a Senior or Lead-level position.
- 6+ years of professional experience in Data Engineering or a closely related field.
- Strong hands-on experience designing and developing solutions with Azure Databricks.
- Advanced proficiency in Python for data processing, transformation, and pipeline development.
- Strong SQL skills, including complex joins, window functions, data transformations, and query performance optimization.
- Hands-on expertise with Apache Spark and PySpark, including distributed data processing and Spark job optimization.
- Solid experience working with Delta Lake, including ACID transactions, time travel, and schema evolution.
- Strong knowledge of Azure Data Lake Storage Gen2 and cloud-based data architecture.
- Understanding of distributed computing concepts and the ability to design solutions for large-scale data workloads.
- Experience using Git for version control and collaborative development.
- Strong understanding of modern ETL/ELT patterns, data pipeline architecture, and data engineering best practices.
- Experience with Azure Data Factory is preferred.
- Exposure to CI/CD practices and tools such as Azure DevOps or GitHub Actions is an advantage.
- Basic understanding of data modeling and its application to analytical workloads.
- Familiarity with Azure security principles, RBAC, access control, and data governance.
- Exposure to real-time and streaming data technologies such as Spark Structured Streaming, Event Hubs, or Kafka is a plus.
- Strong analytical and problem-solving abilities, with a structured approach to debugging and performance optimization.
- Strong communication and collaboration skills, with the ability to work effectively with technical teams, data scientists, analysts, and business stakeholders.
- Ability to take ownership of complex engineering initiatives and provide technical direction at a senior or lead level.
- Opportunity to work on large-scale data engineering and analytics initiatives within an AI and technology-focused environment.
- Hands-on experience with Microsoft Azure, Azure Databricks, Apache Spark, Delta Lake, and modern cloud data platforms.
- Opportunity to design and optimize production-grade data pipelines handling complex and large-scale datasets.
- Exposure to advanced distributed data processing, cloud architecture, and modern ETL/ELT practices.
- Opportunity to contribute to analytics and machine learning use cases in collaboration with data scientists and analysts.
- Exposure to CI/CD, cloud security, data governance, and streaming data technologies.
- Senior/Lead-level scope with opportunities to influence engineering standards and technical architecture.
- Collaborative environment involving data engineering, analytics, data science, and business stakeholders.
- Opportunities for continued technical growth across cloud data engineering, AI, and emerging data technologies.
Requirements:
Benefits:
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply