Ryzlabs
Ryzlabs

Senior Data Engineer

datafull-timeArgentina
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

About the Role

We are looking for a Senior Data Engineer to join a high-performing data engineering team at a leading global company operating at the intersection of sports, technology, and digital commerce.

In this role, you will take ownership of the systems responsible for ingesting, transforming, validating, and publishing data across a large-scale data ecosystem. You will work at the data ingestion boundary, where information from multiple internal and external sources enters the platform, ensuring it is accurate, reliable, and ready to power downstream products, analytics, and services.

You will be responsible for building and maintaining scalable data pipelines, solving complex identity and entity-matching challenges, identifying and addressing data-quality issues, and improving the reliability and observability of data flows across the organization.

This is a highly hands-on engineering role that combines AWS data engineering, Python, SQL, event-driven architectures, data quality, and identity resolution. You will work closely with engineering, product, analytics, and downstream platform teams to troubleshoot complex data challenges, improve existing systems, and build reliable solutions that support data-driven products at scale.

What You'll Do

Own and evolve data ingestion pipelines that bring data from multiple external and internal sources into the platform, including ingestion, cleaning, curation, entity resolution, and event-driven publishing.

Design, build, and maintain scalable batch and event-driven data workflows using Python, SQL, and AWS.

Solve complex identity resolution, entity matching, and deduplication challenges across multiple data sources, ensuring reliable and consistent entity mappings.

Build and maintain data quality and validation frameworks, including freshness monitoring, null-rate and consistency checks, and schema-drift detection.

Monitor data at the ingestion boundary and proactively identify, troubleshoot, and resolve issues before they impact downstream products and services.

Partner with engineering teams responsible for event infrastructure and downstream identity services to trace data and events end-to-end.

Collaborate with Product, Assessments, Analytics, and Engineering teams to understand how data is consumed downstream and ensure pipelines meet evolving product and business requirements.

Work with AWS services such as Glue, Athena, S3, and DynamoDB to build and operate reliable, scalable data infrastructure.

Automate infrastructure and pipeline changes using Infrastructure as Code, primarily Terraform or AWS CDK.

Participate in production incident response, quickly assessing impact, identifying root causes, and implementing effective remediation.

Continuously improve the reliability, observability, scalability, and maintainability of data pipelines and platform infrastructure.

What You'll Bring

5+ years of experience building and operating production-grade data pipelines and data infrastructure.

Strong experience with both batch processing and event-driven architectures.

Hands-on experience with AWS Glue, Athena, and S3-based data lakes, including layered or medallion-style data transformations.

Strong proficiency in Python and SQL for data processing, transformation, and analysis.

Experience solving identity resolution, entity matching, and deduplication problems, including exact and fuzzy matching approaches.

Understanding of durable identifiers and identity-mapping strategies, including the challenges associated with maintaining consistent first-seen or locked mappings over time.

Experience working with event schemas and schema-registry-backed contracts, such as Protobuf, and an understanding of the impact of schema evolution on downstream consumers.

Strong troubleshooting and production incident-response skills, with the ability to assess impact, identify root causes, and drive issues through resolution.

Ability to understand and debug code written in a functional or concurrent programming language, such as Elixir.

Strong communication and collaboration skills, with the ability to work effectively across engineering, product, analytics, and other technical teams.

Nice to Have

Experience with Elixir/Phoenix or other BEAM-based concurrent processing frameworks.

Experience with DynamoDB-backed identity, lookup, or matching services.

Experience with the Snowflake ecosystem, including data modeling, Snowpipe, Streams, and Tasks.

Experience working with sports data providers, licensed data providers, or other complex external data ecosystems.

Experience implementing data observability, including freshness and staleness alerts, null-rate monitoring, schema-drift detection, and data-quality dashboards.

Experience managing infrastructure through Terraform or AWS CDK.

Technologies

Languages: Python, SQL, familiarity with Elixir

AWS: Glue, Athena, S3, DynamoDB

Data & Streaming: Data Lakes, Event-Driven Architecture, Protobuf, Schema Registries

Infrastructure: Terraform, AWS CDK

Data Platforms: Snowflake

Engineering Practices: Data Quality, Data Observability, Identity Resolution, Entity Matching, Incident Response

✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now
Senior Data Engineer at Ryzlabs — Remote