Lodgify
Lodgify

Senior Site Reliability Engineer

engineeringfull-timeSpain
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

⭐ Who we are Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology. Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals.

⭐ How will you make an impact?

  • Define meaningful SLIs, SLOs, and reliability targets for the platform.

  • Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.

  • Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness.

  • Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures.

  • Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana.

  • Implement operational and security best practices through guidelines, policies and automation.

  • Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.

  • Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks.

  • Implement self-service Internal Developer Platform features via APIs and Kubernetes operators.

  • Improve deployment safety, rollbackability, and release observability.

  • Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms.

  • Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements.

  • Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability.

⭐ What makes you a great fit?

  • You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.

  • You understand and apply SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.

  • You can design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals, and are comfortable troubleshooting complex distributed systems to identify systemic reliability improvements.

  • You can write maintainable software to automate operational tasks and reduce manual intervention.

  • You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.

  • You know how to balance reliability, performance, cost, and delivery speed pragmatically.

  • You are comfortable working in a transitional environment where SRE practices are being introduced while critical infrastructure and delivery systems still need hands-on reliability support.

  • You collaborate effectively with Engineering, Platform, Security, and Product stakeholders.

  • You communicate clearly, document well, and enjoy coaching teams toward stronger production ownership.

  • You model initiative and accountability, raising risks early and driving improvements through to completion.

⭐ What does success look like?

  • Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage.

  • Reliability targets are consistently met across critical infrastructure and services.

  • Operational toil and manual intervention are measurably reduced through automation and safer workflows.

  • MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.

  • Post-incident actions are tracked, completed, and used to reduce repeat incidents.

  • Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations.

  • Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience.

✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now