Application Reliability Engineer
About the role
Scope of the Role
We are looking for a hands-on Application Support Engineer to support, maintain, and enhance business-critical enterprise applications built and running on Google App Engine (GAE) and microservices.
The role is focused on application availability, production support, incident response, troubleshooting, and continuous feature enhancement rather than building a new application from the ground up. The ideal candidate can quickly understand an existing microservices-based application landscape, restore service when users are impacted, and deliver incremental enhancements safely across test, pre-production, and production environments.
Experience supporting large-scale, business-critical applications in complex enterprise technology environments is preferred.
What You’ll Own
- Provide production support and maintenance for enterprise applications hosted on Google App Engine, including standard and flexible environments.
- Act as a first point of contact for user-impacting incidents: triage, diagnose, restore service, and drive issues to closure within agreed SLAs/SLOs.
- Troubleshoot application errors, failed requests, latency and performance degradation, service-to-service failures, configuration issues, quota/scaling limits, and dependency or integration failures.
- Design, develop, and deliver feature enhancements and functional improvements to existing applications based on user and business needs.
- Support applications built on microservices architecture, including service boundaries, APIs/contracts, inter-service communication, authentication, and failure/retry behavior.
- Own build, release, and deployment activities across development, test, pre-production, and production environments with appropriate validation, approvals, and rollback plans.
- Manage App Engine deployments including versions, traffic splitting/migration, canary and staged rollouts, rollbacks, service configuration, and scaling settings.
- Perform root-cause analysis for recurring production issues and implement sustainable fixes.
- Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.
- Support platform, framework, library, dependency, and runtime upgrades while maintaining stability, supportability, and compliance.
- Support IAM, service accounts, access controls, secrets management, and operational governance.
- Participate in change management, release-readiness reviews, and on-call/rotational support as required.
- Collaborate with client and cross-functional Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform teams.
- Adapt to established client-specific engineering, security, review, change-management, and operational processes.
- Create and maintain technical documentation, operational runbooks, troubleshooting guides, deployment procedures, and support playbooks.
Microservices Expectations
- Service decomposition and ownership: understand service boundaries, upstream/downstream dependencies, and ownersh