Jobgether
Staff Engineer - Distributed Systems
engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more
About the role
Accountabilities:
- Own the architecture health of large-scale distributed systems, including failure modes, capacity constraints, consistency guarantees, resilience, and interactions across numerous services and deployments.
- Review and influence critical-path technical designs, providing rigorous architectural guidance and making well-reasoned recommendations when teams face complex technical trade-offs.
- Proactively identify systemic risks such as single points of failure, unbounded queues, missing idempotency, thundering-herd effects, capacity constraints, and potential data-loss scenarios, then drive remediation before incidents occur.
- Build and ship solutions for complex architectural problems, including prototypes, reliability improvements, production remediation, and fixes for difficult cross-team issues.
- Design resilience into systems through degradation strategies, backpressure mechanisms, isolation boundaries, capacity models, and failure-handling patterns capable of supporting sustained growth.
- Investigate and resolve the most challenging distributed-system failures by understanding interactions across application services, messaging, databases, infrastructure, and deployment environments.
- Work hands-on with Node.js/TypeScript and/or Go to prototype architectural solutions and implement critical fixes when required.
- Establish and improve engineering practices through design reviews, post-mortems, architectural patterns, documentation, and technical guidance that can be adopted across multiple teams.
- Help define how AI-assisted engineering can be used safely and effectively when developing and operating highly critical distributed systems.
- Work with technologies including GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch across large-scale production environments.
- Help engineering teams prepare systems for substantial traffic growth by identifying capacity limits, improving observability, and validating critical failure modes.
- Influence engineers across teams through technical rigor, collaboration, mentorship, and clear communication rather than relying solely on organizational authority.
- 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances, or systems processing billions of events is strongly preferred.
- Demonstrated experience carrying direct accountability for production systems through significant incidents, migrations, reliability challenges, and operational failures.
- Deep expertise in queueing and asynchronous architectures, including delivery semantics, ordering, backpressure, idempotency, and the practical limitations of exactly-once processing.
- Strong knowledge of multiple storage technologies across SQL and NoSQL environments, with the ability to reason about consistency models, indexing at scale, performance, and appropriate technology selection.
- Expert-level experience with Redis or comparable in-memory systems, including behavior under memory pressure, network partitions, failover scenarios, and other failure conditions.
- Strong production experience operating Kubernetes at scale, including resource limits, autoscaling, capacity planning, node failures, and workload behavior under infrastructure disruption.
- Fluent in Node.js and/or Go, with sufficient hands-on ability to prototype technical proposals and implement production fixes on critical paths.
- Exceptional technical communication skills, including the ability to produce design documents, architectural diagrams, and root-cause analyses that drive decisions across multiple engineering teams.
- Strong systems-thinking ability, with an instinct for evaluating tail behavior, failure modes, network partitions, capacity constraints, and high-load scenarios.
- Proven ability to influence engineering teams through technical depth, constructive disagreement, clear reasoning, and respect.
- Strong ownership, curiosity, judgment, and problem-solving skills, particularly when dealing with ambiguous or cross-functional technical challenges.
- Experience using AI-assisted development tools effectively on complex systems is an advantage, particularly when balancing development speed with code quality, reliability, and operational safety.
- GCP-native experience with technologies such as Pub/Sub, Cloud Tasks, GKE, and Firestore is a strong advantage.
- Previous experience in a Staff, Principal, Architect, or comparable systems-focused engineering capacity is beneficial.
- Experience taking new systems from initial architecture through production while also hardening mature systems is a plus.
- Full-time remote opportunity for engineers based in India.
- Opportunity to work on distributed systems operating at substantial scale, including billions of automation actions and messages and tens of thousands of requests per second at peak.
- Broad technical scope spanning application runtimes, messaging, asynchronous processing, databases, caching, Kubernetes, cloud infrastructure, and observability.
- Opportunity to influence architecture across multiple engineering teams and critical production systems.
- Significant hands-on ownership, with the ability to prototype, build, ship, and remediate solutions rather than working solely in an advisory architecture capacity.
- Exposure to large-scale technologies including Node.js, TypeScript, Go, GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch.
- Opportunity to develop resilience, capacity planning, consistency, and failure-management strategies for systems experiencing continued growth.
- Collaborative environment where technical rigor, clear communication, proactive problem solving, and engineering mentorship are valued.
- Opportunity to establish engineering patterns and practices that influence a broad engineering organization.
- Opportunity to work with AI-assisted engineering tools and help define safe, high-quality practices for their use in critical production systems.
Requirements:
Benefits:
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply