Jobgether
Jobgether

Staff Engineer - Distributed Systems

engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Sick of applying? We apply to roles like this for you, up to 20 a month.
Learn more

About the role

Accountabilities:

    • Own the architecture health of large-scale distributed systems, including failure modes, capacity constraints, consistency guarantees, resilience, and interactions across numerous services and deployments.
    • Review and influence critical-path technical designs, providing rigorous architectural guidance and making well-reasoned recommendations when teams face complex technical trade-offs.
    • Proactively identify systemic risks such as single points of failure, unbounded queues, missing idempotency, thundering-herd effects, capacity constraints, and potential data-loss scenarios, then drive remediation before incidents occur.
    • Build and ship solutions for complex architectural problems, including prototypes, reliability improvements, production remediation, and fixes for difficult cross-team issues.
    • Design resilience into systems through degradation strategies, backpressure mechanisms, isolation boundaries, capacity models, and failure-handling patterns capable of supporting sustained growth.
    • Investigate and resolve the most challenging distributed-system failures by understanding interactions across application services, messaging, databases, infrastructure, and deployment environments.
    • Work hands-on with Node.js/TypeScript and/or Go to prototype architectural solutions and implement critical fixes when required.
    • Establish and improve engineering practices through design reviews, post-mortems, architectural patterns, documentation, and technical guidance that can be adopted across multiple teams.
    • Help define how AI-assisted engineering can be used safely and effectively when developing and operating highly critical distributed systems.
    • Work with technologies including GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch across large-scale production environments.
    • Help engineering teams prepare systems for substantial traffic growth by identifying capacity limits, improving observability, and validating critical failure modes.
    • Influence engineers across teams through technical rigor, collaboration, mentorship, and clear communication rather than relying solely on organizational authority.
    • Requirements:

      • 10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems; experience operating hundreds of services, thousands of instances, or systems processing billions of events is strongly preferred.
      • Demonstrated experience carrying direct accountability for production systems through significant incidents, migrations, reliability challenges, and operational failures.
      • Deep expertise in queueing and asynchronous architectures, including delivery semantics, ordering, backpressure, idempotency, and the practical limitations of exactly-once processing.
      • Strong knowledge of multiple storage technologies across SQL and NoSQL environments, with the ability to reason about consistency models, indexing at scale, performance, and appropriate technology selection.
      • Expert-level experience with Redis or comparable in-memory systems, including behavior under memory pressure, network partitions, failover scenarios, and other failure conditions.
      • Strong production experience operating Kubernetes at scale, including resource limits, autoscaling, capacity planning, node failures, and workload behavior under infrastructure disruption.
      • Fluent in Node.js and/or Go, with sufficient hands-on ability to prototype technical proposals and implement production fixes on critical paths.
      • Exceptional technical communication skills, including the ability to produce design documents, architectural diagrams, and root-cause analyses that drive decisions across multiple engineering teams.
      • Strong systems-thinking ability, with an instinct for evaluating tail behavior, failure modes, network partitions, capacity constraints, and high-load scenarios.
      • Proven ability to influence engineering teams through technical depth, constructive disagreement, clear reasoning, and respect.
      • Strong ownership, curiosity, judgment, and problem-solving skills, particularly when dealing with ambiguous or cross-functional technical challenges.
      • Experience using AI-assisted development tools effectively on complex systems is an advantage, particularly when balancing development speed with code quality, reliability, and operational safety.
      • GCP-native experience with technologies such as Pub/Sub, Cloud Tasks, GKE, and Firestore is a strong advantage.
      • Previous experience in a Staff, Principal, Architect, or comparable systems-focused engineering capacity is beneficial.
      • Experience taking new systems from initial architecture through production while also hardening mature systems is a plus.
      • Benefits:

        • Full-time remote opportunity for engineers based in India.
        • Opportunity to work on distributed systems operating at substantial scale, including billions of automation actions and messages and tens of thousands of requests per second at peak.
        • Broad technical scope spanning application runtimes, messaging, asynchronous processing, databases, caching, Kubernetes, cloud infrastructure, and observability.
        • Opportunity to influence architecture across multiple engineering teams and critical production systems.
        • Significant hands-on ownership, with the ability to prototype, build, ship, and remediate solutions rather than working solely in an advisory architecture capacity.
        • Exposure to large-scale technologies including Node.js, TypeScript, Go, GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch.
        • Opportunity to develop resilience, capacity planning, consistency, and failure-management strategies for systems experiencing continued growth.
        • Collaborative environment where technical rigor, clear communication, proactive problem solving, and engineering mentorship are valued.
        • Opportunity to establish engineering patterns and practices that influence a broad engineering organization.
        • Opportunity to work with AI-assisted engineering tools and help define safe, high-quality practices for their use in critical production systems.
✦ Sick of applying to 40 jobs a month?
I rewrite your resume for ATS by hand first. Once you sign off on it, AutoApply applies to up to 20 roles like this a month, cover letter in your own voice each time. From $14.99/mo, cancel anytime.
Get AutoApply
Apply now