Jobgether
Jobgether

Staff Software Engineer, Alerting Platform

engineeringfull-timeIndia
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Define the architecture for high-fidelity alerting across rule-based thresholds, machine-learning anomaly scores, correlation logic, and customer-defined alert rules.
    • Own the reliability of the alerting pipeline end to end, from alert evaluation through Kafka-based delivery and downstream notification systems.
    • Establish robust guarantees around webhook delivery, idempotency, reliability, and sustained-load performance.
    • Drive correlation strategies that transform security tooling, monitoring telemetry, syslog, OpenTelemetry, and network protocol data into meaningful network and system topologies and dependency maps.
    • Make topology and dependency context usable for AI and LLM-driven reasoning across investigations, troubleshooting, correctness analysis, and remediation workflows.
    • Establish the technical foundations for generating alert definitions from normalized data models using LLM-powered capabilities.
    • Provide architectural leadership across multiple teams and workstreams, stepping into complex initiatives where deep technical expertise is required.
    • Balance strategic investments in detection and correlation capabilities with the long-term scalability, reliability, and maintainability of the alerting platform.
    • Make organization-level architecture and technology decisions and communicate technical direction effectively to engineering and business stakeholders.
    • Mentor engineers across the teams you support, promote strong engineering practices, and raise the technical quality of the overall platform.
    • Contribute to additional strategic initiatives and projects that support the success and evolution of the engineering organization.
    • Requirements:

      • Demonstrated experience designing, building, or operating high-reliability alerting, detection, or notification systems in production at significant scale.
      • Strong experience implementing and operating both rule-based and ML-driven detection or event-processing logic against real-world production traffic.
      • Advanced experience with Apache Kafka, particularly producing and consuming event streams for reliable downstream processing and delivery.
      • Strong understanding of telemetry and protocol data, including security tooling output, monitoring telemetry, syslog, OpenTelemetry, NetFlow or sFlow, SNMP, ICMP, and firewall logs.
      • Ability to transform diverse technical data sources into meaningful network and system topologies, dependency maps, and operational context.
      • Proven experience building systems where webhook reliability, idempotency, delivery guarantees, and performance under sustained load are critical.
      • Expert-level proficiency in Go and/or Python, with strong experience across container orchestration and multi-cloud environments.
      • Demonstrated ability to make organization-level architectural and technology decisions rather than focusing solely on individual implementation tasks.
      • Strong systems-thinking, problem-solving, and technical communication skills, with the ability to influence teams and stakeholders across an organization.
      • Experience mentoring engineers and providing technical direction across multiple teams or workstreams.
      • Experience with Temporal or comparable workflow orchestration platforms is a strong asset.
      • Experience applying LLMs to structured system or network data for troubleshooting, remediation, investigation, or operational workflows is advantageous.
      • Familiarity with stream-processing technologies such as Flink, Protocol Buffers, Argo, or Spacelift is a plus.
      • Knowledge of Memgraph or another graph database for topology data is beneficial.
      • Exposure to osquery, Steampipe, or similar endpoint and cloud inventory technologies is an advantage.
      • Experience operating multi-tenant, cloud-native platforms with large-scale secrets management and TLS is desirable.
      • Benefits:

        • 100% remote work environment.
        • Opportunity to work on a technically challenging alerting and signal-processing platform with significant customer impact.
        • High level of technical ownership and influence over architecture and engineering strategy.
        • Opportunity to work with cloud-native technologies, distributed systems, event streaming, telemetry, and AI/LLM-driven capabilities.
        • Growth opportunities within a high-growth technology environment.
        • Culture that encourages innovation, bold ideas, and technical ownership.
        • Flexible time-off policy.
        • Opportunities to collaborate with distributed engineering teams across multiple countries.
        • Team events and opportunities to build strong professional connections.
        • Employee referral bonus program.
        • Meaningful opportunities to mentor engineers and shape engineering best practices.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now
Staff Software Engineer, Alerting Platform at Jobgether — Remote