Jobgether
Jobgether

Manager, Infrastructure

engineeringfull-timeUS
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities:

    • Lead and develop a distributed infrastructure operations team across US and APAC time zones, including regional data center owners and network engineering.
    • Establish and maintain operational rhythms including DevOps check-ins, alert reviews, on-call coverage, incident reviews, and performance tracking.
    • Own P1/P2 incident response from detection through resolution, with accountability for reducing MTTR, improving runbooks, and maintaining high-quality alerting.
    • Manage capacity planning and the full hardware lifecycle across four data centers, including Dell and Supermicro procurement, GPU expansion, colocation power and space, and remote-hands logistics.
    • Lead annual cloud-versus-colocation evaluations and provide recommendations based on performance, capacity, cost, and operational requirements.
    • Oversee the Linux and infrastructure platform, including Ubuntu/systemd fleets, FreeIPA, Ansible/Salt, MAAS provisioning, and Zabbix monitoring.
    • Manage high-performance spine-leaf networking, BGP and peering, transit providers, and low-latency network optimization.
    • Oversee stateful distributed data platforms such as Aerospike, Kafka, ClickHouse, and Hadoop/HDFS, including capacity management, migrations, evictions, and performance tuning.
    • Own colocation, networking, licensing, procurement, and infrastructure vendor relationships and associated budgets.
    • Partner with security and compliance teams on infrastructure hardening, access reviews, SOC 2 Type 2 evidence, and related controls.
    • Operate effectively across a multi-entity environment and help navigate infrastructure requirements associated with organizational growth and acquisitions.
    • Support and develop infrastructure engineers while maintaining a culture of ownership, proactive communication, technical excellence, and continuous improvement.
    • Requirements:

      • 6–8 years of experience in infrastructure or data center operations, including at least 2 years managing engineers or technical teams.
      • Strong bare-metal and colocation experience, including capacity planning, hardware procurement, owned-cage operations, remote-hands coordination, and physical data center logistics.
      • Solid networking expertise at scale, including spine-leaf architectures, BGP, high-capacity peering, transit providers, and low-latency network optimization.
      • Deep Linux operations experience, including systemd, netplan, FreeIPA, Chrony, Ansible and/or Salt, Zabbix, and management of fleets containing hundreds of servers.
      • Experience operating large-scale stateful distributed systems such as Aerospike, Cassandra, Scylla, Kafka, or ClickHouse within demanding latency and capacity constraints.
      • Demonstrated ownership of P1/P2 incidents, on-call programs, postmortems, operational runbooks, and alert-management practices.
      • Experience leading and developing experienced technical teams across multiple geographic regions and time zones.
      • Strong understanding of hardware lifecycle management, infrastructure procurement, vendor relationships, and operational budgeting.
      • Excellent analytical, troubleshooting, and decision-making abilities, with a practical approach to complex infrastructure problems.
      • Strong written and verbal communication skills and the ability to work effectively with engineering, security, leadership, and external vendors.
      • Ability to operate independently in a fast-moving environment with changing priorities and organizational ambiguity.
      • Experience with AWS, including IAM, Route 53, GuardDuty, and S3, is a plus.
      • Kubernetes exposure and experience with hybrid cloud environments are advantageous.
      • Familiarity with SOC 2, Okta, Vanta, or security-focused infrastructure practices is preferred.
      • Experience in adtech, real-time bidding, or other high-QPS, latency-sensitive environments is a plus.
      • Netris, SDN controller, and MAAS provisioning experience is desirable.
      • Willingness and ability to travel periodically to domestic and international data center locations and headquarters.
      • Must be authorized to work in the United States; the role may involve background screening and does not indicate sponsorship availability.
      • Benefits:

        • Fully remote position for candidates based in the United States.
        • Opportunity to own a complex, globally distributed bare-metal infrastructure environment supporting high-volume, latency-sensitive workloads.
        • Direct ownership of hardware strategy, GPU expansion, capacity planning, and cloud-versus-colocation decisions.
        • Leadership responsibility for an established and experienced global infrastructure team.
        • Direct visibility with senior engineering and infrastructure leadership and strong alignment between infrastructure goals and business objectives.
        • Exposure to cutting-edge infrastructure challenges across high-performance networking, distributed systems, data centers, and on-premise machine learning infrastructure.
        • Periodic travel to data center sites in the United States and internationally, as well as headquarters for operational reviews, team collaboration, and site work.
        • Opportunity to contribute to a rapidly scaling organization and play a significant role in shaping infrastructure strategy and operational excellence.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now