Jobgether
Jobgether

Senior Engineer, Platform Tooling

engineeringfull-timeUS
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
Apply for this position
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more

About the role

Accountabilities

    • Manage and maintain internal and client-facing monitoring and observability platforms, including software support, threshold adjustments, customizations, and special projects.
    • Troubleshoot monitoring issues involving SNMP, WMI, SYSLOG, APIs, and other monitoring mechanisms.
    • Own assigned incidents and service tickets through resolution, ensuring timely communication and adherence to defined SLAs.
    • Diagnose and resolve configuration, customization, and production issues while coordinating vendor support cases through to resolution.
    • Support enterprise infrastructure monitoring across networking, VMware, storage, Windows, and Unix/Linux environments.
    • Develop, test, and implement new monitoring features and functionality, including validation plans for production deployments.
    • Create and maintain clear technical documentation, runbooks, knowledge-base content, service definitions, and operational procedures.
    • Develop and maintain best-practice policies for supported monitoring and observability products.
    • Use Python, Groovy, Bash, or comparable scripting technologies to automate tasks and enhance platform capabilities.
    • Manage personal and team service queues, prioritize competing requests, and allocate time effectively during periods of high operational demand.
    • Communicate effectively with clients, peers, engineering teams, management, and vendors regarding incidents, changes, and technical issues.
    • Provide emergency on-call support as part of a rotating schedule.
    • Remain accessible during assigned shifts through approved communication channels, including instant messaging, phone, email, and other collaboration tools.
    • Identify opportunities for process improvements and provide constructive feedback to management regarding operational challenges and areas of concern.
    • Support cross-functional engineering teams and develop additional technical expertise across other products and technologies as business needs evolve.
    • Handle and escalate high-impact operational issues to management or third-party vendors when necessary.
    • Requirements

      • Bachelor’s degree or equivalent professional or military experience.
      • 3–5 years of experience supporting enterprise IT infrastructure in IT Operations, NOC, Managed Services, Infrastructure Operations, Monitoring, or Observability environments.
      • 2+ years of hands-on experience working with LogicMonitor.
      • Experience creating custom LogicModules in LogicMonitor.
      • Solid understanding of networking fundamentals, including TCP, UDP, IP addressing, routing, switching, VLANs, and firewalls.
      • Technical understanding of VMware ESXi from a monitoring and observability perspective.
      • Experience monitoring enterprise networking and storage infrastructure.
      • Working knowledge of Microsoft Windows and Unix/Linux operating systems.
      • Programming or scripting experience with Python, Groovy, Bash, or similar technologies.
      • Experience with monitoring and analytics platforms such as Nagios, NetXMS, Datadog, New Relic, AppDynamics, Prometheus, Grafana, LogicMonitor, Cribl, or comparable tools.
      • Experience with log handling, parsing, enrichment, and related observability workflows.
      • Ability and willingness to continuously develop technical expertise and stay current with emerging technologies and services.
      • Experience working effectively in multi-vendor environments and coordinating complex technical issues across internal and external teams.
      • Strong troubleshooting, analytical, and problem-solving skills with the ability to manage multiple issues simultaneously.
      • Strong interpersonal and relationship-management skills, with the ability to work effectively with technical and management stakeholders at different levels.
      • Excellent written and verbal communication skills, including the ability to explain technical concepts clearly to non-technical audiences.
      • Strong customer-service orientation and client focus.
      • Ability to work independently, take initiative, recognize priorities, and contribute effectively within a collaborative team environment.
      • Strong dependability, attention to detail, adaptability, persistence, integrity, and stress tolerance.
      • Ability to remain calm, professional, and courteous during high-pressure incidents and operational challenges.
      • Cisco or other relevant vendor certifications are preferred.
      • Experience with Cribl Stream and/or Cribl Edge is a plus.
      • Experience handling and escalating high-impact operational incidents is advantageous.
      • Benefits

        • Remote position available within the Continental United States.
        • Approximately 5% travel associated with the role.
        • Opportunity to work with enterprise monitoring, observability, networking, infrastructure, and managed services technologies.
        • Exposure to a broad range of platforms and vendors in complex technical environments.
        • Opportunities to expand technical expertise across emerging services, automation, and observability technologies.
        • Collaborative environment with engineering teams, clients, vendors, and technical leadership.
        • Opportunity to contribute to operational improvements, best practices, automation, and knowledge-sharing initiatives.
        • Rotating on-call structure providing exposure to complex production environments and incident response.
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply
Apply now