Jobgether
Senior Engineer, Platform Tooling
engineeringfull-timeUS
SALARY
Not listed
WORK TYPE
remote
JOB TYPE
full-time
INDUSTRY
general
✦ AutoApply Let us apply to roles like this on your behalf.
Learn more
About the role
Accountabilities
- Manage and maintain internal and client-facing monitoring and observability platforms, including software support, threshold adjustments, customizations, and special projects.
- Troubleshoot monitoring issues involving SNMP, WMI, SYSLOG, APIs, and other monitoring mechanisms.
- Own assigned incidents and service tickets through resolution, ensuring timely communication and adherence to defined SLAs.
- Diagnose and resolve configuration, customization, and production issues while coordinating vendor support cases through to resolution.
- Support enterprise infrastructure monitoring across networking, VMware, storage, Windows, and Unix/Linux environments.
- Develop, test, and implement new monitoring features and functionality, including validation plans for production deployments.
- Create and maintain clear technical documentation, runbooks, knowledge-base content, service definitions, and operational procedures.
- Develop and maintain best-practice policies for supported monitoring and observability products.
- Use Python, Groovy, Bash, or comparable scripting technologies to automate tasks and enhance platform capabilities.
- Manage personal and team service queues, prioritize competing requests, and allocate time effectively during periods of high operational demand.
- Communicate effectively with clients, peers, engineering teams, management, and vendors regarding incidents, changes, and technical issues.
- Provide emergency on-call support as part of a rotating schedule.
- Remain accessible during assigned shifts through approved communication channels, including instant messaging, phone, email, and other collaboration tools.
- Identify opportunities for process improvements and provide constructive feedback to management regarding operational challenges and areas of concern.
- Support cross-functional engineering teams and develop additional technical expertise across other products and technologies as business needs evolve.
- Handle and escalate high-impact operational issues to management or third-party vendors when necessary.
- Bachelor’s degree or equivalent professional or military experience.
- 3–5 years of experience supporting enterprise IT infrastructure in IT Operations, NOC, Managed Services, Infrastructure Operations, Monitoring, or Observability environments.
- 2+ years of hands-on experience working with LogicMonitor.
- Experience creating custom LogicModules in LogicMonitor.
- Solid understanding of networking fundamentals, including TCP, UDP, IP addressing, routing, switching, VLANs, and firewalls.
- Technical understanding of VMware ESXi from a monitoring and observability perspective.
- Experience monitoring enterprise networking and storage infrastructure.
- Working knowledge of Microsoft Windows and Unix/Linux operating systems.
- Programming or scripting experience with Python, Groovy, Bash, or similar technologies.
- Experience with monitoring and analytics platforms such as Nagios, NetXMS, Datadog, New Relic, AppDynamics, Prometheus, Grafana, LogicMonitor, Cribl, or comparable tools.
- Experience with log handling, parsing, enrichment, and related observability workflows.
- Ability and willingness to continuously develop technical expertise and stay current with emerging technologies and services.
- Experience working effectively in multi-vendor environments and coordinating complex technical issues across internal and external teams.
- Strong troubleshooting, analytical, and problem-solving skills with the ability to manage multiple issues simultaneously.
- Strong interpersonal and relationship-management skills, with the ability to work effectively with technical and management stakeholders at different levels.
- Excellent written and verbal communication skills, including the ability to explain technical concepts clearly to non-technical audiences.
- Strong customer-service orientation and client focus.
- Ability to work independently, take initiative, recognize priorities, and contribute effectively within a collaborative team environment.
- Strong dependability, attention to detail, adaptability, persistence, integrity, and stress tolerance.
- Ability to remain calm, professional, and courteous during high-pressure incidents and operational challenges.
- Cisco or other relevant vendor certifications are preferred.
- Experience with Cribl Stream and/or Cribl Edge is a plus.
- Experience handling and escalating high-impact operational incidents is advantageous.
- Remote position available within the Continental United States.
- Approximately 5% travel associated with the role.
- Opportunity to work with enterprise monitoring, observability, networking, infrastructure, and managed services technologies.
- Exposure to a broad range of platforms and vendors in complex technical environments.
- Opportunities to expand technical expertise across emerging services, automation, and observability technologies.
- Collaborative environment with engineering teams, clients, vendors, and technical leadership.
- Opportunity to contribute to operational improvements, best practices, automation, and knowledge-sharing initiatives.
- Rotating on-call structure providing exposure to complex production environments and incident response.
Requirements
Benefits
✦ Let us apply for you
We find roles like this and apply on your behalf. Cover letter written for each one. Plans from $15/mo. Cancel anytime.
Get AutoApply