At Dragos, the mission is personal. The systems we protect deliver the water you drink, power your home, and keep the hospitals your community depends on running. Those critical infrastructure systems that power our civilization around the world are under attack every day by adversaries. When those systems fail, people are immediately at risk. We are the global leader in xOT cybersecurity, combining technology, threat intelligence, and expert services. The people here chose this work because they understand what is at stake. Here, you will find a remote-first mission-driven team across North America, Europe, the Middle East, and APAC built on authenticity, transparency, and trust. If safeguarding the systems that protect your family, friends, and community is the kind of work that matters to you, you are in the right place.
About the Role:
Dragos Engineering Support is the bridge between our customers and the engineering teams that build the Dragos platform. When something breaks in a customer environment — or before it does — Engineering Support is the team that investigates, diagnoses, and drives resolution. We own Tier 3 technical escalations, application-layer observability and reliability, and the feedback loop that turns production signals into better software.
As a Senior Engineering Support Engineer, you'll serve as a primary escalation point for complex customer issues on a team that spans deep SRE and platform expertise. You'll own deep-dive investigations into the Dragos application stack, collaborate with engineering and product teams to drive resolution, and help build the monitoring and observability practices that keep our platform healthy. You'll also mentor Tier 1/2 Customer Experience support staff, maintain the knowledge base that keeps the team effective, and participate in the on-call rotation that ensures we're always watching the signals that matter.
Responsibilities:
- Lead escalation triage and incident response for complex, multi-component customer support cases — owning investigation and resolution coordination from intake to close, running technical incident command, and coordinating with infrastructure and product teams as diagnosis requires.
- Validate, reproduce, and document confirmed defects emerging from customer escalations — producing clear bug reports with supporting evidence, reproduction steps, and expected behavior.
- Build and maintain application observability — develop and refine Datadog monitors, dashboards, and alerts that provide visibility into customer-facing SLOs, platform health, and the early warning signals that enable proactive response.
- Own Customer Experience communication through the lifecycle of escalated incidents, ensuring accurate, timely updates and appropriate customer visibility into resolution progress.
- Translate field experience into documentation — author and maintain troubleshooting guides, runbooks, and playbooks that Tier 1/2 CX support staff can act on independently.
- Drive pattern recognition across the customer base — identify recurring failure modes, surface trends to product teams, and actively collaborate on permanent fixes rather than workarounds.
- Participate in on-call rotation, including occasional weekend coverage, triaging application-layer alerts within published SLOs and executing documented remediations.
Qualifications:
- 3+ years of experience in technical support engineering, site reliability engineering, or a closely related customer-facing engineering role, with demonstrated ownership of complex issue resolution.
- Demonstrated experience troubleshooting complex distributed systems in production environments — comfortable reading logs, tracing component interactions, and isolating root cause under pressure.
- Strong Linux system administration skills — comfortable with processes, filesystem, networking, and diagnosing application behavior from first principles.
- Experience with containerized application environments (Kubernetes and/or Docker) at a level sufficient to investigate pod health, examine logs, and understand deployment state.
- Familiarity with observability platforms (Datadog or equivalent) and some experience building and maintaining monitors, dashboards, and alerts.
- Experience supporting Elasticsearch, PostgreSQL, or similar database platforms in a production environment.
- Strong written communication skills with a demonstrated ability to produce clear, accurate technical documentation — RCAs, runbooks, postmortems — for both internal and customer-facing audiences.
- Comfort working with AI tools and assistants as part of day-to-day engineering workflow, including prompt engineering and AI-assisted triage or investigation.
- Experience operating in a customer-facing or customer-adjacent engineering role, with the judgment to balance urgency and thoroughness on competing priorities under SLA pressure.
- Experience supporting or securing software in a cybersecurity context — working with security products, operating in security-conscious environments, or supporting customers with security-driven requirements.
- Ability to read and navigate application source code well enough to trace a bug, understand a service's behavior, or follow a data flow.
- Direct experience with OT/ICS cybersecurity environments or familiarity with industrial control systems — the product domain is learnable, but adjacent experience accelerates ramp significantly, preferred.
- Experience with AWS (EKS, RDS, EC2) or comparable cloud infrastructure at the application layer — understanding how infrastructure behaviors manifest as application symptoms, preferred.
- Background in customer success, customer engineering, or professional services roles that required deep technical credibility alongside relationship management, preferred.
Compensation:
- Competitive Equity Package
- Comprehensive Benefits Plan
#LI-MM1 #LI-REMOTE
Dragos is an Equal Opportunity Employer and considers applicants for employment without regard to race, color, religion, sex, orientation, national origin, age, disability, genetics, or any other basis forbidden under federal, state, or local laws. All new hires must pass a background check as a condition of employment.