<p>
<strong>Headquarters:</strong> Remote, United States
</p>
<h1>Director of Production Engineering </h1>
<h2><strong>Remote, United States</strong></h2>
<p><strong>About this Position</strong></p>
<p>Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps & SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance.</p>
<p>This is a hands-on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross-team execution.</p>
<p><strong>Responsibilities</strong></p>
<ul>
<li>Hire and build a globally-distributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.</li>
<li>Own the reliability and infrastructure roadmap for our AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.</li>
<li>Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.</li>
<li>Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.</li>
<li>Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution.</li>
<li>Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and DevOps processes (e.g., automated investigation, remediation, and PR-generation pipelines).</li>
<li>Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.</li>
<li>Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.</li>
<li>Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit-readiness.</li>
<li>Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals.</li>
<li>Stay current on cloud, DevOps, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.</li>
</ul>
<p><strong>Required Qualifications</strong></p>
<ul>
<li>8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience.</li>
<li>Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).</li>
<li>Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues.</li>
<li>5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.</li>
<li>Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.</li>
<li>Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines.</li>
<li>Hands-on experie