Bring your passion, here's what’s needed:
Required Skills and Qualifications
- Experience: 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, including significant production ownership.
- AWS: Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
- Terraform: Advanced proficiency with Terraform, including reusable modules, remote state, dependency management, environment design, code review, and infrastructure lifecycle management.
- Observability: Experience with metrics, logs, traces, dashboards, alerting, and production telemetry using platforms such as Prometheus, Grafana, CloudWatch, New Relic, ELK/OpenSearch, or similar tools.
- Reliability Engineering: Practical knowledge of SRE concepts such as SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and reducing operational toil.
- Kubernetes: Deep understanding of Kubernetes architecture and operations, including EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, and troubleshooting.
- Linux & Systems: Strong Linux systems knowledge and the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.
- Programming & Automation: Proficiency in Python, Bash, Go, or another general-purpose language used to build operational tooling and automation.
- Incident Management: Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective-action tracking.
- CI/CD: Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable tooling.
- Security: Working knowledge of cloud security best practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.
- Communication: Strong written and verbal communication skills, with the ability to collaborate effectively across engineering and business teams.
Preferred Qualifications
- AWS certifications such as AWS Certified Solutions Architect - Professional or AWS Certified DevOps Engineer - Professional.
- Experience with GitOps practices and tools such as Argo CD or Flux.
- Experience designing or participating in formal on-call rotations and incident management programs.
- Familiarity with chaos engineering, resilience testing, or game-day exercises.
- Experience with service meshes, distributed systems, and microservice architectures.
- Knowledge of database operations for technologies such as Amazon RDS, DynamoDB, and PostgreSQL.
- Experience with multi-account AWS environments, landing zones, governance, or large-scale cloud platform design.
- Experience with infrastructure cost optimization, cloud financial management, or FinOps practices.
- Familiarity with security and compliance frameworks such as SOC 2, ISO 27001, PCI DSS, or similar standards.
What Success Looks Like
- Production services become more measurable, reliable, and resilient over time.
- Operational toil and recurring incidents are reduced through engineering and automation.
- Teams have clear SLOs, useful dashboards, actionable alerts, and well-understood operational ownership.
- Infrastructure and deployment changes are repeatable, observable, secure, and low risk.
- Incidents produce meaningful learning and durable improvements rather than recurring fixes.
- Cloud capacity and cost are proactively managed without compromising reliability.
Why Join Us?
- Work on modern cloud and reliability engineering challenges in a fast-paced, innovative environment.
- Help shape reliability standards and engineering practices for mission-critical systems.
- Collaborate with talented engineers across software, cloud, security, and product disciplines.
- Competitive salary, comprehensive benefits, and opportunities for professional growth.
- Flexible remote or hybrid work options.
Be a part of an innovative team shaping the grid of the future through advanced energy intelligence. For more than half a century, Electric Power Engineers (EPE) has partnered with power and energy clients across the globe, providing consulting expertise and energy intelligence software solutions for complex engineering and grid modeling challenges. As leaders in the renewables space, we are focused on building a modern, secure, and resilient grid. Join us in making an impact on the communities we serve and the environment in which we live. Together we can transform the future of energy.
How we support you:
- Comprehensive health and wellness benefits including medical, dental, and vision with 100% premium coverage for you
- Generous PTO and paid holidays
- MyShare Employee Ownership Program
- Work with industry leaders
- 401K, up to a 4% match (100% vested from day 1)
Location: This position will be located in City, State
Travel: Occasional travel may be needed (10% or less)
EPE is an equal opportunity/AA/Disability/Veteran employer. The EEO is the Law poster, and its supplement are available using the following links: EEOC is the Law Poster
Third-Party Recruiting Notification
EPE does not accept unsolicited resumes from third-party recruiters. Any unsolicited third-party resumes forwarded by recruiters to EPE via our career page or to any of our managers or employees will be considered public information, may be treated as a direct application from the person identified in the resume, and will not be eligible for placement fee payment to the agency. EPE will not pay a fee to a third-party recruiter or agency without a previously signed third-party agreement and has not coordinated their recruiting activity with the appropriate member of the Talent Acquisition team.