Systems Engineer- AWS Cloud Watch
Location: Atlanta, Ga
Salary Range: $120K-$130K, plus bonus
Benefits:
Medical, Unlimited PTO, 401K match.
Senior Systems Engineer – Observability & Resilience (AWS / CloudWatch Focus)
About the Position
We are seeking a highly skilled Senior Systems Engineer to join our Observability & Resilience team. This team is in the midst of a major transformation — moving away from a traditional monitoring and alert-management model toward an engineering-led observability and resilience practice focused on intelligent observability, incident correlation, automated response, operational resilience, AI-assisted operations, and engineering enablement.
This is not a monitoring administrator role. This engineer will help move the organization from simply receiving alerts to understanding root cause, improving resilience, reducing operational noise, and helping engineering teams build observability into their systems earlier in the software lifecycle.
The ideal candidate has owned production systems, has experienced real outages, understands what it's like when critical services fail, and can translate those lessons into better observability and operational practices.
Most of our environments are cloud-based, and we currently work with CloudWatch, New Relic, SolarWinds, and Splunk, relying heavily on PagerDuty and ServiceNow ITOM for ingestion. We are not seeking users of these tools — we are seeking thought leaders, owners, developers, and enablers, with an emphasis on back-end capability, adoption, and user experience.
Why This Role Is Different
You're Helping Transform the Team The organization is deliberately moving away from a "ticket-taking" and monitoring administration model. The goal is to build standards, drive adoption, improve resilience, enable engineering teams, and create intelligent automation.
Production Experience Matters We're looking for battle-hardened operators — people who have carried production responsibility, been on incident bridges, supported systems during outages, and worked through high-impact production issues.
You're Not Just Using Tools We're looking for builders, integrators, automators, and platform owners — not simply administrators of monitoring platforms.
AI Is Important — but Not as a Buzzword We're not looking for someone who's built the next AI platform. We want someone who is curious about AI, experimenting with it, passionate about reducing toil through it, and interested in augmenting operational practices.
Spec-Driven Operations For the first time, we can specify how the operations harness is to be used while software is being developed. Help define that standard.
Core Responsibilities
AWS / CloudWatch Observability (Primary Focus)
- Design, build, and improve CloudWatch-based monitoring, metrics, and alerting across cloud-based environments
- Establish and champion cloud observability patterns as the primary standard for engineering teams operating in AWS
- Build synthetic monitoring and alerting that scales across distributed, cloud-native systems
- Drive adoption of AWS monitoring best practices across engineering teams, and help define spec-driven observability standards for cloud workloads
- Move the organization from alert management to incident understanding
- Help engineering teams understand root causes, failure chains, service dependencies, and operational impact across cloud infrastructure
- Design and improve metrics, logs, traces, alerting, and synthetic monitoring patterns
- Build and champion effective observability adoption patterns across engineering teams
- Build scripts, integrations, automation tooling, and AI-enhanced operational solutions
- Own tooling development within our team's GitHub repositories
- Automate alert responses and document monitoring solutions
- Partner directly with release trains, engineering managers, technical leads, and product teams
- Develop and own relationships with engineering teams to drive observability adoption and governance
- Develop a technical specialty area and help train and mentor other engineers
- Spread operational knowledge across the team and broader organization
- Bachelor's degree in a related discipline and 4 years' experience in a related field (or equivalent: master's + 2 years, PhD + up to 1 year, or 16 years' experience in lieu of a degree)
- Deep, hands-on experience with AWS CloudWatch and cloud-native monitoring/observability patterns
- Professional experience optimizing the integration and flow of monitoring and ITIL systems
- Hands-on experience with enterprise tooling such as New Relic, SolarWinds, Splunk, PagerDuty, and ServiceNow ITOM as a builder and integrator, alongside CloudWatch
- Professional experience writing synthetic tests in Python, Ruby, or JavaScript using Playwright, Puppeteer, or Selenium
- Distributed systems expertise and understanding of failure modes
- Deep observability experience — instrumentation, metrics, logs, traces, and alerting at scale
- Experience building internal platforms, developer tools, or automation that scales
- Git/version control and CI/CD pipeline experience
- Infrastructure as code and API design experience
- Track record eliminating toil through intelligent automation
- Production ownership experience (on-call, incident response, observability)
- Systems thinking mindset — understanding how components interact at scale
- Eager to dig into problems and bring proposed solutions to group discussion
- Open to feedback and able to creatively adapt multiple ideas into solutions
- Strong technical writing skills, including high- and low-level diagramming techniques
- Analytical skills and careful attention to detail
- Availability for rotational on-call duties outside standard business hours may be required
Why This Role Is Different (Leadership & Growth)
You'll be a key player transforming a team, developing key relationships with engineering teams and driving a roadmap to enable and govern solid observability and resilience patterns. You'll work with cutting-edge LLM technology to solve real production and observability problems, help define spec-driven operations standards, and grow into technical acumen while gaining exposure to leadership across all levels.
Brilliant Staffing, LLC is an Equal Opportunity Employer and encourages applications from all individuals regardless of race, color, religion, gender, gender identity, sexual orientation, national origin, disability, or veteran status.