Most Loved Workplace® Certified JobSenior Site Reliability Engineer (AWS/EKS and Observability) – EASTERN EUROPE – AWS
Most Loved Workplace® Certified JobAbout the Role
At Distillery Tech, Inc. , a Most Loved Workplace® certified employer in the Information Technology space, Unyielding commitment, relentless pursuit of excellence in nearshore software development.
Rocket Ride is seeking a Senior Site Reliability Engineer to improve the reliability of its production platform. The platform supports an open-source AI runtime and a product website, so website availability is a core product requirement. The engineer will own reliability improvements across infrastructure, observability, incident response, and production change controls, working with application engineers to diagnose failures across the browser, API, runtime, and cloud layers. The environment includes AWS, Amazon EKS, Terraform, GPU infrastructure, and a C++ runtime, plus Python and TypeScript extensions and a JavaScript frontend. Current ops stack uses PagerDuty; a Grafana-based observability approach is under consideration. Good fit: someone who can stabilize an existing system before introducing larger platform changes, and who works well in an early-stage environment with incomplete documentation and changing priorities.
Want to learn more about what it's like to work at Distillery Tech, Inc. ? View our full profile.
Requirements
- 5–7 years of direct SRE work is ideal.
- A mix of SRE, DevOps, and platform engineering experience is also valuable, including 3+ years of direct SRE work.
- Related production, infrastructure, systems, or software engineering experience is valuable when it includes reliability ownership.
- 4+ years operating AWS services in production.
- 3+ years operating Kubernetes in production, including 2+ years with Amazon EKS.
- 3+ years using Terraform, including modules, state, reviews, and safe changes.
- 3+ years working with metrics, logs, traces, dashboards, and actionable alerts.
- 2+ years defining and using service-level indicators and service-level objectives.
- 3+ years responding to production incidents and leading root-cause reviews.
- 5+ years troubleshooting Linux systems and production networks.
- 3+ years using CI/CD controls, automated validation, and safe rollback methods.
- 3+ years writing reliable automation in Python, Go, Bash, or a similar language (strong in one, working knowledge of another).
- Ability to work across infrastructure and application boundaries.
- Clear written and verbal communication in a remote environment.
- A record of ownership from problem detection through verified resolution.
- 2+ years with Grafana or a similar observability platform.
- 1+ year improving on-call operations with PagerDuty or a similar platform, including alert noise and escalation paths.
- 1+ year implementing synthetic monitoring or real-user monitoring.
- 1+ year supporting production C++ services or other high-performance runtimes.
- 2+ years supporting Python applications and TypeScript or JavaScript applications.
- 1+ year supporting GPU infrastructure or AI inference workloads.
- 1+ year supporting OCR, speech, audio, or vision workloads.
- 1+ year supporting identity systems such as Zitadel or another OIDC provider.
- 1+ year supporting SOC 2 or ISO-aligned operational controls, or one complete audit cycle.
- 2+ years in an early-stage company or another fast-changing product environment.
Prefer to apply on the company's career site?
View Original Posting
Why This Is a Most Loved Workplace® Certified Job
What It's Like to Work Here
Frequently Asked Questions About Working at Distillery Tech, Inc.
Common questions candidates ask about this role and Distillery Tech, Inc. 's workplace
Apply for Senior Site Reliability Engineer (AWS/EKS and Observability) – EASTERN EUROPE – AWS
Submit your application directly to the Distillery Tech, Inc. team. Let them know why you'd be a great addition to their loved place to work.

