# Lead Site Reliability Engineer

- Company: [Kontakt](<https://jobstar.asia/company/kontakt>)
- Location: New York
- Remote: Yes
- Team: Engineering
- Employment type: Full Time
- Posted: July 13, 2026

## Job description

**About** [**Kontakt.io**](http://Kontakt.io)

Inside health systems, where every second can matter, operations are still spread across dozens of disconnected tools and platforms.[Kontakt.io](http://Kontakt.io) is changing that.

We combine proprietary hardware, AI-powered intelligence, and deep integrations with the technology health systems already have in place to build real-time understanding of what's happening across their operations. That intelligence becomes the execution layer care teams have been missing, helping them make smarter decisions and deliver better patient care.

Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, AdventHealth, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we’ve more than doubled our revenue and are rapidly scaling with a clear path toward $100M in annual recurring revenue.

If you're excited to solve hard problems and help health systems deliver better care, we'd love to meet you!

We’re looking for a **Lead Site Reliability Engineer** to own the reliability, performance, and automation of our cloud-based, real-time platform. This role will focus on keeping our platform running smoothly 24/7, minimizing downtime, improving observability, incident response, and self-healing automation. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.

## **What You'll Do**

* **Ensure 99.99% uptime** across our cloud platform, meeting strict SLAs for healthcare customers.
* **Design and implement** self-healing, fault-tolerant systems to prevent failures before they happen.
* **Define SLIs, SLOs, and SLAs,** ensuring proactive performance monitoring and incident resolution.
* **Architect and manage** scalable cloud infrastructure (AWS) for massive real-time data processing.
* **Optimize containerized environments** (Kubernetes, Docker) to support multi-region deployments.
* **Lead the adoption of infrastructure as code** (Terraform) to fully automate infrastructure management.
* **Build and refine** a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
* **Lead incident response and on-call operations,** reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
* **Conduct blameless postmortems** and continuously improve system resilience.
* **Reduce manual intervention** through automated deployment, scaling, and failover mechanisms.
* **Partner with Security & Compliance teams** to ensure infrastructure meets HIPAA and SOC 2 standards.
* **Lead disaster recovery** and business continuity planning to ensure critical healthcare services are always available.
* **Drive technical strategy** and roadmap for scalability, monitoring, and reliability engineering.
* **Collaborate** with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.

## **What You Have**

* **10+ years of experience** in Site Reliability Engineering or Cloud Infrastructure.
* **Proven success** scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
* **Deep expertise** in cloud platforms (AWS), Kubernetes, and distributed systems.
* **Strong background** in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
* **Hands-on experience** with incident management, postmortems, and building resilient systems.
* **Deep knowledge** of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
* **A mature leadership approach,** with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.
* **Strong understanding** of network security, access management, and compliance frameworks (HIPAA, SOC 2).

## **Bonus Points If You Have:**

* Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
* Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.
* Prior experience leading on-call rotations and major incident management processes.

## **Logistics, Perks & Benefits**

* Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office
* Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors
* Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work
* Autonomy to solve meaningful problems with work that ships quickly and makes a difference

**Compensation**

The expected salary range for this role is $200,000 – $250,000 for New York-based candidates. Actual compensation within this range will be determined based on relevant experience, skills, and qualifications. In exceptional cases, where a candidate’s experience or qualifications significantly exceed those anticipated for this role, we may consider the candidate for a more senior level. This role may also be eligible for equity and bonus compensation.

## Apply

[Apply on Kontakt](<https://jobs.ashbyhq.com/kontakt/88ad3a25-4395-4b73-aa4a-a2821cc4773d>)

Canonical job page: <https://jobstar.asia/job/lead-site-reliability-engineer-kontakt-new-york-9932c7468ba58d33>
