We’re Heidi.
We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. Our platform has absorbed the administrative chaos of 175 million patient visits and supported 67 million clinical hours. Today, we support 2.8 million patient sessions a week across 190+ countries, in 110 languages and over 200 specialties.
Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s capacity for care and strengthen the human connection at its heart.
We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.
We’ve grown annual recurring revenue from $1 million to $50 million in two years.
To go further, we’ve secured US$340 million: a $100 million Series C led by Blackbird, with Phoenix Court, Point72 Private Investments and Headline, alongside a $240 million growth investment led by General Catalyst’s Customer Value Fund.
If you want to join us in doing work worth shipping, jump in.
The role
This role sits in the core Platform/SRE team that owns production. You’ll work directly on incident response, on-call, system reliability, and day-to-day operations for Heidi’s platform.
We’re open to candidates who are strong mid-level SREs ready to take on more ownership, as well as senior SREs who enjoy being hands-on in operations. The role is intentionally ops-heavy and focused on keeping real systems healthy in production.
What you’ll do
- Participate in on-call and incident response: Respond to production incidents, contribute to service restoration, and support clear communication during incidents. Over time, take increasing responsibility for leading incidents end-to-end.
- Improve operational reliability: Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements.
- Own parts of the production environment: Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing ownership as familiarity increases.
- Strengthen observability: Improve dashboards, alerts, logs, and traces so issues are detected earlier and diagnosed faster, with a strong focus on actionable signals.
- Reduce operational toil: Automate repetitive tasks, simplify runbooks, and improve tooling to make on-call and day-to-day operations easier and safer.
- Support safe change: Improve deployments, rollback mechanisms, and operational readiness to reduce the risk of incidents caused by change.
- Contribute to operational practices: Write and maintain runbooks, participate in blameless post-mortems, and help improve incident response processes over time.
- Collaborate closely with engineers: Work with product and feature teams to improve production readiness, service ownership, and reliability expectations.
What you'll need
- 3–6+ years in SRE, DevOps, Platform, or operations-heavy engineering roles.
- Experience supporting production systems and participating in on-call rotations.
- Comfortable debugging live systems under pressure.
- Experience operating cloud infrastructure (AWS preferred).
- Working knowledge of Kubernetes and containerised workloads.
- Infrastructure as Code experience (Terraform or similar).
- Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc).
- Scripting or automation experience (Python, Bash, or similar).
- Passion for AI, shown through hands-on building, prototyping, or side projects.
- You default to building over requesting.
- You're confident in your thinking and open to being wrong. Great ideas win regardless of who surfaces them.
How we show up
- Build for the