Senior LLMOps Engineer
We’re Heidi. We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries. Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart. We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again. If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi. The role This role sits in the model team, the researchers and engineers who train, deploy, and own the AI models behind every Heidi product. Those models run at serious scale; what they need now is a serious operational layer around them. You'll build it: full visibility into every deployment, a feedback flywheel that turns real clinician signals into better models, and per-model unit economics that tell us which models pay their way. We're looking for someone who has already built this at an AI company operating at or beyond our maturity, and can bring that playbook to Heidi. It's a hands-on senior engineering role: you'll design the data models, build the pipelines, dashboards, and agents, and own them in production. What you’ll do - Build the deployment health dashboard. Give the team live visibility into every model in production: health metrics, monitoring, and proactive incident alerting that catches problems before clinicians feel them. - Make every incident traceable. Build complete session-to-model lineage, so a degradation can be walked from affected sessions to affected user profiles to the exact model ID, scope of impact, and root cause in minutes. - Stand up the improvement flywheel. Start from Intercom tickets and qualitative CSAT feedback, and put AI agents to work identifying the exact session behind each piece of feedback. - Surface the full story. Retrieve the complete execution trace for every flagged session, bring it to the model team for review, and generate summaries that stakeholders outside engineering can act on. - Close the loop. Filter high-value feedback into training data, ship improved models, and monitor post-deployment performance so every release improves on the last. - Own per-model P&L. Measure revenue against inference cost for every model deployment, and turn model selection and deployment strategy into decisions backed by unit economics. - Raise our LLMOps bar. Bring the practices proven at the most mature AI companies (tracing, evaluation, model incident response) and make them how Heidi operates. - Partner across the model team. Work with the researchers and engineers behind our ASR, note generation, Evidence, and Dictate models so observability is built in, not bolted on. What you'll need - You've spent the last 2–3 years hands-on in an LLMOps role, building the observability, tracing, evaluation, and feedback systems around production LLMs and owning them through real incidents. - That experience comes from an AI company operating at or ahead of Heidi's maturity, most likely in the US or China, where LLMOps practice runs deepest. You know what great looks like, and you can build it here. - Proven ability to ship the systems this role owns: monitoring and alerting (Datadog or similar), distributed tracing across multi-step LLM pipelines, and session and event data models that hold up at scale. - Experience building with LLMs, not just operating them. You can put an
Sign in to apply — one profile, every role on PreferHired.
Sign in to apply