Platform Engineer
Build the engineering system our agents work inside: the runtime, delivery pipeline, and observability that let them write, ship, and maintain production software. Infrastructure engineering, where many of your users are agents.
Docflow builds AI agents that run the back office for healthcare providers. Our automation handles the work that buries their teams: document intake, prior authorizations, order entry, and payer portal workflows. Customers start with a paid pilot, see real work getting done within weeks, and expand from there.
We’re a small, scrappy, AI-native founding team, and we’re growing fast. A year ago, agents opened none of the pull requests merged into our codebase. Last month, they opened more than half. Building the system that makes that safe is now the most important engineering work we do. That’s why we’re hiring you.
The Role
You won’t be writing our automations. You’ll be building the system that writes them, runs them, and keeps them working.
Take browser automation. We run hundreds of automations against payer and provider portals we don’t control, following business rules that change whenever a payer or a customer does. Portals get redesigned overnight. A payer adds a step. A customer changes how they want an order handled. No team of people could keep up with that by hand, so agents do most of it. They investigate failures, trace them to a root cause, write the fix, review each other’s code, confirm the fix worked in production, and finish stuck work themselves in a live browser when an automation can’t.
Your job is to build the engineering system all of that happens inside: the runtime agents work in, the pipeline their changes ship through, the observability that shows them and us what’s actually happening, and the guardrails that keep an agent’s change from reaching a customer unchecked. Browser automation is just one example. The same system supports document extraction, eligibility checks, outbound calling, and whatever we build next.
This is infrastructure engineering: distributed systems, durable execution, deployment and versioning, observability, reliability, and security. What’s different is who uses what you build. Your users are agents and the people who supervise them, and agents are demanding users. They work fast, they work in parallel, and they will find every ambiguity in an interface and every gap in a guardrail.
We’re hiring two engineers. This is the one that builds the platform. Our other engineering role, Product Engineer, owns what customers experience: the tools built on top of the platform and the product they use every day. If you’d rather build the system every workflow and every agent depends on than tune any one of them, this is your seat.
We write strictly typed TypeScript in a functional style: algebraic data types and discriminated unions, state reified as explicit values rather than implied by scattered flags, and exhaustive handling the compiler enforces. That matters more when agents write much of the code, not less. When an agent adds a new case, the compiler walks it to every place that has to handle it.
Because we handle protected health information, reliability and care are not optional. An agent moves faster than a person, and so do its mistakes. A big part of this job is making sure the system catches them first.
What You’ll Do
- Build the runtime our agents work in: isolated sessions on on-demand cloud machines, warm pools, sandboxed browsers, and tightly scoped access to the credentials and portals each job needs
- Evolve the loop that carries a problem from first report to verified fix: investigation, root cause, code change, automated review, human approval, deploy, and verification against real business outcomes in production
- Own deployment and versioning, so a workflow change can ship in under a minute, run exactly the code it was built with, and roll back cleanly without anyone watching
- Build the observability that makes a highly autonomous system legible: metrics, alerting, dashboards as code, execution traces, and provenance that ties every result to the build that produced it
- Treat agents as first-class users of everything you build, with CLIs, APIs, and error messages clear enough that an agent can diagnose a problem without a person translating
- Keep long-running distributed workflows correct under failure: idempotent, resumable, bounded, and safe to replay across deploys
- Build the surfaces where people supervise agents, such as approval queues, code review, and live views of what’s running
- Run production: debug the incident, write the postmortem, and fix the class of problem rather than the one instance
- Work directly with our CTO to set the engineering standards the team, human and agent, grows into
What We’re Looking For
Must have:
- Strong professional experience building and operating production systems. We care more about what you’ve built and kept running than a years-of-experience number, though in practice this usually means several years in the field.
- Real distributed systems instincts. Race conditions, retries, idempotency, backpressure, cancellation, and partial failure are things you think about by default, not after a bug.
- You’ve owned the reliability of something that mattered: been paged for it, debugged it from metrics, logs, and traces, and fixed the root cause instead of the symptom.
- Hands-on experience with cloud infrastructure and deployment, including CI/CD, containers, and shipping and rolling back changes safely on platforms such as AWS, GCP, or Azure.
- Deep proficiency in a statically typed language. TypeScript is what we write every day, but strong Go, Rust, or similar is just as welcome if you’re ready to get strong in TypeScript quickly. Either way, you use the type system to make invalid states unrepresentable rather than reaching for escape hatches like
anyor unsafe casts. - You already build with AI coding agents such as Claude Code, Codex, or Cursor every day, and you have clear opinions about where they break down and what the system around them should do about it.
- Comfort across the stack. Most of this role is backend and infrastructure, but you’ll also build the interfaces people use to supervise agents.
- High standards for readable, maintainable code and a low tolerance for needless complexity.
Preferred:
- A strictly typed, functional style: algebraic data types and discriminated unions, exhaustive matching, state reified as explicit values, and pure logic kept apart from side effects. Time in languages built around these ideas, such as Haskell, OCaml, F#, Rust, or Scala, counts
- Experience with durable execution or workflow orchestration, such as Temporal
- Experience building internal developer platforms, build and release systems, or CI/CD infrastructure that other engineers depended on
- Experience running fleets of containers or VMs, including pooling, autoscaling, and graceful rollouts
- Experience with observability tooling such as Prometheus, Grafana, or OpenTelemetry
- Experience building LLM agents in production, including tool design, evaluation, or guardrails
- Experience in regulated or security-sensitive domains such as healthcare or fintech, and handling sensitive data
- Based in Los Angeles, or ready to be
Compensation and Benefits
- Base salary of $150,000 to $185,000 depending on experience
- Founding-team equity
- Health insurance
Apply
No cover letter. Share your LinkedIn and a short Loom video (under five minutes, one or two is plenty) telling us why you think you're a good fit. Loom is free to sign up for at loom.com.
Thanks — your application is in. We review every one and will be in touch if it's a fit.