ML / AI Platform Engineer (SR I / II)
Shipt / Platform
- Build low-latency Rust services for agent memory and feature serving, exposing the same unary API over gRPC and HTTP from one protobuf contract, backed by Postgres, Bigtable, Kafka, and Snowflake.
- Run shared platform infrastructure as horizontally scalable services: the model registry autoscales 2 to 10 pods on Kubernetes with health probes, artifact proxying, and zero-credential client access.
- Own the observability and reliability layer: Langfuse tracing across web and worker roles on a 3-node ClickHouse backend, Prometheus and Datadog APM export, and Promptfoo eval gates in CI.
- Operate the LLM gateway, turning a LiteLLM proxy into a governed gateway with routing, budgets, audit logs, guardrails, and multi-modal support.
- Build agent orchestration foundations on Google ADK, LangGraph, A2A, and MCP: durable sessions and tasks plus reference templates other teams build on.
- Make the operational calls that keep platform services boring: explicit schema-migration guardrails, autoscaling policies, and catching a runaway GenAI job queue before it shipped.