Project
LiveWaypoint
A service for running LLM agents as durable jobs: runs stream to any client, pause for human approval, survive restarts and can be rewound. A research agent with web search, a critic and a citation verifier, measured by an eval harness with a calibrated judge.
- FastAPI
- LangGraph
- Postgres
- Redis
- React
- Evals
Waypoint runs LLM agents as durable jobs. A client starts a run and gets an id back immediately; the agent's work streams to any number of watchers; a run can stop and wait for a person for days, and any past step can be rewound and run down a new branch. It is a FastAPI and LangGraph service with a React client, on Postgres and Redis, deployed behind Caddy on a single host.
It exists to show the parts of agent systems that demos skip: what happens when the connection drops, who owns a paused run, how one tenant's cached answers stay out of another's, and how anyone knows whether a prompt change made the agent better.
Two agents
Research plans a topic into up to five questions and stops for a person to approve or edit the plan. Approved questions fan out to parallel workers that search the web; a critic scores each finding and may send follow-up questions back for a second wave. A writer produces a briefing with numbered citations, and a verifier checks every sentence's citation against the text the search API says the answer quoted.
Chat is a tool-using assistant with long-term memory scoped to tenant and user, so what it learns in one conversation carries into the next.
Decisions that shaped it
A run is not a request. POST starts a task and returns 202. Every event goes to a Redis Stream, and clients read it over server-sent events from any worker. Because stream ids double as SSE event ids, a client that drops reconnects with Last-Event-ID and loses nothing, and sees nothing twice.
Interrupts end the run. The approval pause is a checkpoint in Postgres, not a task waiting in memory. The approval card survives a reload, a deploy or a week, and holds no resources while it waits.
Identity lives in the run's context, never its state. State is checkpointed forever; if the tenant were part of it, a replay months later would run as a snapshot of who the caller used to be. Ownership is in its own table; another tenant's thread is a 404, not a 403, so its existence is never confirmed.
A model per role, chosen per run. Nodes never name a model. Planner, critic and writer use a stronger model; the searcher, which re-reads every search result as input, runs on a model a twentieth of the price. Moving the searcher to it, and capping searches per worker, cut a full evaluation run from about $18.50 to about $7.60.
Rules in code, not prompts. Sources come only from the provider's citation metadata, never from URLs in model text. A finding with no real citation is capped at low confidence whatever the critic says. The source list is appended from metadata, so the writer cannot invent one.
Measuring it
The repository includes an evaluation harness built the way I would want one in production:
- Deterministic checks first, such as "every
[n]resolves to a listed source" and "at least 80% of claim sentences carry a citation". - An LLM judge that grades one criterion per call against written anchors, reasoning before it scores.
- Calibration against human labels. A criterion is trusted only if the judge agrees with labelled examples at least 80% of the time; otherwise it escalates to a stronger judge or is reported as untrusted.
- A baseline gate in CI that fails when a metric drops by more than twice its own trial-to-trial spread.
- Record and replay of every model call, keyed on exactly what was sent, so re-testing a prompt change pays only for the calls that changed: about $0.90 instead of a full run.
The case study walks through using it on a real problem. A low groundedness score turned out to be mostly the judge; after recalibrating the instrument and adding a citation verifier, briefings with at least 80% of claims cited went from 48% to 95%.
Try it
The demo at waypoint.brentdunham.com runs on real models, so access is by personal link; ask me for one. GitHub has the source, a design guide with diagrams and a decision log, and the deployment runbook.