Case study / Project Alpha / 2026
PolicyDesk: the rule that applied, and proof of what was said.
A policy knowledge system for a health plan. Answers use the policy version in force on the member's date of service, only from documents the staff member is cleared to see, and Compliance can reconstruct any disputed answer later. Every one of those decisions is made in code, not by the model.
- right for the member's date of service (6 of 13 without it)
- 13 of 13right for the member's date of service (6 of 13 without it)
- unauthorized chunks reached the model
- 0 of 900unauthorized chunks reached the model
- correct-and-cited, 40 answerable questions
- 95%correct-and-cited, 40 answerable questions
- median end-to-end latency
- 2.7 smedian end-to-end latency
The customer, documents and tickets are synthetic. No real PHI. Every figure comes from the eval in the repo.
Who was the customer?
Cascade Valley Health Plan, a fictional regional payer of about 400 staff, scoped through a discovery brief before any code was written. Three user groups with different clearance:
What was broken?
Wrong version quoted
Superseded PDFs sit next to current ones, and wiki pages copy numbers that go stale. The timely filing window moved from 180 to 90 days; the wiki still said 180.
Wrong audience
Folders were permissioned by department, not by document classification, so a rep could open a fraud referral procedure. The control was a request not to.
No audit trail
When a member disputed what they were told, the only record was a free-text call note that never said which policy text the answer came from.
Current is not applicable
A claim is judged by the rules in force on the date of service. Timely Filing v2 applies to services on or after 2026-01-01, so a November 2025 visit still has 180 days. My first build always answered from the newest version, and got every older claim wrong.
What did I build?
A FastAPI service with one server-rendered page, on Postgres with pgvector and a single model provider. It ingests three source types: 20 policy PDFs (5 with a superseded version), 7 wiki pages and 10 de-identified call notes, into 119 section-level chunks. Each chunk carries its version, effective date, what supersedes it, its access level and a sensitivity flag.
What I deliberately did not automate:
- Answering members directly. Every answer goes to a trained employee. Liability for a benefit misstatement sits with the plan.
- Adjudicating claims. The tool explains policy. Approving or pricing a claim stays in the claims platform and its controls.
- Ingesting PHI. Tickets are de-identified at export and identifier queries are refused, which keeps the pilot out of PHI scope.
How does it work?
The order is the design. Nothing the caller is not cleared for is ever ranked, reasoned over, or generated from.
- 01Resolve role. Role is re-read from the database on every request and mapped to access levels in code. A role change applies on the next request, not the next login.
- 02Refuse early. Queries with member identifiers or claim actions are refused before retrieval, with no model call.
- 03Filter inside SQL. WHERE access_level = ANY(permitted) runs before any ranking. Restricted rows never leave Postgres, so there is nothing to post-filter.
- 04Retrieve and rank. pgvector cosine plus Postgres full-text rank, fused with RRF, then a deterministic rerank.
- 05Resolve versions. With a date of service, every policy chunk is swapped for the version in force that day, in either direction. Without one, superseded sections are swapped for the current version. Conflicts between a policy and a wiki page are flagged, policy wins.
- 06Generate or abstain. Structured JSON from Gemini. Every citation must be a chunk that was in context, retry once, then fail closed.
- 07Trace. User, role, filter, date of service, versions applied, chunk ids, answer, citations with a hash of the cited text, latency and cost, one row per request.
A test spies on the retriever for every role and asserts that no chunk above the caller's clearance comes back from Postgres or reaches the prompt, and that every executed statement carries the permission predicate.
Answers as of the date of service
I found this gap in my own build after it shipped. PolicyDesk resolved superseded policies to the newest version, which looks correct and cites correctly. But claims are paid under the rules in force on the date of service. A member calling in February about a November visit should hear 180 days; PolicyDesk said 90, cited to the current policy.
The ask page now takes an optional date of service. For every policy in the context, code picks the version whose effective date is the latest on or before that date, in both directions: a 2025 date swaps the current version back to the old one. The version lookup runs through the same permission-filtered SQL as retrieval. If no version had taken effect yet, the request is refused in code instead of being answered from a wiki page. Every answer shows the window each cited version was in force, and an answer given without a date warns the rep when the cited policy changed recently.
| 13 questions, 5 changed policies | Without date | With date |
|---|---|---|
| Right for the member's date of service | 6 of 13 | 13 of 13 |
| Questions where an older version applied | 0 of 6 | 6 of 6 |
| Only the version in force reached the model | 50% | 100% |
| Stated that version's value | 58.3% | 100% |
| Date before any version existed | answered | refused in code |
Vertex AI, gemini-3.5-flash, 2026-09-14. Each question has a date under each version, including both sides of an effective-date boundary, plus one date before any version existed. Correct means only the right version reached the model, the answer cited it, and it stated that version's value, all checked in code.
Every question where an older rule applied came back with today's rule and a citation that looked right. That is the failure a member dispute surfaces months later, which is why the next piece exists.
Reconstructing a disputed answer
A member says they were told the wrong deadline. Compliance needs three facts, and none of them should come from a model: what exactly was said, which policy text it relied on, and whether that text applied on the member's date of service. The disputes page, open to compliance reviewers only and enforced in code, answers all three.
- 01Find what was said. Search answers by words, staff member and date asked. Replays and eval runs are excluded.
- 02See it exactly. The stored answer, who gave it, their role at the time, whether a date of service was entered, and each cited section with its version and in-force window.
- 03Check it against the member's date of service. Code decides, from effective dates, whether each cited version applied on that date, and compares a hash of the cited text taken at answer time with the text indexed now.
- 04Replay. Re-run the question with the original role's clearance for that date, to show what should have been said. The replay is its own trace, linked to the original.
In the walkthrough, a rep asked about timely filing without a date and was told 90 days, cited to version 2. Checked against a member date of service of 2025-11-03, the verdict is that the answer used the wrong policy version: version 1 applied, in force 2024-01-01 to 2025-12-31. The replay, with the rep's clearance for that date, answers 180 days cited to version 1. If the version had been right but the section was edited afterwards, the saved hash would say so instead.
Only the replay calls the model, through the same pipeline and permission filter as any question. Every search, open and replay is logged with the reviewer, and other roles get a 403.
How did I evaluate it?
A golden set of 50 questions, 10 per bucket: easy lookup, cross-document synthesis, conflicting sources, stale policy, and must-refuse. Retrieval and generation are measured separately, and a naive baseline (dense retrieval only, no version resolution, no deterministic refusal) runs against the full system on the same model and prompt. Every golden question is also run as every role to count leaks.
| Metric | Baseline | Full system |
|---|---|---|
| recall@5 | 98.3% | 97.1% |
| Correct-and-cited (40 answerable) | 95.0% | 95.0% |
| Groundedness (LLM-judged) | 63.2% | 100% |
| Refusal correctness (50) | 96% | 100% |
| Superseded section in model context (of 50) | 33 | 0 |
| Conflicts flagged in code (of 10) | 0 | 9 |
| Permission leak sweep, unauthorized chunks | 0 of 900 | 0 of 900 |
| Latency p50 / p95 | 3.0 s / 6.2 s | 2.7 s / 3.8 s |
Vertex AI, gemini-3.5-flash, 2026-09-14. Groundedness is graded by the same model family, so treat it as a smoke test. Every item ran on the configured retrieval path. Re-run on the build that includes date-of-service answering; an earlier run on the previous build gave the full system the same accuracy, groundedness, refusal and leak results.
The headline accuracy number does not move. Both configurations score 95.0% correct-and-cited, because the key-fact check passes an answer that quotes a real number with the wrong effective date. The grader does not: it marked 14 of 38 baseline answers ungrounded, every one over an effective date or a superseded version, and none after version resolution. The deterministic layer earns its place in what reaches the model, not in the accuracy column. (On the previous build, the baseline scored 97.5% and the full system 95.0%, so the accuracy column was never where the difference showed.)
What I got wrong first. My first full run looked fine and was wrong. A per-minute embedding quota had quietly pushed 29 of 50 baseline items onto the keyword-search fallback, so it measured a different system than the one I described. I threw the run out, made the harness retry those items, and added a per-item flag so a clean run can prove it.
Five failures were also triggered live: a source changed after ingest, malformed model output, a provider timeout, a role change mid-session, and an out-of-scope query. Each is documented with what the user saw and the captured log.
What would I change before production?
- Key each rule to the right date. Not every rule runs from the date of service: an appeal window runs from the date on the denial notice. The date that decides which version applies should come from each policy's metadata, and the ask page should ask for that date.
- Make degraded retrieval loud. When the embedding call hits a provider quota, the pipeline falls back to keyword search and still answers. That kept users unblocked, but it also let an eval run silently measure a different system. In production a fallback answer should carry a visible flag, page someone, and count against an error budget.
- Refuse permission-denied questions without the model. Restricted text can never reach the model, but some denied questions still depend on the model abstaining. A per-role topic gate would close that without revealing that a restricted document exists.
- Split multi-part questions before retrieval. One embedding for a two-part question is dominated by one part, which is where synthesis recall is lost.
- A second enforcement layer. Row-level security in Postgres behind the SQL predicate, SSO instead of demo passwords, and a retention policy for traces, which contain questions and answers.
- Scheduled staleness checks. The index only learns a source changed when the check runs. It should run from the document system's change feed, with an owner for every wiki page.