An evidence-grounded agent that decides when to retrieve, when to retrieve again, when to call a tool, when to verify — and when to refuse to answer.
Runs with no API key — a deterministic offline engine executes the real graph, retriever and tools.
The live demo is served from a development tunnel and may be offline; the code and the
reasoning-path report are always available, and
the repo runs locally in two commands.
Standard RAG is a straight line: question → vector search → LLM → answer.
It has three failure modes that matter once real people depend on the output.
| Failure | What the user sees | Why one shot can’t fix it |
|---|---|---|
| Facts live in two documents | A confident answer built on half the evidence | One query, one shot — the second fact is never searched for |
| The answer needs arithmetic | A number that looks right and is wrong | LLMs approximate arithmetic, and nothing checks the result |
| The corpus lacks the answer | A fluent hallucination | Nothing ever asks “is this actually enough evidence?” |
Nine LangGraph nodes and six conditional-edge functions — all ordinary, readable Python. None of the control flow hides inside a prompt.
It enumerates the atomic facts required before retrieving anything. That list is what every later routing decision is measured against.
After each search it computes which facts are still missing and writes a query targeting them — one it could not have written earlier.
Arithmetic goes to a deterministic calculator. Every operand must already appear in retrieved evidence or the call is refused.
Citations must resolve to real chunks; every number must trace to evidence or tool output. A fabricated figure fails here, not in front of the user.
If the documents don’t support an answer it says so and names the missing fact, instead of filling the gap from model knowledge.
Every run exposes a structured trace of observable actions and routing decisions — never model chain-of-thought.
The LLM only answers a narrow, checkable question — “does this passage literally state fact X?” The decision itself is arithmetic:
sufficient = all required facts grounded AND no unresolved contradiction AND (if a tool is needed) operands available
Stable across providers. A model-emitted confidence: 0.72 is not.
It composes the draft, audits it against five deterministic checks, and only then does the answer node publish. Nothing unverified can reach the user. A deliberately lying provider — returning “$9,999,999 [Totally Real Source]” — is caught and the run ends in abstention. That case is in the test suite.
Document front-matter (effective_date, status, supersedes) decides authority. When metadata can’t settle it, the agent abstains and reports the conflict rather than quietly picking whichever chunk ranked higher.
Iteration, retrieval and tool budgets; duplicate-query detection over stemmed tokens; tool-argument fingerprinting; a no-progress counter. Every refusal is written into the trace as a GUARD step — you watch the agent choose not to loop.
Six documents for a fictional company, Northstar Labs — deliberately incomplete. Nothing anywhere in the corpus names a founder.
retrieve "annual professional development allowance" → benefits_policy.txt evaluate found 1/2 · MISSING "Maya joining date" → retrieve again retrieve "Maya joining date" → employee_handbook.txt evaluate found 2/2 · coverage 100% → verify → answer
retrieve ×2 "q1 revenue" → $420,000 · "growth target" → 18% tool_router operands 420000 and 18 confirmed in evidence → allow call_tool calculator: 420,000 × (1 + 18/100) = 495,600 verify answer reports the calculator's number, not its own arithmetic
retrieve ×2 "Northstar Labs founded" then a broadened reformulation evaluate found 0/1 · no new facts twice → no-progress limit reached abstain
benefits_policy.txt (current, eff. 2025-01-01) → $750 benefits_policy_2023.txt (archived, superseded) → $500 evaluate conflict → resolved_by_supersession → $750, disclosed in the answer
An agent can reach the right sentence by luck. Each case asserts the route, the documents cited, the tools used (and not used), whether it abstained, and the budgets it respected.
Test Scenario Result ---------------------------------------------------------------- multi_hop_001 Multi-hop across two documents PASS tool_calc_001 Retrieved figures + calculator PASS abstain_001 Unanswerable question PASS conflict_001 Contradictory sources PASS factual_001 Single-hop fact, no extra hops PASS loop_guard_001 Loop prevention PASS ---------------------------------------------------------------- Cases passed 6/6 Route accuracy 100% Citation grounding 100% Correct abstention 100% Unnecessary tool calls 0 Duplicate retrievals 0
Providers swap by environment variable — OpenAI, Anthropic, or a deterministic offline engine.
So do embeddings (hashing / openai / sentence_transformers)
and the vector store (numpy / chroma). Each swap is covered by tests.
eval anywhere — the calculator dispatches named operations to pure functions.