ReasonTrace
Agentic RAG · LangGraph

RAG that knows when to stop guessing.

An evidence-grounded agent that decides when to retrieve, when to retrieve again, when to call a tool, when to verify — and when to refuse to answer.

Runs with no API key — a deterministic offline engine executes the real graph, retriever and tools.
The live demo is served from a development tunnel and may be offline; the code and the reasoning-path report are always available, and the repo runs locally in two commands.

229Tests passing
6/6Eval cases
12/12Path metrics
100%Route accuracy
0Wasted tool calls

Why not just plain RAG?

Standard RAG is a straight line: question → vector search → LLM → answer. It has three failure modes that matter once real people depend on the output.

FailureWhat the user seesWhy one shot can’t fix it
Facts live in two documents A confident answer built on half the evidence One query, one shot — the second fact is never searched for
The answer needs arithmetic A number that looks right and is wrong LLMs approximate arithmetic, and nothing checks the result
The corpus lacks the answer A fluent hallucination Nothing ever asks “is this actually enough evidence?”

How it works

Nine LangGraph nodes and six conditional-edge functions — all ordinary, readable Python. None of the control flow hides inside a prompt.

Analyze→ Retrieve→ Evaluate evidence→ Retrieve again→ Tool→ Verify→ Answer or Insufficient evidence
01

Decides what it needs first

It enumerates the atomic facts required before retrieving anything. That list is what every later routing decision is measured against.

02

Retrieves again, on purpose

After each search it computes which facts are still missing and writes a query targeting them — one it could not have written earlier.

03

Calculates, never guesses

Arithmetic goes to a deterministic calculator. Every operand must already appear in retrieved evidence or the call is refused.

04

Verifies before answering

Citations must resolve to real chunks; every number must trace to evidence or tool output. A fabricated figure fails here, not in front of the user.

05

Refuses when it should

If the documents don’t support an answer it says so and names the missing fact, instead of filling the gap from model knowledge.

06

Shows its whole path

Every run exposes a structured trace of observable actions and routing decisions — never model chain-of-thought.

Three decisions worth defending

Sufficiency is computed, not asked for

The LLM only answers a narrow, checkable question — “does this passage literally state fact X?” The decision itself is arithmetic:

sufficient = all required facts grounded
   AND no unresolved contradiction
   AND (if a tool is needed) operands available

Stable across providers. A model-emitted confidence: 0.72 is not.

The verifier runs before the answer node

It composes the draft, audits it against five deterministic checks, and only then does the answer node publish. Nothing unverified can reach the user. A deliberately lying provider — returning “$9,999,999 [Totally Real Source]” — is caught and the run ends in abstention. That case is in the test suite.

Contradictions are adjudicated, not averaged

Document front-matter (effective_date, status, supersedes) decides authority. When metadata can’t settle it, the agent abstains and reports the conflict rather than quietly picking whichever chunk ranked higher.

Loops are prevented, and the refusal is visible

Iteration, retrieval and tool budgets; duplicate-query detection over stemmed tokens; tool-argument fingerprinting; a no-progress counter. Every refusal is written into the trace as a GUARD step — you watch the agent choose not to loop.

The four demo scenarios

Six documents for a fictional company, Northstar Labs — deliberately incomplete. Nothing anywhere in the corpus names a founder.

Multi-hopWhat annual professional development allowance is Maya eligible for based on her joining date?
retrieve   "annual professional development allowance"  → benefits_policy.txt
evaluate   found 1/2 · MISSING "Maya joining date"       → retrieve again
retrieve   "Maya joining date"                           → employee_handbook.txt
evaluate   found 2/2 · coverage 100%                     → verify → answer
Verified  The second query could not have been written before the first hop returned. That is the hop.
Tool requiredNorthstar generated $420,000 in Q1 revenue. If the growth target is 18%, what revenue would meet it?
retrieve ×2  "q1 revenue" → $420,000 · "growth target" → 18%
tool_router  operands 420000 and 18 confirmed in evidence → allow
call_tool    calculator: 420,000 × (1 + 18/100) = 495,600
verify       answer reports the calculator's number, not its own arithmetic
Verified  The LLM never does the arithmetic, and both operands are checked against the corpus first.
UnanswerableWho founded Northstar Labs?
retrieve ×2  "Northstar Labs founded" then a broadened reformulation
evaluate     found 0/1 · no new facts twice → no-progress limit reached
abstain
Insufficient information  “I can’t answer this from the provided documents.” The overview says the company “has been operating since 2019” — a relevant document that doesn’t contain the required fact. Telling those apart is the whole point.
ContradictionWhat is the annual remote work allowance?
benefits_policy.txt          (current,  eff. 2025-01-01) → $750
benefits_policy_2023.txt     (archived, superseded)      → $500
evaluate  conflict → resolved_by_supersession → $750, disclosed in the answer
Verified  Silently dropping the $500 would fail the evaluation just as picking it would.

Evaluated on the path, not the string

An agent can reach the right sentence by luck. Each case asserts the route, the documents cited, the tools used (and not used), whether it abstained, and the budgets it respected.

Test              Scenario                            Result
----------------------------------------------------------------
multi_hop_001     Multi-hop across two documents      PASS
tool_calc_001     Retrieved figures + calculator      PASS
abstain_001       Unanswerable question               PASS
conflict_001      Contradictory sources               PASS
factual_001       Single-hop fact, no extra hops      PASS
loop_guard_001    Loop prevention                     PASS
----------------------------------------------------------------
Cases passed  6/6      Route accuracy        100%
Citation grounding 100%      Correct abstention    100%
Unnecessary tool calls  0    Duplicate retrievals    0

Stack & engineering

Python 3.11+LangGraph Pydantic v2FastAPI StreamlitNumPy ChromaDBsentence-transformers pytestDocker

Providers swap by environment variable — OpenAI, Anthropic, or a deterministic offline engine. So do embeddings (hashing / openai / sentence_transformers) and the vector store (numpy / chroma). Each swap is covered by tests.

  • Hybrid retrieval — dense vectors + BM25 → reciprocal rank fusion → reranking.
  • Section-aware chunking with page numbers and front-matter metadata preserved for citations.
  • No eval anywhere — the calculator dispatches named operations to pure functions.
  • Structured logging of every routing decision, keyed by question id. Secrets are never logged.
  • 229 tests, including an adversarial lying-provider case and a Streamlit run that clicks through the demos.