Self-RAG
Self-RAG writes an answer from documents it has retrieved, and checks the draft before returning it. I built the control flow as a LangGraph state machine rather than a chain, because a draft that fails its checks has to be able to go back round.
the loop
The graph has four nodes. Retrieve pulls candidates from a Chroma vector store, and grade_documents scores each one against the question. Anything the grader turns down is dropped there. If what is left is too thin, the routing goes through web_search, a Tavily lookup, before it reaches generate. Otherwise it goes straight there, and the search never runs.
Generate writes the draft, and the graph decides where it goes from there. A chain runs in one direction, so there is nowhere in it to put the edge that hands a failing draft back to generate.
three graders
Three graders decide all of this, each one a small chain of its own with structured output.
The retrieval grader decides whether a document is relevant to the question. Hallucination grading happens after the draft exists, and compares it against the documents in hand to see whether they support what it says. A draft can be fully supported by its sources and still miss the question, so the third grader asks only whether it addresses what was asked.
A draft that fails the hallucination grader goes back through generate. One that keeps failing is not returned.
why it is tested this hard
I wrote a test for every chain and every node, and the graph routing has a set of tests of its own, separate from those.
That is more testing than a project this size needs. The graders are the only reason to trust what comes out, and one that quietly stops working leaves something that still looks fine from outside. The routing tests cover the case where every node works and the graph still sends a draft the wrong way.
- Python
- LangGraph
- LangChain
- Chroma
- Tavily
- OpenAI