Building a gut-health advisor with LightRAG, LangGraph and Claude on Amazon Bedrock
A walkthrough of the retrieval-augmented advisor behind a microbiome diagnostic service: how domain tables become a knowledge graph, how a LangGraph conversation graph handles memory, rewriting, retrieval and streaming, and why a first version with grader nodes on a self-hosted model gave way to a four-call design on Amazon Bedrock — with simplified code for each step.
SmileGut is a microbiome-based diagnostic and health service. Customers who receive a gut-microbiome report want to know what their results mean and what to do next. The knowledge needed to answer is spread across report-interpretation guides, a microbe dictionary, ingredient and nutrient tables, recipes and an FAQ. An FAQ page cannot personalize, and a general chatbot answers from its own memory, which a health product cannot rely on for accuracy, currency or stability. Most real questions also cross sources — a microbe, the nutrients it uses, the foods that contain them — and flat text chunks lose those links.
In this post we walk through the advisor we built. Domain tables are rendered into documents and indexed as a LightRAG knowledge graph, with Amazon Nova Micro extracting entities and relations and Titan Text Embeddings v2 producing vectors. A LangGraph conversation graph loads memory from MongoDB, short-circuits exact FAQ matches, rewrites the question, retrieves in hybrid graph mode and streams the answer from Claude 3.5 Sonnet on Amazon Bedrock. We also cover how we got there: a first version on a self-hosted model with corrective and self-RAG grader nodes, replaced after per-node latency profiling by a design with four LLM calls per question instead of about fifteen. The advisor launched commercially as Dr. SmileGut on April 9, 2025.
Solution overview
The system has an offline path that turns the service's domain tables into a knowledge graph, and an online path — a LangGraph state graph behind a FastAPI streaming endpoint — that answers each question in a session. Every model runs as a managed service on Amazon Bedrock, so the application itself runs on an ordinary cloud VM without GPUs. Conversation history and per-turn summaries are kept in MongoDB.
Render domain tables into documents
Knowledge graph and vectors
Load history, match FAQ
Standalone question
Hybrid graph retrieval
Stream, then summarize the turn
The numbered steps in the diagram:
- FAQ, report-interpretation guides, a microbe dictionary, ingredient and food-to-nutrient tables, gut-type recommendations, recipes and site content are rendered into documents through per-table templates with pandas.
- LightRAG chunks the documents with overlap, extracts entities and relations with Amazon Nova Micro, embeds them with Titan Text Embeddings v2, and stores vectors in FAISS and the graph in networkx; content-hash IDs make rebuilds idempotent.
- A LangGraph state graph with a checkpointer loads the session's history from MongoDB; a question that exactly matches a curated FAQ returns the stored answer directly.
- Claude 3.5 Sonnet rewrites the question into a standalone one using the last four messages.
- LightRAG's hybrid mode retrieves entities, relations, their graph neighbourhood and the source chunks behind them, using search keywords that Claude extracts from the rewritten question.
- Claude streams the answer through a FastAPI endpoint to a Streamlit front end, and Nova Micro summarizes the turn into MongoDB for the next one.
Technology stack
| Layer | Technology | What it does here |
|---|---|---|
| Knowledge prep | pandas · per-table document templates | Domain tables rendered into retrievable documents |
| Indexing | LightRAG · Amazon Nova Micro · Titan Text Embeddings v2 | Entities, relations and chunks; content-hash document IDs |
| Stores | FAISS (vectors) · networkx (graph) | Entity, relation and chunk vectors; the knowledge graph |
| Orchestration | LangGraph StateGraph · MemorySaver checkpointer | Load memory → FAQ or rewrite → retrieve → generate → save |
| Generation | Claude 3.5 Sonnet on Amazon Bedrock (Converse streaming) | Rewrite, keywords and the streamed answer |
| Memory | MongoDB | Full turns archived; turn summaries for prompts |
| Serving | FastAPI streaming endpoint · Streamlit · cloud VM · tenacity | Streamed answers, retries with backoff, per-node timing |
Step 1: Render domain tables into documents
The knowledge behind the advisor was not a document collection but a set of structured tables, built with microbiome, diet and nutrition experts: FAQ, report-interpretation guides, a microbe dictionary, an ingredient master, food-to-nutrient tables, gut-type ingredient recommendations, recipes and site content. A table row on its own is a poor retrieval unit — a number without its column name means little to an embedding model, or to an LLM reading the context.
Each table type therefore has a template that renders a row, or a group of related rows, into a short, self-contained document that says what every value is. Documents of the same type share one shape, which helps the extraction step that follows. Each document's ID is a hash of its content, so re-running the build skips documents that are already indexed instead of duplicating them.
import hashlib
import pandas as pd
# One template per table type. Fields and wording here are invented examples.
TEMPLATES = {
"microbe": "{name} is a gut microbe. What it does: {role}. Related foods: {foods}.",
"nutrient": "{food} contains {amount} {unit} of {nutrient} per {serving}.",
"recipe": "{title} is a recipe for {gut_type}. Main ingredients: {ingredients}.",
}
def render_table(kind: str, df: pd.DataFrame):
template = TEMPLATES[kind]
for row in df.fillna("").to_dict(orient="records"):
text = template.format(**row)
doc_id = "doc-" + hashlib.md5(text.encode("utf-8")).hexdigest() # content hash
yield {"id": doc_id, "kind": kind, "text": text}
def render_all(tables: dict[str, pd.DataFrame]) -> list[dict]:
docs = {d["id"]: d for kind, df in tables.items() for d in render_table(kind, df)}
return list(docs.values()) # identical rows collapse to one documentSimplified. Real templates, table names and fields are not shown.
Step 2: Index a knowledge graph with LightRAG
LightRAG turns the documents into a knowledge graph. It splits each document into overlapping chunks, asks an LLM to extract entities — microbes, nutrients, foods, ingredients, gut types — and the relations between them, merges entities that recur across documents, and embeds entities, relations and chunks. The graph is stored with networkx and the three vector indices with FAISS.
Model choice follows cost. Extraction runs over every chunk on every rebuild, so it uses Amazon Nova Micro, the smallest model in the stack; vectors come from Titan Text Embeddings v2. Both are called through the Bedrock runtime API.
import json
import boto3
import numpy as np
from lightrag import LightRAG
from lightrag.utils import EmbeddingFunc
bedrock = boto3.client("bedrock-runtime")
async def titan_embed(texts: list[str]) -> np.ndarray:
vectors = []
for t in texts:
r = bedrock.invoke_model(modelId="amazon.titan-embed-text-v2:0",
body=json.dumps({"inputText": t, "dimensions": 1024}))
vectors.append(json.loads(r["body"].read())["embedding"])
return np.array(vectors, dtype=np.float32)
async def nova_micro(prompt, system_prompt=None, history_messages=[], **kwargs) -> str:
return converse(NOVA_MICRO_ID, prompt, system_prompt) # Bedrock Converse API
async def build_index(docs: list[dict]) -> LightRAG:
rag = LightRAG(
working_dir=KG_DIR,
llm_model_func=nova_micro, # entity/relation extraction
embedding_func=EmbeddingFunc(embedding_dim=1024, max_token_size=8192, func=titan_embed),
vector_storage="FaissVectorDBStorage",
graph_storage="NetworkXStorage",
chunk_token_size=CHUNK_TOKENS, chunk_overlap_token_size=OVERLAP_TOKENS,
)
await rag.initialize_storages()
await rag.ainsert([d["text"] for d in docs], ids=[d["id"] for d in docs])
return ragSimplified. converse() wraps the Bedrock Converse API; chunk sizes and the model ID constant are placeholders.
Why a graph? A question such as "my Bifidobacterium is low — what should I eat?" connects a microbe to the nutrients it uses, the foods that contain them and the recipes that use those foods, and each link lives in a different source table. Search over flat chunks tends to return one table's view; the graph keeps the links.
Step 3: Model the conversation as a LangGraph state graph
Each turn runs through a LangGraph StateGraph. The first node loads the session's history from MongoDB and looks the question up in the curated FAQ. On an exact match the graph returns the stored answer straight away — no LLM call, and wording the service has already vetted. Otherwise the turn runs through rewrite, retrieve, generate and save.
Graph state is checkpointed per session with LangGraph's MemorySaver, using the session ID as the thread ID. The durable record is MongoDB, and memory there has two tracks: the full text of every turn is archived, while prompts see only a short summary of each earlier turn. That keeps prompts short as a conversation grows.
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import MemorySaver
class Turn(TypedDict, total=False):
session_id: str
question: str
history: list # recent turns, as summaries, from MongoDB
faq_answer: str # curated answer on an exact FAQ match, else None
query: str # standalone rewrite of the question
context: str # entity, relation and source blocks
answer: str
async def load_memory(s: Turn) -> dict:
return {"history": await load_history(s["session_id"]),
"faq_answer": faq_lookup(s["question"])}
def route(s: Turn) -> str:
return "faq" if s.get("faq_answer") else "rewrite"
g = StateGraph(Turn)
g.add_node("load_memory", load_memory)
g.add_node("faq", serve_faq)
g.add_sequence([("rewrite", rewrite), ("retrieve", retrieve),
("generate", generate), ("save", save_turn)])
g.add_edge(START, "load_memory")
g.add_conditional_edges("load_memory", route, {"faq": "faq", "rewrite": "rewrite"})
g.add_edge("faq", END)
g.add_edge("save", END)
advisor = g.compile(checkpointer=MemorySaver())Simplified. load_history() and faq_lookup() stand in for the MongoDB and FAQ lookups; the other nodes appear in the next steps.
Step 4: Rewrite the question, then retrieve in hybrid mode
A follow-up such as "and what about breakfast?" cannot be retrieved on its own. Claude 3.5 Sonnet first rewrites each question into a standalone one using the last four messages, and retrieval then works on the question together with the person's diagnostic context rather than on the bare words of the last message.
Retrieval uses LightRAG's hybrid mode. Claude extracts two kinds of search keywords from the rewritten question: specific ones, such as a microbe or a food, which are matched against entity vectors (LightRAG's local view), and thematic ones, such as fibre intake or diversity, which are matched against relation vectors (the global view). The graph is expanded around what was found and the source chunks behind those entities and relations are attached. The context handed to the answer step is split into three blocks: entities, relations and sources. In production, a dense MMR search over a FAISS index ran alongside the graph retrieval.
from lightrag import QueryParam
async def rewrite(s: Turn) -> dict:
prompt = render_prompt("rewrite", question=s["question"], recent=s["history"][-4:])
return {"query": await claude(prompt)}
async def retrieve(s: Turn) -> dict:
kw = await claude_json(render_prompt("keywords", query=s["query"]))
param = QueryParam(
mode="hybrid", # local (entities) + global (relations)
only_need_context=True, # return the context; we generate ourselves
top_k=TOP_K,
ll_keywords=kw["specific"], # e.g. a microbe or a food -> entity vectors
hl_keywords=kw["thematic"], # e.g. fibre intake -> relation vectors
)
return {"context": await rag.aquery(s["query"], param=param)}Simplified. claude() and claude_json() wrap Bedrock calls; prompts are not shown.
What about re-ranking and keyword search? Hybrid keyword-plus-vector search, re-ranking and knowledge-graph retrieval were all evaluated during design. The pipeline that shipped uses LightRAG's graph-hybrid retrieval with dense retrieval alongside it; it has no separate keyword index and no re-ranking stage.
Step 5: Generate and stream the answer with Claude on Bedrock
The answer step sends the standalone question, the summaries of recent turns and the three context blocks to Claude 3.5 Sonnet through the Bedrock converse_stream API and forwards tokens as they arrive. Bedrock calls are retried with exponential backoff using tenacity.
There is no separate classifier node. The answer prompt itself has branches: answer from the context when the question is on topic; say so when it is unrelated; ask a clarifying question when it is ambiguous; and respond briefly to general chat. Questions beyond what a wellness service should answer are handled by those branches and by the model's own refusals. Answers are plain text rather than JSON, so they can stream.
import boto3
from botocore.exceptions import ClientError
from langgraph.config import get_stream_writer
from tenacity import retry, retry_if_exception_type, stop_after_attempt, wait_exponential
bedrock = boto3.client("bedrock-runtime")
@retry(retry=retry_if_exception_type(ClientError),
wait=wait_exponential(multiplier=1, max=MAX_WAIT), stop=stop_after_attempt(MAX_TRIES))
def open_stream(system_prompt: str, user_text: str):
return bedrock.converse_stream(
modelId=SONNET_ID,
system=[{"text": system_prompt}],
messages=[{"role": "user", "content": [{"text": user_text}]}],
inferenceConfig={"temperature": TEMPERATURE, "maxTokens": MAX_TOKENS},
)["stream"]
def generate(s: Turn) -> dict:
write = get_stream_writer() # LangGraph custom stream
prompt = render_prompt("answer", question=s["query"], history=s["history"],
context=s["context"]) # scope branches live in this prompt
parts = []
for event in open_stream(SYSTEM_PROMPT, prompt):
delta = event.get("contentBlockDelta", {}).get("delta", {})
if "text" in delta:
write({"token": delta["text"]})
parts.append(delta["text"])
return {"answer": "".join(parts)}Simplified. Model ID, retry limits and inference settings are placeholders; the answer prompt is not shown.
Step 6: Save the turn and serve it as a stream
After the answer, Nova Micro condenses the turn — question and answer — into a short summary, stored in MongoDB next to the full text. The summarizer runs on every turn, so it is the cheap model; the strong model is kept for the steps that decide answer quality.
The graph is served from a FastAPI endpoint that streams tokens to a Streamlit front end. In the sketch below, the answer node writes tokens to LangGraph's custom stream and the endpoint relays them as a streaming HTTP response. Because every model is managed on Bedrock, the service runs on a cloud VM with no GPUs.
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from langgraph.config import get_stream_writer
from pydantic import BaseModel
async def save_turn(s: Turn) -> dict:
summary = await nova_micro(render_prompt("turn_summary", q=s["question"], a=s["answer"]))
await archive_turn(s["session_id"], s["question"], s["answer"], summary) # MongoDB
return {}
def serve_faq(s: Turn) -> dict:
get_stream_writer()({"token": s["faq_answer"]}) # curated answer as one chunk
return {"answer": s["faq_answer"]}
app = FastAPI()
class ChatRequest(BaseModel):
session_id: str
question: str
@app.post("/chat")
async def chat(req: ChatRequest):
config = {"configurable": {"thread_id": req.session_id}}
inputs = {"session_id": req.session_id, "question": req.question}
async def tokens():
async for chunk in advisor.astream(inputs, config, stream_mode="custom"):
yield chunk["token"]
return StreamingResponse(tokens(), media_type="text/plain; charset=utf-8")Simplified. archive_turn() stands in for the MongoDB write; authentication and error handling are omitted.
Step 7: Profile every node, then remove what does not pay
The first version looked different. It ran a self-hosted EXAONE 3.5 (7.8B) model on vLLM with a Korean fine-tune of BGE-m3 embeddings and FAISS, inside an adaptive, corrective and self-RAG graph: a router chose the path, a document grader judged each retrieved document, web search filled gaps, and a hallucination grader checked the answer. A question could take about fifteen LLM calls, and answers came back as JSON so that the graders could parse them.
We recorded the latency of every node. The graders and classifiers cost more time than they added in quality, so they were removed: classification moved into the answer prompt, the answer became streamed plain text, and the models moved to Amazon Bedrock. The final design makes four LLM calls per question — rewrite, keyword extraction, answer and turn summary — and needs no GPU server.
import functools
import inspect
import logging
import time
log = logging.getLogger("advisor.timing")
def timed(name, node):
"""Wrap a LangGraph node so every call logs its latency."""
def done(t0):
log.info("node=%s ms=%.0f", name, (time.perf_counter() - t0) * 1000)
if inspect.iscoroutinefunction(node):
@functools.wraps(node)
async def wrapper(state):
t0 = time.perf_counter()
try:
return await node(state)
finally:
done(t0)
else:
@functools.wraps(node)
def wrapper(state):
t0 = time.perf_counter()
try:
return node(state)
finally:
done(t0)
return wrapperSimplified. In this sketch every node is registered through the wrapper, e.g. g.add_node("retrieve", timed("retrieve", retrieve)), so each call logs its latency.
| First version | Final design | |
|---|---|---|
| Models | Self-hosted EXAONE 3.5 (7.8B) on vLLM; BGE-m3 Korean embeddings | Claude 3.5 Sonnet, Nova Micro and Titan v2 on Amazon Bedrock |
| Retrieval | FAISS dense search, web search as fallback | LightRAG knowledge graph (hybrid) with dense MMR alongside |
| Control | Router, document grader, hallucination grader | Exact-FAQ shortcut; scope branches inside the answer prompt |
| LLM calls per question | About 15 | 4 |
| Output | JSON for the graders | Plain text, streamed |
| Infrastructure | GPU server for the model | Cloud VM, no GPUs |
Try the live model
The live model below runs a small version of the workflow on a fictional report, so you can compare a generic chatbot's answer with one grounded in the person's own results and in retrieved guidance.
Live model on a fictional report and 14 invented general-wellness passages. Report-field selection, retrieval (a BM25 keyword retriever and a character n-gram retriever fused by reciprocal rank, with a small boost for passages that match the report) and a rule-based check on treatment questions run live in your browser; they are simple stand-ins for the production knowledge-graph retrieval and the answer step's scope branches. The answers were written in advance and are replayed, with citation numbers bound to whatever was actually retrieved. Not medical advice. Open the live model on its own page ↗
Results
The advisor launched commercially as Dr. SmileGut on April 9, 2025. Answer quality was assessed by people against a rubric on a multi-turn question set — accuracy, fluency, consistency and fit to the conversation — rather than by a single automatic score.
The project established a personal-advisor model that combines validated domain knowledge with personal health data, strengthened the customer experience of the diagnostic service, and laid a technical foundation for follow-on services in diet, probiotics, supplements and personalized health management.
Lessons learned
- Profile before adding intelligence. Router and grader nodes looked like quality features; per-node timing showed what they cost. Removing them cut LLM calls per question from about fifteen to four and made streaming possible.
- Tier models by how often they run. Extraction and turn summaries run constantly and work with a small model; the rewrite, keywords and answer decide quality and get the strong one.
- Render tables deliberately. Templates that turn rows into self-describing documents mattered for retrieval as much as the choice of index.
- Let curated answers stay curated. An exact FAQ match returns the vetted answer with no model in the loop.
- Streaming changes the design. Once answers stream, JSON output and after-the-fact graders no longer fit; checks move into the prompt or ahead of generation.
Conclusion
A health advisor has to answer from validated knowledge and from the person's own results, consistently, and fast enough for a chat. A knowledge graph built from the service's domain tables, a small LangGraph graph with explicit memory, an FAQ shortcut, hybrid graph retrieval and a streamed answer from a strong managed model met those requirements with four model calls per question.
The path there is the more general lesson: start with the patterns that look safest — routers, graders, verifiers — measure what each node costs, and keep only what earns its latency.
Limitations
- The advisor provides general wellness information; it does not diagnose or treat, and it says so.
- Answer quality is bounded by the knowledge base and by the extracted graph; a relation the extraction model missed or got wrong is invisible to retrieval.
- Out-of-scope questions are handled by prompt branches and the model's own refusals rather than by a separate guardrail service.
- The live model's retrievers, fusion and rule-based treatment check are simple in-browser stand-ins for LightRAG graph retrieval and the answer step's scope branches, and its answers are pre-written and replayed.
About the demo and confidentiality
The user, report values, knowledge passages and answers in the embedded model are invented general-wellness text. No customer data or conversation, knowledge-source text, prompt, credential or infrastructure detail from the production service appears in this post; code is simplified and written for illustration.