GenAI · RAG · Knowledge graphCJ AI Center, with the SmileGut serviceTechnical write-up · October 2026 · 13 min read

Building a gut-health advisor with LightRAG, LangGraph and Claude on Amazon Bedrock

A walkthrough of the retrieval-augmented advisor behind a microbiome diagnostic service: how domain tables become a knowledge graph, how a LangGraph conversation graph handles memory, rewriting, retrieval and streaming, and why a first version with grader nodes on a self-hosted model gave way to a four-call design on Amazon Bedrock — with simplified code for each step.

Built withPythonLightRAGFAISSnetworkxLangGraphMongoDBAmazon BedrockClaude 3.5 SonnetNova Micro · Titan Embeddings v2FastAPIStreamlit

SmileGut is a microbiome-based diagnostic and health service. Customers who receive a gut-microbiome report want to know what their results mean and what to do next. The knowledge needed to answer is spread across report-interpretation guides, a microbe dictionary, ingredient and nutrient tables, recipes and an FAQ. An FAQ page cannot personalize, and a general chatbot answers from its own memory, which a health product cannot rely on for accuracy, currency or stability. Most real questions also cross sources — a microbe, the nutrients it uses, the foods that contain them — and flat text chunks lose those links.

In this post we walk through the advisor we built. Domain tables are rendered into documents and indexed as a LightRAG knowledge graph, with Amazon Nova Micro extracting entities and relations and Titan Text Embeddings v2 producing vectors. A LangGraph conversation graph loads memory from MongoDB, short-circuits exact FAQ matches, rewrites the question, retrieves in hybrid graph mode and streams the answer from Claude 3.5 Sonnet on Amazon Bedrock. We also cover how we got there: a first version on a self-hosted model with corrective and self-RAG grader nodes, replaced after per-node latency profiling by a design with four LLM calls per question instead of about fifteen. The advisor launched commercially as Dr. SmileGut on April 9, 2025.

Solution overview

The system has an offline path that turns the service's domain tables into a knowledge graph, and an online path — a LangGraph state graph behind a FastAPI streaming endpoint — that answers each question in a session. Every model runs as a managed service on Amazon Bedrock, so the application itself runs on an ordinary cloud VM without GPUs. Conversation history and per-turn summaries are kept in MongoDB.

Architecture
1Offline

Render domain tables into documents

pandastemplates
2Index

Knowledge graph and vectors

LightRAGNova MicroTitan v2
3Memory

Load history, match FAQ

LangGraphMongoDB
4Rewrite

Standalone question

Claude 3.5 Sonnet
5Retrieve

Hybrid graph retrieval

LightRAG hybridFAISSnetworkx
6Answer

Stream, then summarize the turn

Claude 3.5 SonnetNova MicroFastAPI
Offline, domain tables become documents and a knowledge graph with vector indices. Online, each turn loads memory and either returns a curated FAQ answer or rewrites, retrieves, streams an answer and saves a summary of the turn.

The numbered steps in the diagram:

  1. FAQ, report-interpretation guides, a microbe dictionary, ingredient and food-to-nutrient tables, gut-type recommendations, recipes and site content are rendered into documents through per-table templates with pandas.
  2. LightRAG chunks the documents with overlap, extracts entities and relations with Amazon Nova Micro, embeds them with Titan Text Embeddings v2, and stores vectors in FAISS and the graph in networkx; content-hash IDs make rebuilds idempotent.
  3. A LangGraph state graph with a checkpointer loads the session's history from MongoDB; a question that exactly matches a curated FAQ returns the stored answer directly.
  4. Claude 3.5 Sonnet rewrites the question into a standalone one using the last four messages.
  5. LightRAG's hybrid mode retrieves entities, relations, their graph neighbourhood and the source chunks behind them, using search keywords that Claude extracts from the rewritten question.
  6. Claude streams the answer through a FastAPI endpoint to a Streamlit front end, and Nova Micro summarizes the turn into MongoDB for the next one.

Technology stack

LayerTechnologyWhat it does here
Knowledge preppandas · per-table document templatesDomain tables rendered into retrievable documents
IndexingLightRAG · Amazon Nova Micro · Titan Text Embeddings v2Entities, relations and chunks; content-hash document IDs
StoresFAISS (vectors) · networkx (graph)Entity, relation and chunk vectors; the knowledge graph
OrchestrationLangGraph StateGraph · MemorySaver checkpointerLoad memory → FAQ or rewrite → retrieve → generate → save
GenerationClaude 3.5 Sonnet on Amazon Bedrock (Converse streaming)Rewrite, keywords and the streamed answer
MemoryMongoDBFull turns archived; turn summaries for prompts
ServingFastAPI streaming endpoint · Streamlit · cloud VM · tenacityStreamed answers, retries with backoff, per-node timing

Step 1: Render domain tables into documents

The knowledge behind the advisor was not a document collection but a set of structured tables, built with microbiome, diet and nutrition experts: FAQ, report-interpretation guides, a microbe dictionary, an ingredient master, food-to-nutrient tables, gut-type ingredient recommendations, recipes and site content. A table row on its own is a poor retrieval unit — a number without its column name means little to an embedding model, or to an LLM reading the context.

Each table type therefore has a template that renders a row, or a group of related rows, into a short, self-contained document that says what every value is. Documents of the same type share one shape, which helps the extraction step that follows. Each document's ID is a hash of its content, so re-running the build skips documents that are already indexed instead of duplicating them.

kb/render.py
import hashlib
import pandas as pd

# One template per table type. Fields and wording here are invented examples.
TEMPLATES = {
    "microbe": "{name} is a gut microbe. What it does: {role}. Related foods: {foods}.",
    "nutrient": "{food} contains {amount} {unit} of {nutrient} per {serving}.",
    "recipe": "{title} is a recipe for {gut_type}. Main ingredients: {ingredients}.",
}

def render_table(kind: str, df: pd.DataFrame):
    template = TEMPLATES[kind]
    for row in df.fillna("").to_dict(orient="records"):
        text = template.format(**row)
        doc_id = "doc-" + hashlib.md5(text.encode("utf-8")).hexdigest()   # content hash
        yield {"id": doc_id, "kind": kind, "text": text}

def render_all(tables: dict[str, pd.DataFrame]) -> list[dict]:
    docs = {d["id"]: d for kind, df in tables.items() for d in render_table(kind, df)}
    return list(docs.values())          # identical rows collapse to one document

Simplified. Real templates, table names and fields are not shown.

Step 2: Index a knowledge graph with LightRAG

LightRAG turns the documents into a knowledge graph. It splits each document into overlapping chunks, asks an LLM to extract entities — microbes, nutrients, foods, ingredients, gut types — and the relations between them, merges entities that recur across documents, and embeds entities, relations and chunks. The graph is stored with networkx and the three vector indices with FAISS.

Model choice follows cost. Extraction runs over every chunk on every rebuild, so it uses Amazon Nova Micro, the smallest model in the stack; vectors come from Titan Text Embeddings v2. Both are called through the Bedrock runtime API.

kb/index.py
import json
import boto3
import numpy as np
from lightrag import LightRAG
from lightrag.utils import EmbeddingFunc
bedrock = boto3.client("bedrock-runtime")

async def titan_embed(texts: list[str]) -> np.ndarray:
    vectors = []
    for t in texts:
        r = bedrock.invoke_model(modelId="amazon.titan-embed-text-v2:0",
                                 body=json.dumps({"inputText": t, "dimensions": 1024}))
        vectors.append(json.loads(r["body"].read())["embedding"])
    return np.array(vectors, dtype=np.float32)

async def nova_micro(prompt, system_prompt=None, history_messages=[], **kwargs) -> str:
    return converse(NOVA_MICRO_ID, prompt, system_prompt)   # Bedrock Converse API

async def build_index(docs: list[dict]) -> LightRAG:
    rag = LightRAG(
        working_dir=KG_DIR,
        llm_model_func=nova_micro,                           # entity/relation extraction
        embedding_func=EmbeddingFunc(embedding_dim=1024, max_token_size=8192, func=titan_embed),
        vector_storage="FaissVectorDBStorage",
        graph_storage="NetworkXStorage",
        chunk_token_size=CHUNK_TOKENS, chunk_overlap_token_size=OVERLAP_TOKENS,
    )
    await rag.initialize_storages()
    await rag.ainsert([d["text"] for d in docs], ids=[d["id"] for d in docs])
    return rag

Simplified. converse() wraps the Bedrock Converse API; chunk sizes and the model ID constant are placeholders.

Why a graph? A question such as "my Bifidobacterium is low — what should I eat?" connects a microbe to the nutrients it uses, the foods that contain them and the recipes that use those foods, and each link lives in a different source table. Search over flat chunks tends to return one table's view; the graph keeps the links.

Step 3: Model the conversation as a LangGraph state graph

Each turn runs through a LangGraph StateGraph. The first node loads the session's history from MongoDB and looks the question up in the curated FAQ. On an exact match the graph returns the stored answer straight away — no LLM call, and wording the service has already vetted. Otherwise the turn runs through rewrite, retrieve, generate and save.

Graph state is checkpointed per session with LangGraph's MemorySaver, using the session ID as the thread ID. The durable record is MongoDB, and memory there has two tracks: the full text of every turn is archived, while prompts see only a short summary of each earlier turn. That keeps prompts short as a conversation grows.

advisor/graph.py
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import MemorySaver

class Turn(TypedDict, total=False):
    session_id: str
    question: str
    history: list        # recent turns, as summaries, from MongoDB
    faq_answer: str      # curated answer on an exact FAQ match, else None
    query: str           # standalone rewrite of the question
    context: str         # entity, relation and source blocks
    answer: str

async def load_memory(s: Turn) -> dict:
    return {"history": await load_history(s["session_id"]),
            "faq_answer": faq_lookup(s["question"])}

def route(s: Turn) -> str:
    return "faq" if s.get("faq_answer") else "rewrite"

g = StateGraph(Turn)
g.add_node("load_memory", load_memory)
g.add_node("faq", serve_faq)
g.add_sequence([("rewrite", rewrite), ("retrieve", retrieve),
                ("generate", generate), ("save", save_turn)])
g.add_edge(START, "load_memory")
g.add_conditional_edges("load_memory", route, {"faq": "faq", "rewrite": "rewrite"})
g.add_edge("faq", END)
g.add_edge("save", END)
advisor = g.compile(checkpointer=MemorySaver())

Simplified. load_history() and faq_lookup() stand in for the MongoDB and FAQ lookups; the other nodes appear in the next steps.

Step 4: Rewrite the question, then retrieve in hybrid mode

A follow-up such as "and what about breakfast?" cannot be retrieved on its own. Claude 3.5 Sonnet first rewrites each question into a standalone one using the last four messages, and retrieval then works on the question together with the person's diagnostic context rather than on the bare words of the last message.

Retrieval uses LightRAG's hybrid mode. Claude extracts two kinds of search keywords from the rewritten question: specific ones, such as a microbe or a food, which are matched against entity vectors (LightRAG's local view), and thematic ones, such as fibre intake or diversity, which are matched against relation vectors (the global view). The graph is expanded around what was found and the source chunks behind those entities and relations are attached. The context handed to the answer step is split into three blocks: entities, relations and sources. In production, a dense MMR search over a FAISS index ran alongside the graph retrieval.

advisor/retrieve.py
from lightrag import QueryParam

async def rewrite(s: Turn) -> dict:
    prompt = render_prompt("rewrite", question=s["question"], recent=s["history"][-4:])
    return {"query": await claude(prompt)}

async def retrieve(s: Turn) -> dict:
    kw = await claude_json(render_prompt("keywords", query=s["query"]))
    param = QueryParam(
        mode="hybrid",                 # local (entities) + global (relations)
        only_need_context=True,        # return the context; we generate ourselves
        top_k=TOP_K,
        ll_keywords=kw["specific"],    # e.g. a microbe or a food -> entity vectors
        hl_keywords=kw["thematic"],    # e.g. fibre intake -> relation vectors
    )
    return {"context": await rag.aquery(s["query"], param=param)}

Simplified. claude() and claude_json() wrap Bedrock calls; prompts are not shown.

What about re-ranking and keyword search? Hybrid keyword-plus-vector search, re-ranking and knowledge-graph retrieval were all evaluated during design. The pipeline that shipped uses LightRAG's graph-hybrid retrieval with dense retrieval alongside it; it has no separate keyword index and no re-ranking stage.

Step 5: Generate and stream the answer with Claude on Bedrock

The answer step sends the standalone question, the summaries of recent turns and the three context blocks to Claude 3.5 Sonnet through the Bedrock converse_stream API and forwards tokens as they arrive. Bedrock calls are retried with exponential backoff using tenacity.

There is no separate classifier node. The answer prompt itself has branches: answer from the context when the question is on topic; say so when it is unrelated; ask a clarifying question when it is ambiguous; and respond briefly to general chat. Questions beyond what a wellness service should answer are handled by those branches and by the model's own refusals. Answers are plain text rather than JSON, so they can stream.

advisor/generate.py
import boto3
from botocore.exceptions import ClientError
from langgraph.config import get_stream_writer
from tenacity import retry, retry_if_exception_type, stop_after_attempt, wait_exponential

bedrock = boto3.client("bedrock-runtime")

@retry(retry=retry_if_exception_type(ClientError),
       wait=wait_exponential(multiplier=1, max=MAX_WAIT), stop=stop_after_attempt(MAX_TRIES))
def open_stream(system_prompt: str, user_text: str):
    return bedrock.converse_stream(
        modelId=SONNET_ID,
        system=[{"text": system_prompt}],
        messages=[{"role": "user", "content": [{"text": user_text}]}],
        inferenceConfig={"temperature": TEMPERATURE, "maxTokens": MAX_TOKENS},
    )["stream"]

def generate(s: Turn) -> dict:
    write = get_stream_writer()                   # LangGraph custom stream
    prompt = render_prompt("answer", question=s["query"], history=s["history"],
                           context=s["context"])  # scope branches live in this prompt
    parts = []
    for event in open_stream(SYSTEM_PROMPT, prompt):
        delta = event.get("contentBlockDelta", {}).get("delta", {})
        if "text" in delta:
            write({"token": delta["text"]})
            parts.append(delta["text"])
    return {"answer": "".join(parts)}

Simplified. Model ID, retry limits and inference settings are placeholders; the answer prompt is not shown.

Step 6: Save the turn and serve it as a stream

After the answer, Nova Micro condenses the turn — question and answer — into a short summary, stored in MongoDB next to the full text. The summarizer runs on every turn, so it is the cheap model; the strong model is kept for the steps that decide answer quality.

The graph is served from a FastAPI endpoint that streams tokens to a Streamlit front end. In the sketch below, the answer node writes tokens to LangGraph's custom stream and the endpoint relays them as a streaming HTTP response. Because every model is managed on Bedrock, the service runs on a cloud VM with no GPUs.

service/api.py
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from langgraph.config import get_stream_writer
from pydantic import BaseModel

async def save_turn(s: Turn) -> dict:
    summary = await nova_micro(render_prompt("turn_summary", q=s["question"], a=s["answer"]))
    await archive_turn(s["session_id"], s["question"], s["answer"], summary)   # MongoDB
    return {}

def serve_faq(s: Turn) -> dict:
    get_stream_writer()({"token": s["faq_answer"]})      # curated answer as one chunk
    return {"answer": s["faq_answer"]}

app = FastAPI()

class ChatRequest(BaseModel):
    session_id: str
    question: str

@app.post("/chat")
async def chat(req: ChatRequest):
    config = {"configurable": {"thread_id": req.session_id}}
    inputs = {"session_id": req.session_id, "question": req.question}

    async def tokens():
        async for chunk in advisor.astream(inputs, config, stream_mode="custom"):
            yield chunk["token"]

    return StreamingResponse(tokens(), media_type="text/plain; charset=utf-8")

Simplified. archive_turn() stands in for the MongoDB write; authentication and error handling are omitted.

Step 7: Profile every node, then remove what does not pay

The first version looked different. It ran a self-hosted EXAONE 3.5 (7.8B) model on vLLM with a Korean fine-tune of BGE-m3 embeddings and FAISS, inside an adaptive, corrective and self-RAG graph: a router chose the path, a document grader judged each retrieved document, web search filled gaps, and a hallucination grader checked the answer. A question could take about fifteen LLM calls, and answers came back as JSON so that the graders could parse them.

We recorded the latency of every node. The graders and classifiers cost more time than they added in quality, so they were removed: classification moved into the answer prompt, the answer became streamed plain text, and the models moved to Amazon Bedrock. The final design makes four LLM calls per question — rewrite, keyword extraction, answer and turn summary — and needs no GPU server.

advisor/timing.py
import functools
import inspect
import logging
import time

log = logging.getLogger("advisor.timing")

def timed(name, node):
    """Wrap a LangGraph node so every call logs its latency."""
    def done(t0):
        log.info("node=%s ms=%.0f", name, (time.perf_counter() - t0) * 1000)

    if inspect.iscoroutinefunction(node):
        @functools.wraps(node)
        async def wrapper(state):
            t0 = time.perf_counter()
            try:
                return await node(state)
            finally:
                done(t0)
    else:
        @functools.wraps(node)
        def wrapper(state):
            t0 = time.perf_counter()
            try:
                return node(state)
            finally:
                done(t0)
    return wrapper

Simplified. In this sketch every node is registered through the wrapper, e.g. g.add_node("retrieve", timed("retrieve", retrieve)), so each call logs its latency.

First versionFinal design
ModelsSelf-hosted EXAONE 3.5 (7.8B) on vLLM; BGE-m3 Korean embeddingsClaude 3.5 Sonnet, Nova Micro and Titan v2 on Amazon Bedrock
RetrievalFAISS dense search, web search as fallbackLightRAG knowledge graph (hybrid) with dense MMR alongside
ControlRouter, document grader, hallucination graderExact-FAQ shortcut; scope branches inside the answer prompt
LLM calls per questionAbout 154
OutputJSON for the gradersPlain text, streamed
InfrastructureGPU server for the modelCloud VM, no GPUs

Try the live model

The live model below runs a small version of the workflow on a fictional report, so you can compare a generic chatbot's answer with one grounded in the person's own results and in retrieved guidance.

Live model on a fictional report and 14 invented general-wellness passages. Report-field selection, retrieval (a BM25 keyword retriever and a character n-gram retriever fused by reciprocal rank, with a small boost for passages that match the report) and a rule-based check on treatment questions run live in your browser; they are simple stand-ins for the production knowledge-graph retrieval and the answer step's scope branches. The answers were written in advance and are replayed, with citation numbers bound to whatever was actually retrieved. Not medical advice. Open the live model on its own page ↗

Results

Launchedcommercially as Dr. SmileGut on April 9, 2025
End to enddata processing, RAG design, LangGraph workflow, prompts and architecture
Foundationfor diet, probiotic, supplement and personalized health services

The advisor launched commercially as Dr. SmileGut on April 9, 2025. Answer quality was assessed by people against a rubric on a multi-turn question set — accuracy, fluency, consistency and fit to the conversation — rather than by a single automatic score.

The project established a personal-advisor model that combines validated domain knowledge with personal health data, strengthened the customer experience of the diagnostic service, and laid a technical foundation for follow-on services in diet, probiotics, supplements and personalized health management.

Lessons learned

Conclusion

A health advisor has to answer from validated knowledge and from the person's own results, consistently, and fast enough for a chat. A knowledge graph built from the service's domain tables, a small LangGraph graph with explicit memory, an FAQ shortcut, hybrid graph retrieval and a streamed answer from a strong managed model met those requirements with four model calls per question.

The path there is the more general lesson: start with the patterns that look safest — routers, graders, verifiers — measure what each node costs, and keep only what earns its latency.

Limitations

About the demo and confidentiality

The user, report values, knowledge passages and answers in the embedded model are invented general-wellness text. No customer data or conversation, knowledge-source text, prompt, credential or infrastructure detail from the production service appears in this post; code is simplified and written for illustration.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterData processing, RAG design, LangGraph workflow, prompt optimization and system architecture. Demo built on fictional data for this site.