GenAI · RAG · Document AICJ AI Center, with CJ Olive YoungTechnical write-up · October 2026 · 12 min read

Building an ad-copy legal-risk reviewer with OCR, retrieval and an independent LLM critic

A walkthrough of the system that reviews beauty and food advertising for legal risk: how ad images become located text, how a policy knowledge base is built and searched, how an LLM judge and a separate critic make the call, and how it all runs as a service — with simplified code for each step.

Built withPythonPyMuPDFPlaywrightCloud document OCRText embeddingsChromaDBManaged LLMpyahocorasickFastAPIPostgreSQLDocker

Every phrase in a cosmetics or food advertisement has to clear advertising and labelling law, and the same words can be acceptable in one product category and a violation in another. Review at a large retailer had relied on people reading every ad, keyword lists and past cases scattered across documents. Keyword filters miss paraphrases and flag innocent uses; a general-purpose LLM does not know industry-specific rules, cannot show its basis, and does not say where on the ad the problem is.

In this post we walk through the system we built instead. It reads an ad image with OCR that keeps coordinates, retrieves the most similar violation examples for the product's category from a policy knowledge base, lets an LLM judge decide with those examples in front of it, and runs an independent critic that may withdraw weak flags before anyone sees them. A keyword path runs in parallel for terms the client prohibits outright. The result is a report that marks each risky line on the original image with its category, risk level, reason and the case it resembles. The system is in production and is patented, with the author as lead inventor.

Solution overview

The system has an offline path that builds and updates the policy knowledge base, and an online path that reviews each ad. Both are Python services; the online path is exposed through an asynchronous API and stores jobs in PostgreSQL so they survive restarts.

Architecture
1Offline

Build the policy knowledge base

Managed LLMEmbeddingsChromaDB
2Ingest

Render images, PDFs and web pages

PyMuPDFPlaywright
3OCR

Text blocks with coordinates

Document OCRunion-find
4Retrieve

Recall-first candidate search

ChromaDBsubstring match
5Judge

LLM decision per category

Managed LLMJSON schema
6Verify

One-way critic

Managed LLM
7Serve

Located report and API

FastAPIPostgreSQLDocker
Offline, policy examples are enriched and indexed per risk category. Online, each ad is rendered, read with coordinates, matched against the index, judged, audited and returned as a located report. A keyword path (not drawn) runs in parallel with steps 4–6.

The numbered steps in the diagram:

  1. Policy examples for each risk category are enriched with definitions, category boundaries, core expressions and keywords, embedded and indexed in ChromaDB with cosine HNSW. The index is updated incrementally every day.
  2. Images are taken as-is, PDFs are rendered at 200 dpi with PyMuPDF, and web pages are captured full-page with Playwright; very tall images are split at gaps between content.
  3. Document OCR returns words with bounding boxes. Words are rebuilt into lines and blocks, nearby blocks are merged with union-find, and footnote markers are linked to the lines they qualify.
  4. For every sentence, dense search over the index and exact substring matching against examples produce scored candidates; the best score per example is kept, then the top categories and their evidence.
  5. An LLM judges each candidate category at temperature 0 with the retrieved examples and returns a decision and a reason as structured JSON.
  6. A separate critic audits violation flags that are not high-confidence and can only clear them — never create new ones.
  7. Results are merged with the keyword path into block-level JSON and an HTML overlay and returned through async, batch and sync APIs.

Technology stack

LayerTechnologyWhat it does here
IngestionPyMuPDF · Playwright · image preprocessingTurns images, PDFs and web pages into OCR-ready images
OCRManaged document-text OCR · union-find mergingText blocks with coordinates; footnotes attached to their lines
Knowledge baseManaged text embeddings · ChromaDB (cosine HNSW)Per-category policy examples and enriched expressions
JudgmentManaged lightweight LLM · temperature 0 · JSON schemaPer-category decision with a written reason
Keywordspyahocorasick · Unicode NFKC normalizationProhibited-keyword path running in parallel
ServiceFastAPI · PostgreSQL (JSONB) · object storage · Docker ComposeAsync, batch and sync APIs with job recovery and callbacks
EvaluationSciPy linear_sum_assignment · pytestBlock-level matching and micro/macro precision, recall and F1

Step 1: Build a per-category policy knowledge base

The client's policy tables map violating phrases to a risk category and a risk level, separately for food, cosmetics and functional cosmetics. Raw examples are too terse to retrieve well, so an offline job asks an LLM to enrich each policy: a definition, what the category covers and where it ends relative to neighbouring categories, the core expressions that signal it, and hard keywords. Generic expressions that show up in three or more policies are dropped because they discriminate nothing.

Everything is embedded and written to ChromaDB with the category, domain and risk level as metadata, so a search can be restricted to the ad's product domain.

kb/build.py
import chromadb

client = chromadb.PersistentClient(path="kb")
kb = client.get_or_create_collection(
    "policy_examples", metadata={"hnsw:space": "cosine"})

def upsert_policy(policy):
    """One policy = a risk category with examples and enriched expressions."""
    items = policy.examples + policy.core_expressions
    kb.upsert(
        ids=[f"{policy.id}:{i}" for i in range(len(items))],
        documents=[it.text for it in items],
        embeddings=embed([it.text for it in items]),
        metadatas=[{"category": policy.category, "domain": policy.domain,
                    "risk": it.risk, "kind": it.kind} for it in items],
    )

Simplified. embed() wraps the managed embedding model; enrichment prompts are not shown.

Updates are incremental. A policy is regenerated only when enough of its examples change; otherwise only the changed vectors are patched, and keyword lists are merged only when their fingerprint changes. That keeps a daily sync with the client's policy system cheap.

Step 2: Read the ad as located text

Review needs to point at the line, so OCR output keeps coordinates end to end. PDFs are rendered page by page; web pages are captured full-page in a headless browser. Images are upscaled and contrast-boosted before OCR, with a fallback to the original if that fails, and very tall product pages are split at whitespace gaps.

OCR returns words with boxes. The system rebuilds them into lines and paragraphs, then merges blocks that are vertically adjacent and overlap horizontally — a union-find pass over block pairs. Footnote markers such as * are linked to the body line they qualify, because a footnote can turn a violation into an acceptable claim.

ocr/blocks.py
def merge_blocks(blocks, gap):
    """Union-find over OCR blocks that touch vertically and overlap horizontally."""
    parent = list(range(len(blocks)))

    def find(i):
        while parent[i] != i:
            parent[i] = parent[parent[i]]
            i = parent[i]
        return i

    for i, a in enumerate(blocks):
        for j in range(i + 1, len(blocks)):
            b = blocks[j]
            touches = 0 <= b.top - a.bottom <= gap
            overlaps = min(a.right, b.right) > max(a.left, b.left)
            if touches and overlaps:
                parent[find(j)] = find(i)

    groups = {}
    for i, blk in enumerate(blocks):
        groups.setdefault(find(i), []).append(blk)
    return [Block.union(g) for g in groups.values()]

Simplified block merging. Real thresholds depend on font size and layout.

Step 3: Retrieve candidates for recall

Retrieval is tuned for recall: it should surface every plausible violation and leave precision to the judge. Each block and each sentence inside it is searched densely in the ad's domain, and every example that appears verbatim in the sentence gets a perfect score. Scores are merged by keeping the maximum per example — no weighted sum to tune.

score(e | s) = 1.0                    if example e occurs in sentence s
             = 1 − cos_dist(s, e)     otherwise              keep the max per example

evidence     = top-3 categories by max score,  ≤ 3 examples each,  score ≥ τ
review/retrieve.py
from collections import defaultdict

def candidates(sentence, domain, tau, top_cats=3, per_cat=3):
    best = {}
    res = kb.query(query_embeddings=embed([sentence]), n_results=20,
                   where={"domain": domain})
    for ex_id, dist, meta in zip(res["ids"][0], res["distances"][0],
                                 res["metadatas"][0]):
        best[ex_id] = max(best.get(ex_id, (0.0, meta)), (1 - dist, meta),
                          key=lambda x: x[0])
    for ex in examples_in(sentence, domain):          # verbatim hits
        best[ex.id] = (1.0, ex.meta)

    by_cat = defaultdict(list)
    for ex_id, (score, meta) in best.items():
        if score >= tau:
            by_cat[meta["category"]].append((score, ex_id))
    ranked = sorted(by_cat.items(), key=lambda kv: -max(kv[1])[0])[:top_cats]
    return {cat: sorted(hits, reverse=True)[:per_cat] for cat, hits in ranked}

Simplified candidate search. examples_in() is an exact-substring lookup over the domain's examples.

Step 4: Judge with an LLM, then audit with a one-way critic

The judge sees a handful of sentences, one candidate category, that category's rules and the retrieved evidence, and must answer in a fixed JSON schema at temperature 0. Batching a few sentences per call keeps cost proportional to the ad's length; a failed batch is retried one sentence at a time. When the best evidence is a verbatim match the judge runs in a lenient mode, otherwise in a strict one.

The critic is a separate call with a separate job: look at each violation flag that is not high-confidence and ask whether its reason is supported, whether it is consistent with similar lines on the same ad that were cleared, and whether it matched only a topic rather than a claim. It can clear a flag; it can never raise one. On any error the original decision stands.

review/judge.py
JUDGE = {"type": "object", "required": ["violation", "reason"],
         "properties": {"violation": {"type": "boolean"},
                        "reason": {"type": "string"}}}

def judge(sentences, category, evidence):
    return llm.generate(render_judge(sentences, category, evidence),
                        schema=JUDGE, temperature=0)

def critic(flags, siblings):
    """One-way audit: may clear a flag, never create one."""
    for f in flags:
        if not f.violation or f.high_confidence:
            continue
        try:
            v = llm.generate(render_critic(f, siblings), schema=CRITIC,
                             temperature=0)
        except Exception:
            continue                      # keep the judge's decision
        if v["clear"]:
            f.violation, f.cleared_because = False, v["why"]
    return flags

Simplified. Prompts are rendered from templates that are not shown; llm.generate enforces the JSON schema.

Why one-way? A second model that can add flags doubles the ways to be wrong. Restricting the critic to removals bounds its effect on recall while still cutting false alarms — the design choice behind the patent's independent verification unit.

Step 5: Run prohibited keywords in parallel

Some terms are prohibited outright by the client regardless of context. These go through two Aho-Corasick automata — one over normalized text, one with whitespace removed so spaced-out spellings still match — and a single LLM call per image removes only matches that are clearly out of context. If that call fails, every match is kept.

review/keywords.py
import unicodedata
import ahocorasick

def norm(s, compact=False):
    s = unicodedata.normalize("NFKC", s).lower()
    return "".join(s.split()) if compact else s

def build(keywords, compact=False):
    a = ahocorasick.Automaton()
    for kw in keywords:
        a.add_word(norm(kw, compact), kw)
    a.make_automaton()
    return a

def hits(automaton, text, compact=False):
    t = norm(text, compact)
    return [(end - len(norm(kw, compact)) + 1, kw)
            for end, kw in automaton.iter(t)]

Simplified. Two automata over NFKC-normalized text, built once per keyword-list version.

Step 6: Serve it as an asynchronous API

The service accepts single ads, batches and a synchronous mode for small inputs. Each job is persisted in PostgreSQL (results as JSONB) so work resumes after a restart, inputs and overlays go to object storage, and finished jobs trigger a callback to the client. The retrieval path and the keyword path for one ad run concurrently.

service/api.py
import asyncio
from fastapi import FastAPI

app = FastAPI()

@app.post("/reviews")
async def review(req: ReviewRequest) -> Report:
    pages = await render(req.source)                 # image | PDF | URL
    blocks = await ocr_with_boxes(pages)
    rag, kw = await asyncio.gather(
        rag_path(blocks, req.domain),                # steps 3-4
        keyword_path(blocks),                        # step 5
    )
    return Report.merge(blocks, rag, kw)             # JSON + HTML overlay

Simplified endpoint; authentication, limits and job persistence are omitted.

Concurrency is capped per process, judge and critic calls run in small thread pools, embeddings are cached, and duplicate sentences across an ad are processed once. Token usage and cost are metered per job.

Step 7: Measure it the way reviewers work

Evaluation matches labelled ground-truth phrases to OCR blocks (exact, then whitespace-insensitive, then fuzzy), and then scores three things separately: was the phrase detected, was the category right, and was the risk level right. Category matching is an assignment problem — each predicted category is paired with at most one labelled category — solved with the Hungarian algorithm, and micro and macro precision, recall and F1 are reported.

eval/match.py
import numpy as np
from scipy.optimize import linear_sum_assignment

def match_categories(pred, gold, sim):
    """Pair predictions with labels to maximize total similarity."""
    reward = np.array([[sim(p, g) for g in gold] for p in pred])
    rows, cols = linear_sum_assignment(reward, maximize=True)
    return [(pred[r], gold[c]) for r, c in zip(rows, cols) if reward[r, c] > 0]

Simplified. Assignment between predicted and labelled categories on one block.

Try the live model

The live model below reproduces the decision logic on two invented ads, so you can watch each stage change the result: a keyword filter alone, retrieval plus LLM judgment, and the independent verifier.

Retrieval (character-trigram TF-IDF over a small invented case store), keyword matching and the precision/recall at each stage are computed in your browser. The LLM judgments and the critic's notes were produced in advance for these two ads and are replayed. Open the live model on its own page ↗

Results

81.14%precision on 499 production ad images
94.96%recall on the same images
Patentfiled with the author as lead inventor; in production

Measured on 499 ad images from operation, the system reached 81.14% precision and 94.96% recall. Every flag comes with its category, risk level, reason, the case it resembles and its location on the original image, which is what makes the output usable by reviewers and defensible afterwards.

Because policies live in the knowledge base rather than in model weights, a new category or a changed rule is handled by adding or updating references — no retraining.

Lessons learned

Conclusion

Ad review is a regulated, category-dependent judgment that neither keyword lists nor a general LLM handle well on their own. Combining located OCR, a per-category knowledge base, recall-first retrieval, a structured LLM judge and a one-way critic produced a system that explains and locates every flag and adapts to new rules without retraining.

The same pattern — retrieval for recall, an LLM for precision, an independent check before a person sees the result — applies to other compliance reviews where the rules change and the basis for each decision has to be shown.

Limitations

About the demo and confidentiality

All products, ad copy and cases in the embedded model are invented. No client policy text, keyword list, prompt, model or vendor configuration, threshold or production data appears in this post; code is simplified and written for illustration.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterLead inventor on the patent; system design, retrieval and verification architecture, evaluation framework. Demo built on invented ads for this site.