Building an ad-copy legal-risk reviewer with OCR, retrieval and an independent LLM critic
A walkthrough of the system that reviews beauty and food advertising for legal risk: how ad images become located text, how a policy knowledge base is built and searched, how an LLM judge and a separate critic make the call, and how it all runs as a service — with simplified code for each step.
Every phrase in a cosmetics or food advertisement has to clear advertising and labelling law, and the same words can be acceptable in one product category and a violation in another. Review at a large retailer had relied on people reading every ad, keyword lists and past cases scattered across documents. Keyword filters miss paraphrases and flag innocent uses; a general-purpose LLM does not know industry-specific rules, cannot show its basis, and does not say where on the ad the problem is.
In this post we walk through the system we built instead. It reads an ad image with OCR that keeps coordinates, retrieves the most similar violation examples for the product's category from a policy knowledge base, lets an LLM judge decide with those examples in front of it, and runs an independent critic that may withdraw weak flags before anyone sees them. A keyword path runs in parallel for terms the client prohibits outright. The result is a report that marks each risky line on the original image with its category, risk level, reason and the case it resembles. The system is in production and is patented, with the author as lead inventor.
Solution overview
The system has an offline path that builds and updates the policy knowledge base, and an online path that reviews each ad. Both are Python services; the online path is exposed through an asynchronous API and stores jobs in PostgreSQL so they survive restarts.
Build the policy knowledge base
Render images, PDFs and web pages
Text blocks with coordinates
Recall-first candidate search
LLM decision per category
One-way critic
Located report and API
The numbered steps in the diagram:
- Policy examples for each risk category are enriched with definitions, category boundaries, core expressions and keywords, embedded and indexed in ChromaDB with cosine HNSW. The index is updated incrementally every day.
- Images are taken as-is, PDFs are rendered at 200 dpi with PyMuPDF, and web pages are captured full-page with Playwright; very tall images are split at gaps between content.
- Document OCR returns words with bounding boxes. Words are rebuilt into lines and blocks, nearby blocks are merged with union-find, and footnote markers are linked to the lines they qualify.
- For every sentence, dense search over the index and exact substring matching against examples produce scored candidates; the best score per example is kept, then the top categories and their evidence.
- An LLM judges each candidate category at temperature 0 with the retrieved examples and returns a decision and a reason as structured JSON.
- A separate critic audits violation flags that are not high-confidence and can only clear them — never create new ones.
- Results are merged with the keyword path into block-level JSON and an HTML overlay and returned through async, batch and sync APIs.
Technology stack
| Layer | Technology | What it does here |
|---|---|---|
| Ingestion | PyMuPDF · Playwright · image preprocessing | Turns images, PDFs and web pages into OCR-ready images |
| OCR | Managed document-text OCR · union-find merging | Text blocks with coordinates; footnotes attached to their lines |
| Knowledge base | Managed text embeddings · ChromaDB (cosine HNSW) | Per-category policy examples and enriched expressions |
| Judgment | Managed lightweight LLM · temperature 0 · JSON schema | Per-category decision with a written reason |
| Keywords | pyahocorasick · Unicode NFKC normalization | Prohibited-keyword path running in parallel |
| Service | FastAPI · PostgreSQL (JSONB) · object storage · Docker Compose | Async, batch and sync APIs with job recovery and callbacks |
| Evaluation | SciPy linear_sum_assignment · pytest | Block-level matching and micro/macro precision, recall and F1 |
Step 1: Build a per-category policy knowledge base
The client's policy tables map violating phrases to a risk category and a risk level, separately for food, cosmetics and functional cosmetics. Raw examples are too terse to retrieve well, so an offline job asks an LLM to enrich each policy: a definition, what the category covers and where it ends relative to neighbouring categories, the core expressions that signal it, and hard keywords. Generic expressions that show up in three or more policies are dropped because they discriminate nothing.
Everything is embedded and written to ChromaDB with the category, domain and risk level as metadata, so a search can be restricted to the ad's product domain.
import chromadb
client = chromadb.PersistentClient(path="kb")
kb = client.get_or_create_collection(
"policy_examples", metadata={"hnsw:space": "cosine"})
def upsert_policy(policy):
"""One policy = a risk category with examples and enriched expressions."""
items = policy.examples + policy.core_expressions
kb.upsert(
ids=[f"{policy.id}:{i}" for i in range(len(items))],
documents=[it.text for it in items],
embeddings=embed([it.text for it in items]),
metadatas=[{"category": policy.category, "domain": policy.domain,
"risk": it.risk, "kind": it.kind} for it in items],
)Simplified. embed() wraps the managed embedding model; enrichment prompts are not shown.
Updates are incremental. A policy is regenerated only when enough of its examples change; otherwise only the changed vectors are patched, and keyword lists are merged only when their fingerprint changes. That keeps a daily sync with the client's policy system cheap.
Step 2: Read the ad as located text
Review needs to point at the line, so OCR output keeps coordinates end to end. PDFs are rendered page by page; web pages are captured full-page in a headless browser. Images are upscaled and contrast-boosted before OCR, with a fallback to the original if that fails, and very tall product pages are split at whitespace gaps.
OCR returns words with boxes. The system rebuilds them into lines and paragraphs, then merges blocks that are vertically adjacent and overlap horizontally — a union-find pass over block pairs. Footnote markers such as * are linked to the body line they qualify, because a footnote can turn a violation into an acceptable claim.
def merge_blocks(blocks, gap):
"""Union-find over OCR blocks that touch vertically and overlap horizontally."""
parent = list(range(len(blocks)))
def find(i):
while parent[i] != i:
parent[i] = parent[parent[i]]
i = parent[i]
return i
for i, a in enumerate(blocks):
for j in range(i + 1, len(blocks)):
b = blocks[j]
touches = 0 <= b.top - a.bottom <= gap
overlaps = min(a.right, b.right) > max(a.left, b.left)
if touches and overlaps:
parent[find(j)] = find(i)
groups = {}
for i, blk in enumerate(blocks):
groups.setdefault(find(i), []).append(blk)
return [Block.union(g) for g in groups.values()]Simplified block merging. Real thresholds depend on font size and layout.
Step 3: Retrieve candidates for recall
Retrieval is tuned for recall: it should surface every plausible violation and leave precision to the judge. Each block and each sentence inside it is searched densely in the ad's domain, and every example that appears verbatim in the sentence gets a perfect score. Scores are merged by keeping the maximum per example — no weighted sum to tune.
score(e | s) = 1.0 if example e occurs in sentence s
= 1 − cos_dist(s, e) otherwise keep the max per example
evidence = top-3 categories by max score, ≤ 3 examples each, score ≥ τ
from collections import defaultdict
def candidates(sentence, domain, tau, top_cats=3, per_cat=3):
best = {}
res = kb.query(query_embeddings=embed([sentence]), n_results=20,
where={"domain": domain})
for ex_id, dist, meta in zip(res["ids"][0], res["distances"][0],
res["metadatas"][0]):
best[ex_id] = max(best.get(ex_id, (0.0, meta)), (1 - dist, meta),
key=lambda x: x[0])
for ex in examples_in(sentence, domain): # verbatim hits
best[ex.id] = (1.0, ex.meta)
by_cat = defaultdict(list)
for ex_id, (score, meta) in best.items():
if score >= tau:
by_cat[meta["category"]].append((score, ex_id))
ranked = sorted(by_cat.items(), key=lambda kv: -max(kv[1])[0])[:top_cats]
return {cat: sorted(hits, reverse=True)[:per_cat] for cat, hits in ranked}Simplified candidate search. examples_in() is an exact-substring lookup over the domain's examples.
Step 4: Judge with an LLM, then audit with a one-way critic
The judge sees a handful of sentences, one candidate category, that category's rules and the retrieved evidence, and must answer in a fixed JSON schema at temperature 0. Batching a few sentences per call keeps cost proportional to the ad's length; a failed batch is retried one sentence at a time. When the best evidence is a verbatim match the judge runs in a lenient mode, otherwise in a strict one.
The critic is a separate call with a separate job: look at each violation flag that is not high-confidence and ask whether its reason is supported, whether it is consistent with similar lines on the same ad that were cleared, and whether it matched only a topic rather than a claim. It can clear a flag; it can never raise one. On any error the original decision stands.
JUDGE = {"type": "object", "required": ["violation", "reason"],
"properties": {"violation": {"type": "boolean"},
"reason": {"type": "string"}}}
def judge(sentences, category, evidence):
return llm.generate(render_judge(sentences, category, evidence),
schema=JUDGE, temperature=0)
def critic(flags, siblings):
"""One-way audit: may clear a flag, never create one."""
for f in flags:
if not f.violation or f.high_confidence:
continue
try:
v = llm.generate(render_critic(f, siblings), schema=CRITIC,
temperature=0)
except Exception:
continue # keep the judge's decision
if v["clear"]:
f.violation, f.cleared_because = False, v["why"]
return flagsSimplified. Prompts are rendered from templates that are not shown; llm.generate enforces the JSON schema.
Why one-way? A second model that can add flags doubles the ways to be wrong. Restricting the critic to removals bounds its effect on recall while still cutting false alarms — the design choice behind the patent's independent verification unit.
Step 5: Run prohibited keywords in parallel
Some terms are prohibited outright by the client regardless of context. These go through two Aho-Corasick automata — one over normalized text, one with whitespace removed so spaced-out spellings still match — and a single LLM call per image removes only matches that are clearly out of context. If that call fails, every match is kept.
import unicodedata
import ahocorasick
def norm(s, compact=False):
s = unicodedata.normalize("NFKC", s).lower()
return "".join(s.split()) if compact else s
def build(keywords, compact=False):
a = ahocorasick.Automaton()
for kw in keywords:
a.add_word(norm(kw, compact), kw)
a.make_automaton()
return a
def hits(automaton, text, compact=False):
t = norm(text, compact)
return [(end - len(norm(kw, compact)) + 1, kw)
for end, kw in automaton.iter(t)]Simplified. Two automata over NFKC-normalized text, built once per keyword-list version.
Step 6: Serve it as an asynchronous API
The service accepts single ads, batches and a synchronous mode for small inputs. Each job is persisted in PostgreSQL (results as JSONB) so work resumes after a restart, inputs and overlays go to object storage, and finished jobs trigger a callback to the client. The retrieval path and the keyword path for one ad run concurrently.
import asyncio
from fastapi import FastAPI
app = FastAPI()
@app.post("/reviews")
async def review(req: ReviewRequest) -> Report:
pages = await render(req.source) # image | PDF | URL
blocks = await ocr_with_boxes(pages)
rag, kw = await asyncio.gather(
rag_path(blocks, req.domain), # steps 3-4
keyword_path(blocks), # step 5
)
return Report.merge(blocks, rag, kw) # JSON + HTML overlaySimplified endpoint; authentication, limits and job persistence are omitted.
Concurrency is capped per process, judge and critic calls run in small thread pools, embeddings are cached, and duplicate sentences across an ad are processed once. Token usage and cost are metered per job.
Step 7: Measure it the way reviewers work
Evaluation matches labelled ground-truth phrases to OCR blocks (exact, then whitespace-insensitive, then fuzzy), and then scores three things separately: was the phrase detected, was the category right, and was the risk level right. Category matching is an assignment problem — each predicted category is paired with at most one labelled category — solved with the Hungarian algorithm, and micro and macro precision, recall and F1 are reported.
import numpy as np
from scipy.optimize import linear_sum_assignment
def match_categories(pred, gold, sim):
"""Pair predictions with labels to maximize total similarity."""
reward = np.array([[sim(p, g) for g in gold] for p in pred])
rows, cols = linear_sum_assignment(reward, maximize=True)
return [(pred[r], gold[c]) for r, c in zip(rows, cols) if reward[r, c] > 0]Simplified. Assignment between predicted and labelled categories on one block.
Try the live model
The live model below reproduces the decision logic on two invented ads, so you can watch each stage change the result: a keyword filter alone, retrieval plus LLM judgment, and the independent verifier.
Retrieval (character-trigram TF-IDF over a small invented case store), keyword matching and the precision/recall at each stage are computed in your browser. The LLM judgments and the critic's notes were produced in advance for these two ads and are replayed. Open the live model on its own page ↗
Results
Measured on 499 ad images from operation, the system reached 81.14% precision and 94.96% recall. Every flag comes with its category, risk level, reason, the case it resembles and its location on the original image, which is what makes the output usable by reviewers and defensible afterwards.
Because policies live in the knowledge base rather than in model weights, a new category or a changed rule is handled by adding or updating references — no retraining.
Lessons learned
- Split recall and precision across components. Retrieval casts a wide net; the LLM decides with the category's rules in front of it. Tuning each for one job was easier than tuning one component for both.
- Keep coordinates from the first step. Location is not a reporting nicety — it lets reviewers verify a flag in seconds and lets footnotes qualify the right line.
- Constrain the verifier. A critic that can only remove flags improved precision without putting recall at risk.
- Make the knowledge base cheap to update. Incremental vector patches and fingerprinted keyword merges made a daily policy sync practical.
Conclusion
Ad review is a regulated, category-dependent judgment that neither keyword lists nor a general LLM handle well on their own. Combining located OCR, a per-category knowledge base, recall-first retrieval, a structured LLM judge and a one-way critic produced a system that explains and locates every flag and adapts to new rules without retraining.
The same pattern — retrieval for recall, an LLM for precision, an independent check before a person sees the result — applies to other compliance reviews where the rules change and the basis for each decision has to be shown.
Limitations
- Decisions are only as good as the reference store; categories with few examples are judged less reliably.
- OCR errors on stylized typography propagate to every later step.
- The live model's ads, products and cases are invented, its retrieval is a simple in-browser stand-in, and its LLM steps are replayed.
About the demo and confidentiality
All products, ad copy and cases in the embedded model are invented. No client policy text, keyword list, prompt, model or vendor configuration, threshold or production data appears in this post; code is simplified and written for illustration.