Building a self-verifying recruitment assistant with LangGraph and a self-hosted GPT-OSS model on vLLM
A walkthrough of the multi-agent assistant that reads application essays for group-wide hiring: how sentences are indexed so every claim can be traced, how summaries are verified and regenerated in a loop, how interview questions are judged for coverage and screened by policy, and how it all runs on a self-hosted open-weight model — with simplified code for each step.
CJ Group's open recruitment brings in roughly 73,000 applications across three recent cycles, and recruiters read well over 60,000 application essays a year at 15 to 20 minutes each. External recruitment-AI tools are built on general-purpose data and quantitative scoring, so they do not reflect the group's talent model, and buying separate tools for summaries, question generation and chat adds up at this volume. Hiring is also auditable: a summary that credits an experience to the wrong company, or a quote the applicant never wrote, cannot be used for a fair evaluation or defended in an appeal. A single prompt to a general LLM produces fluent text with no evidence trail and no check.
In this post we walk through the assistant we built instead. It indexes every sentence of the essay with its character offsets and runs three engines in parallel: a summarization graph that generates, verifies and regenerates; an evidence engine that returns sentence indices and rebuilds each quote from the source; and a three-stage interview-question graph with a coverage judge and a policy filter. All of them call an open-weight GPT-OSS model served by vLLM behind an OpenAI-compatible API. Verification is not one extra model call at the end — it is spread across the engines, and the recruiter stays the judge. The work is filed as a patent and moved into production for the first-half 2026 open recruitment.
Solution overview
The assistant is a batch pipeline that runs inside the group's open-recruitment process rather than a chat service. A command-line job reads the essay tables, processes applicants concurrently through three engines, and writes one JSON result per applicant, which the recruitment system shows to evaluators. Every model call goes to a self-hosted open-weight model, and nothing is retrieved from outside the applicant's own essay.
Load essays, keep Q&A aligned
Sentences with character offsets
Generate → verify → regenerate
Sentence indices, not quotes
Generate, judge coverage, filter
Self-hosted model, JSON per applicant
The numbered steps in the diagram:
- Essay tables (CSV or XLSX) are loaded with pandas, and an integrity check makes sure every answer is still paired with its question before anything is generated.
- A regex sentence splitter tuned for Korean records each sentence's character offsets and numbers sentences globally across all answers; these numbers are the only thing later steps may cite.
- A LangGraph graph extracts job keywords per answer, then writes the overall and per-question summaries; a separate verifier passes or fails each one, and failures are regenerated with the reason fed back, up to five attempts.
- One call per talent-model competency, run in parallel with asyncio, returns the indices of supporting sentences; the quote shown to the recruiter is cut from the essay by offset, never taken from the model.
- A second LangGraph graph writes question sets across eight behavioural areas, a coverage judge decides whether the essay already answers each one and must cite 2–4 sentence indices, and a policy filter deletes or rewrites follow-ups on prohibited topics.
- Every engine calls an open-weight GPT-OSS model served by vLLM through its OpenAI-compatible API; outputs pass through LangChain's JsonOutputParser and a repair chain, and each applicant's result is written as one JSON file.
Technology stack
| Layer | Technology | What it does here |
|---|---|---|
| Ingestion | pandas · integrity checks | Essay tables loaded with questions and answers kept aligned |
| Grounding | Regex sentence splitter with character offsets | Global sentence index; evidence returned as positions |
| Orchestration | LangGraph StateGraph · conditional edges | Verify-and-regenerate loop; staged question pipeline |
| Model serving | vLLM · OpenAI-compatible API · openai SDK · GPT-OSS | Self-hosted inference; applicant data stays in-house |
| Structured output | LangChain JsonOutputParser · repair chain · retries with backoff | Valid JSON from every call at batch scale |
| Concurrency | asyncio (applicants × engines × calls) | Keeps enough requests in flight for the model server |
| Evaluation | QA-FactEval · recruiter pilot | Factual consistency against commercial and open-weight baselines |
Step 1: Load the essays and index every sentence
Essays arrive as tables: one row per applicant, with the role or division applied for and the answers to several essay questions. They are read with pandas, and an integrity check confirms that every answer is still paired with the right question. A shifted column would not crash anything — it would produce confident summaries of the wrong answer — so the check runs before any model call.
Each answer is then split into sentences by a regular-expression splitter tuned for Korean, which records the start and end character offset of every sentence. Sentences are numbered globally across all answers. From here on, the evidence engine and the question engine refer to the essay only by these numbers.
import re
from dataclasses import dataclass
# Simplified: the production splitter is tuned for Korean sentence endings.
SENTENCE = re.compile(r"[^.!?\n]+(?:[.!?]+|$)", re.M)
@dataclass(frozen=True)
class Sentence:
idx: int # global number across all answers, starting at 1
answer: int # which essay question the sentence belongs to
start: int # character offsets inside that answer
end: int
def index_sentences(answers: list[str]) -> list[Sentence]:
out = []
for a, text in enumerate(answers):
for m in SENTENCE.finditer(text):
s, e = m.start(), m.end()
while s < e and text[s].isspace():
s += 1
if s < e:
out.append(Sentence(len(out) + 1, a, s, e))
return out
def quote(answers: list[str], s: Sentence) -> str:
return answers[s.answer][s.start:s.end].strip()
def numbered(answers: list[str], sentences: list[Sentence]) -> str:
return "\n".join(f"[{s.idx}] {quote(answers, s)}" for s in sentences)Simplified. The pattern is a generic stand-in for the production splitter; offsets are what matter.
Why no retrieval? Everything a recruiter needs to check is in the applicant's own essay. Retrieving other documents would add a second source of claims the applicant never made. The essay is passed whole, there is no vector store, and every statement can be traced back to a sentence number.
Step 2: Serve an open-weight model behind an OpenAI-compatible API
The core models run on an open-weight GPT-OSS reasoning model that we host ourselves instead of calling an external commercial API. Three reasons drove the choice: application essays are personal data and stay on infrastructure the group controls; per-call vendor cost would keep accumulating at tens of thousands of essays per cycle; and owning the model made it possible to benchmark it against commercial and other open-weight models on the same essays (see Results).
The model is served with vLLM, which exposes an OpenAI-compatible endpoint, so the pipeline uses the standard openai SDK and LangChain parsers and the engines do not depend on which model sits behind the endpoint. A model of this class needs a large-memory GPU or tensor parallelism across several; vLLM supports both and schedules concurrent requests together, which is why the pipeline is built to keep many requests in flight (Step 7).
Temperature is set per stage. Generation stages run warmer, so a regenerated summary can differ from the one that failed; judging stages run cold, so the same input tends to get the same verdict.
import asyncio
from openai import AsyncOpenAI, APIConnectionError, APITimeoutError
# vLLM serves an OpenAI-compatible API, so the standard SDK works unchanged.
client = AsyncOpenAI(base_url=LLM_BASE_URL, api_key=LLM_API_KEY)
# Generation runs warmer than judging; the values are not shown.
TEMPERATURE = {"keywords": T_GEN, "summary": T_GEN, "questions": T_GEN,
"verify": T_JUDGE, "evidence": T_JUDGE, "coverage": T_JUDGE,
"policy": T_JUDGE, "repair": T_JUDGE}
async def chat(stage: str, prompt: str, retries: int = 3) -> str:
for attempt in range(retries):
try:
resp = await client.chat.completions.create(
model=MODEL_NAME,
messages=[{"role": "user", "content": prompt}],
temperature=TEMPERATURE[stage],
)
return resp.choices[0].message.content
except (APIConnectionError, APITimeoutError):
if attempt == retries - 1:
raise
await asyncio.sleep(2 ** attempt) # exponential backoff
async def chat_json(stage: str, prompt: str) -> dict:
return await parse_json(await chat(stage, prompt)) # recovery chain, step 6Simplified. Endpoint, model name and temperatures are placeholders; the stage-to-temperature mapping is illustrative.
Step 3: Summarize with a generate → verify → regenerate loop
The summarization engine first extracts job-core keywords from each answer, in parallel, and then fans out to an overall summary and one summary per essay question, written in a recruitment-analyst persona. Every summary goes through the same loop: a generator writes a draft, a separate verifier call reads the draft against the essay and returns PASS or FAIL with a reason, and a failed draft is regenerated with that reason in the prompt.
for attempt t = 1 … 5
s_t ← LLM_gen(essay, persona, feedback_(t−1))
(v_t, r_t) ← LLM_verify(essay, s_t) v ∈ {PASS, FAIL}, r = reason
accept s_t if v_t = PASS and |s_t| ≤ 1.1 · L
feedback_t ← r_t otherwise
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
MAX_ATTEMPTS, LENGTH_SLACK = 5, 1.1
class SummaryState(TypedDict):
essay: str
limit: int # target length in characters
draft: str
passed: bool
feedback: str # the verifier's reason, fed into the next attempt
attempts: int
async def generate(s: SummaryState) -> dict:
prompt = render_prompt("summary", essay=s["essay"], feedback=s["feedback"])
return {"draft": await chat("summary", prompt), "attempts": s["attempts"] + 1}
async def verify(s: SummaryState) -> dict:
v = await chat_json("verify", render_prompt("verify", essay=s["essay"], draft=s["draft"]))
too_long = len(s["draft"]) > LENGTH_SLACK * s["limit"]
reason = "over the length limit" if too_long else v.get("reason", "")
return {"passed": v["verdict"] == "PASS" and not too_long, "feedback": reason}
def next_step(s: SummaryState) -> str:
return "done" if s["passed"] or s["attempts"] >= MAX_ATTEMPTS else "retry"
g = StateGraph(SummaryState)
g.add_node("generate", generate); g.add_node("verify", verify)
g.add_edge(START, "generate"); g.add_edge("generate", "verify")
g.add_conditional_edges("verify", next_step, {"retry": "generate", "done": END})
summarize = g.compile()Simplified. render_prompt() fills templates that are not shown; keyword extraction and the fan-out to per-question summaries are omitted.
If five attempts do not produce a passing summary, a deterministic fallback decides what is written — truncate an over-long draft, keep the last attempt, or return nothing — instead of looping indefinitely. Each accepted summary is finally split into exactly two sentences for display.
Why a separate verifier? A model asked to check its own output in the same call tends to agree with itself. A separate call with its own instructions and a low temperature behaves more like an independent reviewer, and its reason gives the next attempt something specific to fix.
Step 4: Return evidence as sentence indices, then rebuild the quote
The evidence engine maps each competency in the talent model to the sentences that support it. It makes one call per competency, all in parallel. Each call sees the numbered essay and that competency's definition and returns sentence indices. If the model also writes out a quote, the text is discarded: the quote the recruiter sees is cut from the original essay using the stored offsets.
Integers are easy to check — an index either exists in this essay or it does not — and out-of-range or duplicate indices are dropped in code. A paraphrased quote is much harder to catch, and in an auditable process a near-quote is worse than none.
import asyncio
async def evidence_for(factor, answers, sentences):
prompt = render_prompt("evidence", factor=factor, essay=numbered(answers, sentences))
out = await chat_json("evidence", prompt)
by_idx = {s.idx: s for s in sentences}
picked = []
for i in out.get("indices", []): # any quote the model writes is ignored
if isinstance(i, int) and i in by_idx and i not in picked:
picked.append(i)
return [{"sentence": i,
"answer": by_idx[i].answer,
"span": [by_idx[i].start, by_idx[i].end],
"quote": quote(answers, by_idx[i])} # cut from the essay itself
for i in picked]
async def highlight(factors, answers, sentences):
"""One call per competency, all in parallel."""
found = await asyncio.gather(*(evidence_for(f, answers, sentences) for f in factors))
return {f.id: ev for f, ev in zip(factors, found)}Simplified. Competency definitions and prompts are not shown.
Step 5: Generate interview questions, judge coverage, filter by policy
The question engine is a three-stage LangGraph pipeline. Generation writes question sets across eight behavioural areas, covering both what the essay answers and what it leaves open. Each set has a main question and probing questions, each tied to the sentence indices it was built from. The coverage judge then decides, for every main question, whether the essay already answers it. If it does, the judge writes the answer it expects in the applicant's own voice and must cite two to four sentence indices as evidence; it also proposes follow-ups.
Code, not the model, has the last word on consistency: deterministic rules keep main and probing questions consistent with each other, and every evidence index is range-checked. The policy filter finally checks each follow-up against the categories of questions an interviewer must not ask, and deletes or rewrites any that match.
import asyncio
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
class QState(TypedDict):
essay: str # numbered sentences from step 1
n: int # number of sentences
sets: list # main question + probing questions + source indices
async def generate_sets(s: QState) -> dict:
out = await chat_json("questions", render_prompt("questions", essay=s["essay"]))
return {"sets": out["sets"]}
async def judge_one(qs: dict, s: QState) -> dict:
v = await chat_json("coverage", render_prompt("coverage", question=qs["main"], essay=s["essay"]))
ev = [i for i in v.get("evidence", []) if isinstance(i, int) and 1 <= i <= s["n"]][:4]
answered = v.get("answered") == "Y" and len(ev) >= 2 # example of a code-side rule
return {**qs, "answered": answered, "evidence": ev, "followups": v.get("followups", []),
"expected": v.get("expected_answer") if answered else None}
async def judge_coverage(s: QState) -> dict:
return {"sets": list(await asyncio.gather(*(judge_one(q, s) for q in s["sets"])))}
async def policy_filter(s: QState) -> dict:
return {"sets": [{**q, "followups": await screen_followups(q["followups"])} for q in s["sets"]]}
g = StateGraph(QState)
g.add_sequence([("generate", generate_sets), ("coverage", judge_coverage), ("policy", policy_filter)])
g.add_edge(START, "generate"); g.add_edge("policy", END)
question_graph = g.compile()Simplified. screen_followups() deletes or rewrites follow-ups that touch a prohibited category; the override rule shown is an example of a code-side check, not the production rule set.
Taken together, the engines place each kind of check where that error can be caught most directly:
| Error | Caught by | How |
|---|---|---|
| Off tone, or a key experience left out of a summary | Summary verifier (Step 3) | PASS/FAIL with a reason; regenerate |
| A claim the essay does not support | Coverage judge (Step 5) | Y/N with 2–4 cited sentence indices, checked in code |
| An altered or invented quote | Evidence engine (Step 4) | Model text discarded; quote rebuilt by offset |
| A question on a prohibited topic | Policy filter (Step 5) | Follow-up deleted or rewritten |
| Malformed output | JSON recovery chain (Step 6) | Cheap fixes first, then one repair call |
Step 6: Recover structured output at batch scale
Every engine expects JSON, and one applicant needs many calls. At that volume even a small rate of malformed output becomes a steady stream of failures, so parsing is a chain that tries cheap fixes before spending another model call: parse with LangChain's JsonOutputParser; strip markdown code fences; extract the outermost braces; drop trailing commas; and only then ask the model once to repair its own output. Transport errors are retried with exponential backoff (Step 2).
import json
import re
from langchain_core.exceptions import OutputParserException
from langchain_core.output_parsers import JsonOutputParser
parser = JsonOutputParser()
def cheap_fixes(text: str):
yield text
text = re.sub(r"^```(?:json)?\s*|\s*```$", "", text.strip()) # code fences
yield text
a, b = text.find("{"), text.rfind("}")
if a != -1 and b > a:
text = text[a:b + 1] # outermost braces
yield text
yield re.sub(r",\s*([}\]])", r"\1", text) # trailing commas
async def parse_json(raw: str) -> dict:
for candidate in cheap_fixes(raw):
try:
return parser.parse(candidate)
except OutputParserException:
continue
fixed = await chat("repair", render_prompt("repair", broken=raw)) # one repair call
return json.loads(fixed)Simplified. Cheap fixes first, one model repair last; the repair prompt is not shown.
Step 7: Run applicants, engines and calls concurrently
Concurrency is nested three levels deep. Applicants are processed in parallel; for each applicant the three engines run in parallel; and inside the engines the per-answer keyword calls, the per-competency evidence calls and the per-question coverage calls run in parallel too. A self-hosted server is a fixed resource whose throughput comes from serving many requests at once, so keeping enough requests in flight matters more than the speed of any single call.
import asyncio
import json
from pathlib import Path
async def process(applicant, profile, gate: asyncio.Semaphore) -> dict:
async with gate: # bound applicants in flight
answers = applicant.answers
sentences = index_sentences(answers) # step 1
summary, evidence, questions = await asyncio.gather(
summarize_all(answers, applicant.role), # step 3: keywords, then summaries
highlight(profile.factors, answers, sentences), # step 4
question_graph.ainvoke({"essay": numbered(answers, sentences),
"n": len(sentences), "sets": []}), # step 5
)
return {"id": applicant.id, "summary": summary,
"evidence": evidence, "questions": questions["sets"]}
async def run_batch(applicants, profile, out_dir: Path, max_parallel: int):
gate = asyncio.Semaphore(max_parallel)
results = await asyncio.gather(*(process(a, profile, gate) for a in applicants))
for r in results:
(out_dir / f"{r['id']}.json").write_text(
json.dumps(r, ensure_ascii=False, indent=2), encoding="utf-8")Simplified. summarize_all() wraps the keyword step and the summary graph of Step 3; the semaphore is an illustrative bound on applicants in flight.
The batch job writes one JSON file per applicant — summaries, keywords, evidence with offsets, and question sets with their coverage verdicts and cited indices. The recruitment system shows these to evaluators, who make every decision themselves; there is no automatic score.
Try the live model
The live model below replays the three engines on one invented essay, so you can watch a flawed first summary get caught, see evidence highlighted by sentence number, and see a follow-up removed by the policy filter.
Live model on an invented essay. The agent outputs — including a deliberately flawed first summary — were produced in advance and are replayed; the support check (does each summary line's wording appear in the essay, and do its numbers match) runs live in your browser as a simple stand-in for the production verifiers. Switch between the first draft and the verified version, and pick a competency to see its evidence. Open the live model on its own page ↗
Results
On QA-FactEval — a question-answering-based measure of whether a summary's statements are supported by its source — the internalized summarization model scored 83.45% on 380 application essays, against 75.95% for a commercial small model and 68.69% for an open-weight baseline, with stable reproducibility across repeated runs (standard deviation 0.09). Job-core keyword extraction reached 85.36% versus 76.79% for an external model. On the issue set, the verification loops improved 21.6% of whole-summary cases, 18.7% of per-question cases and 37.3% of interview-question cases.
In a pilot with 31 recruiters and the open-recruitment task force on 490 samples, positive-plus-neutral responses exceeded 94%; across 21,142 evaluator responses, 87.14% rated the outputs helpful. First-pass review time fell by about half and interview preparation by about two-thirds. The work is filed as a patent — a recruitment document reading assistant using evidence-candidate generation and multi-agent verification — and moved into production for the first-half 2026 open recruitment.
Lessons learned
- Return positions, not text. Asking for sentence indices instead of quotes turned the hardest hallucination to spot — a slightly altered quote — into an integer that code can check.
- Give the verifier its own call. A separate verifier with its own instructions and a low temperature flagged problems the generator would not flag in itself, and its reason made each retry targeted rather than random.
- Spread verification across the pipeline. Tone, unsupported claims, quotes and policy each have a check where that error can be caught most directly, instead of one catch-all reviewer at the end.
- Keep deterministic rules in code. Range checks and consistency rules are cheaper and more reliable in code than in a prompt, and they behave the same on every run.
- Self-hosting is a throughput problem. On a fixed model server, nested concurrency and a robust JSON chain mattered as much as the choice of model.
Conclusion
Reading application essays at group scale is a high-stakes, auditable task that a single LLM prompt handles poorly: it produces fluent text without evidence or checks. Indexing the essay by sentence, verifying and regenerating summaries, returning evidence as positions, judging question coverage against cited sentences and filtering by policy produced an assistant whose outputs can be traced to the applicant's own words — on a model the group hosts itself.
The same pattern — make the source addressable, let models point rather than quote, and place a specific check after each generator — applies to other document-reading tasks where the output has to be defended later.
Limitations
- QA-FactEval measures the factual consistency of summaries, not the quality of a hiring decision.
- Grounding checks catch claims the essay does not support; they cannot tell whether the essay itself is truthful.
- The verifiers are themselves LLM calls and can pass a flawed output or fail a good one; attempts are capped and the recruiter remains the judge.
- The live model's support check is a simple word-coverage and number test standing in for the production verifiers, and its agent outputs are pre-written for one invented essay and replayed.
About the demo and confidentiality
The applicant, essay, competencies, questions and outputs in the embedded model are invented. No applicant data, talent-model wording or competency definitions, prohibited-question categories, interview check items, prompts, question sets or deployment configuration from the production system appear in this post; code is simplified and written for illustration.