GenAI · Multi-agent · Self-hosted LLMCJ AI Center, with CJ Group HRTechnical write-up · October 2026 · 13 min read

Building a self-verifying recruitment assistant with LangGraph and a self-hosted GPT-OSS model on vLLM

A walkthrough of the multi-agent assistant that reads application essays for group-wide hiring: how sentences are indexed so every claim can be traced, how summaries are verified and regenerated in a loop, how interview questions are judged for coverage and screened by policy, and how it all runs on a self-hosted open-weight model — with simplified code for each step.

Built withPythonLangGraphLangChainvLLMGPT-OSS · self-hostedOpenAI-compatible APIasynciopandasRegex sentence index

CJ Group's open recruitment brings in roughly 73,000 applications across three recent cycles, and recruiters read well over 60,000 application essays a year at 15 to 20 minutes each. External recruitment-AI tools are built on general-purpose data and quantitative scoring, so they do not reflect the group's talent model, and buying separate tools for summaries, question generation and chat adds up at this volume. Hiring is also auditable: a summary that credits an experience to the wrong company, or a quote the applicant never wrote, cannot be used for a fair evaluation or defended in an appeal. A single prompt to a general LLM produces fluent text with no evidence trail and no check.

In this post we walk through the assistant we built instead. It indexes every sentence of the essay with its character offsets and runs three engines in parallel: a summarization graph that generates, verifies and regenerates; an evidence engine that returns sentence indices and rebuilds each quote from the source; and a three-stage interview-question graph with a coverage judge and a policy filter. All of them call an open-weight GPT-OSS model served by vLLM behind an OpenAI-compatible API. Verification is not one extra model call at the end — it is spread across the engines, and the recruiter stays the judge. The work is filed as a patent and moved into production for the first-half 2026 open recruitment.

Solution overview

The assistant is a batch pipeline that runs inside the group's open-recruitment process rather than a chat service. A command-line job reads the essay tables, processes applicants concurrently through three engines, and writes one JSON result per applicant, which the recruitment system shows to evaluators. Every model call goes to a self-hosted open-weight model, and nothing is retrieved from outside the applicant's own essay.

Architecture
1Ingest

Load essays, keep Q&A aligned

pandas
2Index

Sentences with character offsets

regex splitter
3Summarize

Generate → verify → regenerate

LangGraphLLM verifier
4Evidence

Sentence indices, not quotes

asyncioLLM
5Questions

Generate, judge coverage, filter

LangGraphLLM judge
6Serve

Self-hosted model, JSON per applicant

vLLMGPT-OSSJsonOutputParser
Essays are loaded and every sentence is numbered; three engines then work on the same essay in parallel, each with its own check, and all of them call one self-hosted model. Steps 3–5 run concurrently for each applicant, and applicants run concurrently with each other.

The numbered steps in the diagram:

  1. Essay tables (CSV or XLSX) are loaded with pandas, and an integrity check makes sure every answer is still paired with its question before anything is generated.
  2. A regex sentence splitter tuned for Korean records each sentence's character offsets and numbers sentences globally across all answers; these numbers are the only thing later steps may cite.
  3. A LangGraph graph extracts job keywords per answer, then writes the overall and per-question summaries; a separate verifier passes or fails each one, and failures are regenerated with the reason fed back, up to five attempts.
  4. One call per talent-model competency, run in parallel with asyncio, returns the indices of supporting sentences; the quote shown to the recruiter is cut from the essay by offset, never taken from the model.
  5. A second LangGraph graph writes question sets across eight behavioural areas, a coverage judge decides whether the essay already answers each one and must cite 2–4 sentence indices, and a policy filter deletes or rewrites follow-ups on prohibited topics.
  6. Every engine calls an open-weight GPT-OSS model served by vLLM through its OpenAI-compatible API; outputs pass through LangChain's JsonOutputParser and a repair chain, and each applicant's result is written as one JSON file.

Technology stack

LayerTechnologyWhat it does here
Ingestionpandas · integrity checksEssay tables loaded with questions and answers kept aligned
GroundingRegex sentence splitter with character offsetsGlobal sentence index; evidence returned as positions
OrchestrationLangGraph StateGraph · conditional edgesVerify-and-regenerate loop; staged question pipeline
Model servingvLLM · OpenAI-compatible API · openai SDK · GPT-OSSSelf-hosted inference; applicant data stays in-house
Structured outputLangChain JsonOutputParser · repair chain · retries with backoffValid JSON from every call at batch scale
Concurrencyasyncio (applicants × engines × calls)Keeps enough requests in flight for the model server
EvaluationQA-FactEval · recruiter pilotFactual consistency against commercial and open-weight baselines

Step 1: Load the essays and index every sentence

Essays arrive as tables: one row per applicant, with the role or division applied for and the answers to several essay questions. They are read with pandas, and an integrity check confirms that every answer is still paired with the right question. A shifted column would not crash anything — it would produce confident summaries of the wrong answer — so the check runs before any model call.

Each answer is then split into sentences by a regular-expression splitter tuned for Korean, which records the start and end character offset of every sentence. Sentences are numbered globally across all answers. From here on, the evidence engine and the question engine refer to the essay only by these numbers.

grounding/sentences.py
import re
from dataclasses import dataclass

# Simplified: the production splitter is tuned for Korean sentence endings.
SENTENCE = re.compile(r"[^.!?\n]+(?:[.!?]+|$)", re.M)

@dataclass(frozen=True)
class Sentence:
    idx: int        # global number across all answers, starting at 1
    answer: int     # which essay question the sentence belongs to
    start: int      # character offsets inside that answer
    end: int

def index_sentences(answers: list[str]) -> list[Sentence]:
    out = []
    for a, text in enumerate(answers):
        for m in SENTENCE.finditer(text):
            s, e = m.start(), m.end()
            while s < e and text[s].isspace():
                s += 1
            if s < e:
                out.append(Sentence(len(out) + 1, a, s, e))
    return out

def quote(answers: list[str], s: Sentence) -> str:
    return answers[s.answer][s.start:s.end].strip()

def numbered(answers: list[str], sentences: list[Sentence]) -> str:
    return "\n".join(f"[{s.idx}] {quote(answers, s)}" for s in sentences)

Simplified. The pattern is a generic stand-in for the production splitter; offsets are what matter.

Why no retrieval? Everything a recruiter needs to check is in the applicant's own essay. Retrieving other documents would add a second source of claims the applicant never made. The essay is passed whole, there is no vector store, and every statement can be traced back to a sentence number.

Step 2: Serve an open-weight model behind an OpenAI-compatible API

The core models run on an open-weight GPT-OSS reasoning model that we host ourselves instead of calling an external commercial API. Three reasons drove the choice: application essays are personal data and stay on infrastructure the group controls; per-call vendor cost would keep accumulating at tens of thousands of essays per cycle; and owning the model made it possible to benchmark it against commercial and other open-weight models on the same essays (see Results).

The model is served with vLLM, which exposes an OpenAI-compatible endpoint, so the pipeline uses the standard openai SDK and LangChain parsers and the engines do not depend on which model sits behind the endpoint. A model of this class needs a large-memory GPU or tensor parallelism across several; vLLM supports both and schedules concurrent requests together, which is why the pipeline is built to keep many requests in flight (Step 7).

Temperature is set per stage. Generation stages run warmer, so a regenerated summary can differ from the one that failed; judging stages run cold, so the same input tends to get the same verdict.

llm/client.py
import asyncio
from openai import AsyncOpenAI, APIConnectionError, APITimeoutError

# vLLM serves an OpenAI-compatible API, so the standard SDK works unchanged.
client = AsyncOpenAI(base_url=LLM_BASE_URL, api_key=LLM_API_KEY)

# Generation runs warmer than judging; the values are not shown.
TEMPERATURE = {"keywords": T_GEN, "summary": T_GEN, "questions": T_GEN,
               "verify": T_JUDGE, "evidence": T_JUDGE, "coverage": T_JUDGE,
               "policy": T_JUDGE, "repair": T_JUDGE}

async def chat(stage: str, prompt: str, retries: int = 3) -> str:
    for attempt in range(retries):
        try:
            resp = await client.chat.completions.create(
                model=MODEL_NAME,
                messages=[{"role": "user", "content": prompt}],
                temperature=TEMPERATURE[stage],
            )
            return resp.choices[0].message.content
        except (APIConnectionError, APITimeoutError):
            if attempt == retries - 1:
                raise
            await asyncio.sleep(2 ** attempt)          # exponential backoff

async def chat_json(stage: str, prompt: str) -> dict:
    return await parse_json(await chat(stage, prompt))   # recovery chain, step 6

Simplified. Endpoint, model name and temperatures are placeholders; the stage-to-temperature mapping is illustrative.

Step 3: Summarize with a generate → verify → regenerate loop

The summarization engine first extracts job-core keywords from each answer, in parallel, and then fans out to an overall summary and one summary per essay question, written in a recruitment-analyst persona. Every summary goes through the same loop: a generator writes a draft, a separate verifier call reads the draft against the essay and returns PASS or FAIL with a reason, and a failed draft is regenerated with that reason in the prompt.

for attempt t = 1 … 5
    s_t        ← LLM_gen(essay, persona, feedback_(t−1))
    (v_t, r_t) ← LLM_verify(essay, s_t)             v ∈ {PASS, FAIL},  r = reason
    accept s_t   if v_t = PASS  and  |s_t| ≤ 1.1 · L
    feedback_t ← r_t                               otherwise
summary/graph.py
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
MAX_ATTEMPTS, LENGTH_SLACK = 5, 1.1

class SummaryState(TypedDict):
    essay: str
    limit: int          # target length in characters
    draft: str
    passed: bool
    feedback: str       # the verifier's reason, fed into the next attempt
    attempts: int

async def generate(s: SummaryState) -> dict:
    prompt = render_prompt("summary", essay=s["essay"], feedback=s["feedback"])
    return {"draft": await chat("summary", prompt), "attempts": s["attempts"] + 1}

async def verify(s: SummaryState) -> dict:
    v = await chat_json("verify", render_prompt("verify", essay=s["essay"], draft=s["draft"]))
    too_long = len(s["draft"]) > LENGTH_SLACK * s["limit"]
    reason = "over the length limit" if too_long else v.get("reason", "")
    return {"passed": v["verdict"] == "PASS" and not too_long, "feedback": reason}

def next_step(s: SummaryState) -> str:
    return "done" if s["passed"] or s["attempts"] >= MAX_ATTEMPTS else "retry"

g = StateGraph(SummaryState)
g.add_node("generate", generate); g.add_node("verify", verify)
g.add_edge(START, "generate"); g.add_edge("generate", "verify")
g.add_conditional_edges("verify", next_step, {"retry": "generate", "done": END})
summarize = g.compile()

Simplified. render_prompt() fills templates that are not shown; keyword extraction and the fan-out to per-question summaries are omitted.

If five attempts do not produce a passing summary, a deterministic fallback decides what is written — truncate an over-long draft, keep the last attempt, or return nothing — instead of looping indefinitely. Each accepted summary is finally split into exactly two sentences for display.

Why a separate verifier? A model asked to check its own output in the same call tends to agree with itself. A separate call with its own instructions and a low temperature behaves more like an independent reviewer, and its reason gives the next attempt something specific to fix.

Step 4: Return evidence as sentence indices, then rebuild the quote

The evidence engine maps each competency in the talent model to the sentences that support it. It makes one call per competency, all in parallel. Each call sees the numbered essay and that competency's definition and returns sentence indices. If the model also writes out a quote, the text is discarded: the quote the recruiter sees is cut from the original essay using the stored offsets.

Integers are easy to check — an index either exists in this essay or it does not — and out-of-range or duplicate indices are dropped in code. A paraphrased quote is much harder to catch, and in an auditable process a near-quote is worse than none.

evidence/highlight.py
import asyncio

async def evidence_for(factor, answers, sentences):
    prompt = render_prompt("evidence", factor=factor, essay=numbered(answers, sentences))
    out = await chat_json("evidence", prompt)
    by_idx = {s.idx: s for s in sentences}
    picked = []
    for i in out.get("indices", []):          # any quote the model writes is ignored
        if isinstance(i, int) and i in by_idx and i not in picked:
            picked.append(i)
    return [{"sentence": i,
             "answer": by_idx[i].answer,
             "span": [by_idx[i].start, by_idx[i].end],
             "quote": quote(answers, by_idx[i])}      # cut from the essay itself
            for i in picked]

async def highlight(factors, answers, sentences):
    """One call per competency, all in parallel."""
    found = await asyncio.gather(*(evidence_for(f, answers, sentences) for f in factors))
    return {f.id: ev for f, ev in zip(factors, found)}

Simplified. Competency definitions and prompts are not shown.

Step 5: Generate interview questions, judge coverage, filter by policy

The question engine is a three-stage LangGraph pipeline. Generation writes question sets across eight behavioural areas, covering both what the essay answers and what it leaves open. Each set has a main question and probing questions, each tied to the sentence indices it was built from. The coverage judge then decides, for every main question, whether the essay already answers it. If it does, the judge writes the answer it expects in the applicant's own voice and must cite two to four sentence indices as evidence; it also proposes follow-ups.

Code, not the model, has the last word on consistency: deterministic rules keep main and probing questions consistent with each other, and every evidence index is range-checked. The policy filter finally checks each follow-up against the categories of questions an interviewer must not ask, and deletes or rewrites any that match.

questions/graph.py
import asyncio
from typing import TypedDict
from langgraph.graph import StateGraph, START, END

class QState(TypedDict):
    essay: str          # numbered sentences from step 1
    n: int              # number of sentences
    sets: list          # main question + probing questions + source indices

async def generate_sets(s: QState) -> dict:
    out = await chat_json("questions", render_prompt("questions", essay=s["essay"]))
    return {"sets": out["sets"]}

async def judge_one(qs: dict, s: QState) -> dict:
    v = await chat_json("coverage", render_prompt("coverage", question=qs["main"], essay=s["essay"]))
    ev = [i for i in v.get("evidence", []) if isinstance(i, int) and 1 <= i <= s["n"]][:4]
    answered = v.get("answered") == "Y" and len(ev) >= 2      # example of a code-side rule
    return {**qs, "answered": answered, "evidence": ev, "followups": v.get("followups", []),
            "expected": v.get("expected_answer") if answered else None}

async def judge_coverage(s: QState) -> dict:
    return {"sets": list(await asyncio.gather(*(judge_one(q, s) for q in s["sets"])))}

async def policy_filter(s: QState) -> dict:
    return {"sets": [{**q, "followups": await screen_followups(q["followups"])} for q in s["sets"]]}

g = StateGraph(QState)
g.add_sequence([("generate", generate_sets), ("coverage", judge_coverage), ("policy", policy_filter)])
g.add_edge(START, "generate"); g.add_edge("policy", END)
question_graph = g.compile()

Simplified. screen_followups() deletes or rewrites follow-ups that touch a prohibited category; the override rule shown is an example of a code-side check, not the production rule set.

Taken together, the engines place each kind of check where that error can be caught most directly:

ErrorCaught byHow
Off tone, or a key experience left out of a summarySummary verifier (Step 3)PASS/FAIL with a reason; regenerate
A claim the essay does not supportCoverage judge (Step 5)Y/N with 2–4 cited sentence indices, checked in code
An altered or invented quoteEvidence engine (Step 4)Model text discarded; quote rebuilt by offset
A question on a prohibited topicPolicy filter (Step 5)Follow-up deleted or rewritten
Malformed outputJSON recovery chain (Step 6)Cheap fixes first, then one repair call

Step 6: Recover structured output at batch scale

Every engine expects JSON, and one applicant needs many calls. At that volume even a small rate of malformed output becomes a steady stream of failures, so parsing is a chain that tries cheap fixes before spending another model call: parse with LangChain's JsonOutputParser; strip markdown code fences; extract the outermost braces; drop trailing commas; and only then ask the model once to repair its own output. Transport errors are retried with exponential backoff (Step 2).

llm/json_repair.py
import json
import re
from langchain_core.exceptions import OutputParserException
from langchain_core.output_parsers import JsonOutputParser

parser = JsonOutputParser()

def cheap_fixes(text: str):
    yield text
    text = re.sub(r"^```(?:json)?\s*|\s*```$", "", text.strip())    # code fences
    yield text
    a, b = text.find("{"), text.rfind("}")
    if a != -1 and b > a:
        text = text[a:b + 1]                                         # outermost braces
        yield text
    yield re.sub(r",\s*([}\]])", r"\1", text)                        # trailing commas

async def parse_json(raw: str) -> dict:
    for candidate in cheap_fixes(raw):
        try:
            return parser.parse(candidate)
        except OutputParserException:
            continue
    fixed = await chat("repair", render_prompt("repair", broken=raw))  # one repair call
    return json.loads(fixed)

Simplified. Cheap fixes first, one model repair last; the repair prompt is not shown.

Step 7: Run applicants, engines and calls concurrently

Concurrency is nested three levels deep. Applicants are processed in parallel; for each applicant the three engines run in parallel; and inside the engines the per-answer keyword calls, the per-competency evidence calls and the per-question coverage calls run in parallel too. A self-hosted server is a fixed resource whose throughput comes from serving many requests at once, so keeping enough requests in flight matters more than the speed of any single call.

pipeline/run.py
import asyncio
import json
from pathlib import Path

async def process(applicant, profile, gate: asyncio.Semaphore) -> dict:
    async with gate:                                      # bound applicants in flight
        answers = applicant.answers
        sentences = index_sentences(answers)              # step 1
        summary, evidence, questions = await asyncio.gather(
            summarize_all(answers, applicant.role),       # step 3: keywords, then summaries
            highlight(profile.factors, answers, sentences),   # step 4
            question_graph.ainvoke({"essay": numbered(answers, sentences),
                                    "n": len(sentences), "sets": []}),   # step 5
        )
    return {"id": applicant.id, "summary": summary,
            "evidence": evidence, "questions": questions["sets"]}

async def run_batch(applicants, profile, out_dir: Path, max_parallel: int):
    gate = asyncio.Semaphore(max_parallel)
    results = await asyncio.gather(*(process(a, profile, gate) for a in applicants))
    for r in results:
        (out_dir / f"{r['id']}.json").write_text(
            json.dumps(r, ensure_ascii=False, indent=2), encoding="utf-8")

Simplified. summarize_all() wraps the keyword step and the summary graph of Step 3; the semaphore is an illustrative bound on applicants in flight.

The batch job writes one JSON file per applicant — summaries, keywords, evidence with offsets, and question sets with their coverage verdicts and cited indices. The recruitment system shows these to evaluators, who make every decision themselves; there is no automatic score.

Try the live model

The live model below replays the three engines on one invented essay, so you can watch a flawed first summary get caught, see evidence highlighted by sentence number, and see a follow-up removed by the policy filter.

Live model on an invented essay. The agent outputs — including a deliberately flawed first summary — were produced in advance and are replayed; the support check (does each summary line's wording appear in the essay, and do its numbers match) runs live in your browser as a simple stand-in for the production verifiers. Switch between the first draft and the verified version, and pick a competency to see its evidence. Open the live model on its own page ↗

Results

83.45%QA-FactEval of the internalized summarizer on 380 essays (75.95% commercial small model, 68.69% open-weight)
−50%first-pass essay review time, 15 min → under 7
−66%interview preparation time, 30 min → 10

On QA-FactEval — a question-answering-based measure of whether a summary's statements are supported by its source — the internalized summarization model scored 83.45% on 380 application essays, against 75.95% for a commercial small model and 68.69% for an open-weight baseline, with stable reproducibility across repeated runs (standard deviation 0.09). Job-core keyword extraction reached 85.36% versus 76.79% for an external model. On the issue set, the verification loops improved 21.6% of whole-summary cases, 18.7% of per-question cases and 37.3% of interview-question cases.

In a pilot with 31 recruiters and the open-recruitment task force on 490 samples, positive-plus-neutral responses exceeded 94%; across 21,142 evaluator responses, 87.14% rated the outputs helpful. First-pass review time fell by about half and interview preparation by about two-thirds. The work is filed as a patent — a recruitment document reading assistant using evidence-candidate generation and multi-agent verification — and moved into production for the first-half 2026 open recruitment.

Lessons learned

Conclusion

Reading application essays at group scale is a high-stakes, auditable task that a single LLM prompt handles poorly: it produces fluent text without evidence or checks. Indexing the essay by sentence, verifying and regenerating summaries, returning evidence as positions, judging question coverage against cited sentences and filtering by policy produced an assistant whose outputs can be traced to the applicant's own words — on a model the group hosts itself.

The same pattern — make the source addressable, let models point rather than quote, and place a specific check after each generator — applies to other document-reading tasks where the output has to be defended later.

Limitations

About the demo and confidentiality

The applicant, essay, competencies, questions and outputs in the embedded model are invented. No applicant data, talent-model wording or competency definitions, prohibited-question categories, interview check items, prompts, question sets or deployment configuration from the production system appear in this post; code is simplified and written for illustration.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterLed design and development; model internalization and benchmarking; patent inventor. Demo built on an invented essay for this site.