← All work

Upwork client (confidential)Software company · Upwork engagement2024–2025Sole engineer, delivered as an SDK

Catching an AI's made-up claims before the user does

Five stars, eighteen thousand dollars, and a component you can drop into any pipeline.

An AI answer is split into individual claims; each claim is checked against the reference documents, and the answer is rewritten around exactly the claims that failed.

  • EngagementNovember 2024 – March 2025
  • Rating5★, $18K
  • Delivered asPython SDK, wheel + sdist
  • EvaluationSynthetic hallucination injection, confusion-matrix scoring

What they needed

Their product answered questions from reference documents, and sometimes the answers contained things the documents never said. They needed to know which sentence was wrong, automatically, and to fix the answer rather than throw it away.

What I built

A pipeline that turns an answer into subject–predicate–object claims and checks each claim against the reference. When too many fail, it re-prompts with the exact failure map, so the rewrite fixes what was wrong and keeps what was right. Delivered as an installable SDK with baselines and an evaluation kit so the client can measure it on their own data.

What changed

The client left a five-star review. The component is transferable: it plugs into any question-answering system that has reference text.

Under the hood
  • Verdicts per claim: each answer triplet is checked against all reference triplets in one constrained LLM request whose output is forced to triplet_idx:result, then merged into a verdict map.
  • Challenge-and-revise: above a threshold, a reprompter regenerates from question + answer + references + the per-triplet failure map.
  • Swappable parts by CLI flag (answer generator, triplet extractor, checker) with non-LLM baselines — Stanford OpenIE, exact and partial match — for comparison.
  • Evaluation kit: synthetic hallucination injection and confusion-matrix scoring.
  • Packaging: LLMTripletValidator public API, built as wheel and sdist.
Rendered map of claims and verdicts
A rendered map of one answer's claims and their verdicts
  • Python
  • LLM APIs
  • Stanford OpenIE (baseline)