AI engineer · 8 years · TwoWeeks

In two weeks,I turn your idea into AIyou can use right away.

I build the systems behind the feature: the data, the models, the checks and the product around them. One person takes it from the first call to operations, and every two weeks you get something you can actually use.

  • 1M+people have used AI I built
  • $500M+in yearly business my AI supports
  • 36%average performance gain in the models I built
  • 2wksfrom idea to prototype

Everyone says you need AI. So why are you still stuck?

  • 01

    No idea where to start

    You hear a lot about what AI can do, but nobody tells you concretely how it fits your own work.

  • 02

    It works in the demo, not in real use

    The chatbot that looked great in the demo gives odd answers to real customer questions and real data.

  • 03

    Outsourcing that leaves nothing behind

    When the contract ends, the code, the documents and the know-how stay with the agency.

  • 04

    Quotes with no reasoning behind them

    You get a single number, with no way to see why it costs that or what dropping a feature would save.

I only promise what works, and show it every two weeks.

  1. 01

    Tested on your data first

    I build a small version on your real documents, calls and data, and only the features that prove useful go into full development.

  2. 02

    Something that works, every two weeks

    You see progress as something you can click through, not as a status report.

  3. 03

    A quote broken down by feature

    You see what each feature costs, so you can drop what you do not need and pay only for what you do.

What I build for you.

Eight things, each one shipped before. Click a card for how it works and where it has run.

RAG · Retrieval-augmented generation

An AI assistant that understands and uses your documents

Your team asks in plain language; the assistant reads your own files, answers with the source shown, and can act on the answer — draft the reply, fill the form, open the ticket. When the source is not there, it says so.

  • RAG
  • hybrid search
  • citations
  • streaming chat
  • KakaoTalk/Slack

Running live inside the production app of a $500M+ revenue company.

How it worksClose

What you get. An ingestion pipeline for your documents (PDF, Word, Excel, slides, Korean office formats, scans), a search layer that combines meaning and keywords, and a chat interface — web or inside a messenger — that answers with citations. Access control, memory and deployment on your infrastructure are part of the build, not extras.

How I keep it honest. Before launch we agree a set of real questions and the answers a good employee would give. The assistant is measured against that set, and every later change is measured again. An answer without a source is not an answer; the system says so.

In technical terms. Retrieval-augmented generation (RAG): semantic chunking, hybrid search (vector + keyword) with reranking, citation-backed answers, streaming chat, and an evaluation set scored by an LLM judge.

GraphRAG · Knowledge graph

A knowledge graph of your expert domain

When your knowledge is about how things relate — crops and treatments, parts and failures, rules and cases — a graph answers questions a document search cannot.

  • Neo4j
  • LLM extraction
  • text-to-Cypher
  • expert review loop

Built and shipped for a Fortune 500 subsidiary — the ontology, the extraction, the database and the agent on top.

How it worksClose

What you get. A schema designed with your experts, an extraction pipeline that turns your documents into the graph with quality control (conflicts flagged and resolved against human-decided cases), the database deployed with build/snapshot/reset tooling, and an assistant that queries the graph in natural language. Method documentation so your team can extend it after handover.

Why most vendors cannot show this. The hard part is not the database — it is constructing the data with measured quality. That construction is what I have done in production.

In technical terms. GraphRAG: ontology design, LLM extraction with quality control, Neo4j multi-database, text-to-Cypher agents, and hybrid graph + vector retrieval.

AI agents · Multi-agent orchestration

Agents that research and write, with a person in the loop

A team of AI agents does the research, scoring and drafting for one business workflow; a human approves before anything leaves the building.

  • LangGraph
  • CrewAI
  • approval gates
  • resumable runs
  • cost per agent

Export-buyer reports live at FederationLabs; a 20-agent marketing pipeline with two approval gates at STIA.

How it worksClose

What you get. One workflow — market research, buyer scoring, campaign strategy, proposal drafting — designed as a pipeline of specialist agents with validation steps: grounding against real data, a critic that revises the draft, and approval gates where a person must click. Outputs land in your database as structured records, not loose text.

What I insist on. Long runs must survive a crash and resume. Approvals must be enforced in code, not in a prompt. Every model call must have a cost attached. These are the parts that make an agent system usable a year later.

In technical terms. AI agents and LLM workflows: memory, tool use, guardrails, multi-agent orchestration (LangGraph, CrewAI), deterministic pipelines for multi-step business logic and structured output, human-in-the-loop gates, resumable runs.

LLM workflows · Document generation

Finished documents, in your template

Turn source material — recordings, PDFs, web data, a curriculum — into real PowerPoint, Word and Excel files that match your templates and open ready to edit.

  • PPTX/DOCX/XLSX generation
  • OCR
  • transcription
  • schema-validated output

Classroom decks for paying teachers at MangoFactory; consulting reports from phone calls for an agricultural enterprise.

How it worksClose

What you get. Ingestion for messy inputs (scans, Korean office formats, audio), generation that fills your template in place so branding and layout survive, image selection by meaning, and delivery by e-mail, storage or inside your app. Structured output is validated against a schema before it is written, so a bad generation fails loudly instead of shipping quietly.

What it is not. Not a markdown dump pasted into a slide. The output is the same kind of file your team already makes — only faster.

In technical terms. LLM workflows with schema-validated structured output, in-place PPTX / DOCX / XLSX generation, OCR and transcription for ingestion, semantic image retrieval.

LLM fine-tuning · Model deployment · Voice AI

Your own LLM model — fine-tuned and deployed

When an API is too expensive, too slow, too inaccurate or not private enough, I fine-tune a model on your data and deploy it on your hardware — the training and the serving, both, for text and for speech.

  • LoRA / full fine-tuning
  • vLLM / ONNX serving
  • speech-to-text
  • diarization
  • GPU + CPU deployment
  • annotation tooling

In-house models replaced GPT-4 in a 1M-user product; fine-tuned speech recognition at 10.45% error on real farm calls where the commercial API scores 16.39%.

How it worksClose

Fine-tuning. A data pipeline from your documents, logs or recordings, and, where needed, an annotation tool your team uses to build the training set. Base-model selection (open-weight LLMs, speech models), LoRA or full fine-tuning on a GPU cluster, and an evaluation against the API you use today on your own test set — so the decision to switch is a number on paper, not an opinion. If the API wins, I say so.

Deployment. The model served on your infrastructure: vLLM or ONNX Runtime, GPU or CPU, containerised, behind an OpenAI-compatible endpoint so your application swaps by URL. Batch and streaming modes, autoscaling where it pays, monitoring, and a CI/CD path so the next version ships the same way. Cost per request and latency are measured before and after.

In technical terms. Fine-tuning LLMs (LoRA and full) and training proprietary models for performance, cost and privacy constraints; domain-adapted STT, diarization and streaming transcription on a GPU cluster; vLLM and ONNX Runtime serving with an OpenAI-compatible surface.

LLM evaluation · Optimisation · MCP · Agent harness

Making your AI measurable — and systematically better

“It doesn’t work well and we can’t tell why.” I turn that into numbers, then into an evaluation-and-optimisation loop that improves them release after release — the same loop that took my own models to state of the art.

  • Evaluation harnesses
  • optimisation loop
  • LLM-as-judge
  • hallucination detection
  • MCP servers
  • agent harness

A state-of-the-art scoring model built with this loop (arXiv 2309.02740); a $18K, five-star fact-checking engagement; my own unattended agent runner in daily use.

How it worksClose

Measure. An evaluation framework — criteria agreed with you, scored datasets, judge models, versioned reports — and a regression suite wired to your CI, so a model or prompt change cannot silently degrade production. Where the problem is made-up claims, a per-claim fact-checking component.

Optimise. Then the loop: error analysis on the failing cases, one change at a time (prompt, retrieval, data, fine-tuning), re-measure, keep only what moves the number, and record what did not. This is how performance improves on a schedule instead of by luck. The same loop is how I took an essay-scoring model to state of the art before it served a million learners.

Run unattended. Where the goal is autonomy: MCP servers over your data and tools, a harness with permission rules, and a scheduled runner with locking, timeouts, bounded retries, cost caps and one aggregated alert per run.

The rule behind it. Verification runs in a separate process with a different model, because a system reviewing its own work misses its own mistakes.

In technical terms. Eval harnesses, LLM-as-judge scoring, hallucination detection, error-analysis-driven optimisation; MCP servers that expose your data and APIs as agent tools; the agent harness — scheduling, retries, permissions, human-in-the-loop — so agents run unattended.

End-to-end delivery · Full stack

The whole product, not just the AI

Frontend, backend, database, deployment. No engineering team yet? I build the product around the AI as well, and hand it over in a form a team can take on later.

  • React / Next.js
  • FastAPI / NestJS
  • PostgreSQL / Supabase
  • Docker
  • AWS / GCP
  • CI/CD

MangoFactory — a paying SaaS built solo, end to end, in five months; the crop-monitoring dashboard and gateway at an LG subsidiary.

How it worksClose

What you get. A user-ready product: the app your customers log in to, the backend and database behind it, payments and exports where needed, and deployment on your own cloud — from one person who has shipped the whole thing before, so the AI does not wait for a team that is not there yet.

In technical terms. React, Next.js or Svelte frontends; FastAPI or NestJS backends; PostgreSQL, Supabase or Firestore; Docker, AWS or GCP; CI/CD and infrastructure as code from day one.

AI transformation · Consulting

Finding where AI pays off in your business

Before anything is built, we find which of your processes gain the most from AI, prove it with a small pilot on your own data, and only then scale.

  • Process diagnosis
  • two-week pilot
  • success metrics
  • roadmap

Every case study on this page started this way. The two-week prototype package is the pilot.

How it worksClose

What you get. A short diagnosis of your workflows: where time goes, what data already exists, and which steps an AI could take over with a measurable result. Then one pilot, two weeks, on your data, with a number at the end that says whether it worked. Then a roadmap that orders the rest by payoff and risk.

What it is not. Not a slide deck about “AI strategy”. The output is a working pilot and a written plan you can hand to any engineer, including one who is not me.

In technical terms. Process and data audit, candidate-workflow scoring, evaluation criteria agreed before the pilot, and a phased roadmap with the scope and quote for each phase.

From the first call to operations, built as if it were my own.

  1. STEP 1

    Requirements & data

    Together we look at the problem and the data you already have. Call recordings, documents, databases: anything works.

  2. STEP 2

    Scope & quote document

    Hours and cost per feature, plus milestones, in one document a CEO can read and decide on.

  3. STEP 3

    Two-week build cycles

    Every two weeks you get something to use, and your feedback goes into the next two weeks. Each milestone has a written test you run yourself; payment follows the test.

  4. STEP 4

    Verify, hand over, operate

    I verify performance with measured numbers and hand over all code, documents and runbooks. I stay on after launch.

AI support chatbotSample
M1 · Data cleanup · first working buildDone
M2 · Field test · accuracy workIn progress
M3 · App integration · handoverPlanned
Next demoFriday, week 2

Progress, always on one screen

  • Milestone documentWhat finishes when, on one page
  • Two-week demoA working screen, not a promise
  • Project document repoPlans, designs and decisions shared with you

Only the features you need, with a quote you can read.

  1. 1

    Hours × rate, per feature

    Every feature is broken into the hours it takes and shown in a table, so you can see what each line costs. Parts I have built before cost their integration hours, and the quote says which rows those are.

  2. 2

    Three stacking packages

    Core, then extend, then refine. You choose how far to go.

  3. 3

    Scope that fits the budget

    A smaller budget does not mean lower quality. I cut scope instead, and design it so the rest can be added later. Every scope also lists what is not included.

Ask for a quote
Sample quote (actual quotes follow a consultation)
PackageFeatureHours
1 · CoreDocument collection & cleaning pipeline24h
1 · CoreChatbot that cites its sources40h
2 · ExtendAdmin console & answer-quality dashboard32h
3 · RefineKakaoTalk & app integration24h
Start here

Two-week prototype package

Find out in two weeks, on your own data, whether the idea actually works, before you commit to a big budget.

  1. DAY 1–3

    Goals & data

    Agree on the goal and success criteria, and gather the data.

  2. DAY 4–7

    First working build

    Get one core feature actually working.

  3. DAY 8–11

    Tested on real data

    Run it on real data, measure, and fix.

  4. DAY 12–14

    Review & next scope

    Review the results together and settle the scope and quote for full development.

What you keep after two weeks
  • A prototype you can use
  • A verification report with numbers
  • Scope & quote for full development
Ask about the package

What I built, and what changed.

Eight case studies, names used with permission. Click one for the full story, including the technical detail.

  • 1M+people have used AI I built

    TestGlider, an English-test practice service. I built and served the essay and speaking scoring models.

  • $500M+in yearly business my AI supports

    Farmhannong, an LG subsidiary. My crop-consulting AI runs inside its production app and has been rolling out to all farms since May 2026.

  • 36%average performance gain in the models I built

    For example, on farm phone calls a commercial speech API got 16 of every 100 characters wrong. My model gets 10.

  • 2wksfrom idea to prototype

    A first version running on your own data, ready for you to try within two weeks.

  • 5mofrom plan to commercial launch

    MangoFactory. A paying SaaS for teachers, built alone from the first line of code to launch.

  • $5M+raised by a client on my technology

    TestGlider raised this with my scoring model as its core technology. The model was shown to be state of the art in a published paper.

The agronomy knowledge graph in Neo4j — coloured nodes for conditions, diseases, cultivation types and more, joined by relationships

Farmhannong (FarmsAll)LG subsidiary · agriculture · $500M+ revenue2025–2026

Consulting AI inside a farming enterprise's app

Problem
Expert guidelines, past reports and consultants' calls were scattered, so nobody could reach them when a farmer asked.
What I built
A knowledge graph of the company's know-how, an assistant that answers with citations, and a copilot that drafts the consulting report from the call.
  • Live in production · pilot expanding to all farms from May 2026
  • Answers cite their sources; reports cite the transcript
  • One architecture map, one gateway, sixteen repositories kept in line
Full case studyClose

Every sentence the AI writes points to where it came from. If it cannot, the sentence is dropped.

  • ClientFortune 500 subsidiary, $500M+ annual revenue
  • StatusLive in the production app
  • ScopeKnowledge graph · assistant · report copilot · dashboard
  • Codebase16 repositories, ~1,000 commits on the core two

What they needed

A large agricultural company runs a consulting service for farms. Its best knowledge lived in expert guidelines, years of consulting reports and the consultants’ own phone calls. None of it was reachable at the moment a farmer or a junior consultant needed it. Writing up a consultation took as long as the visit itself.

What I built

Three things that share one brain. A knowledge graph of the company’s agronomy know-how, built from their documents with quality control by their own experts. An assistant in the app (web and KakaoTalk) that answers with citations and refuses when the source is not there. And a report copilot that turns a recorded consultation into a per-topic report, then exports it to PDF. Every bullet points to the sentence in the call it came from.

What changed

The assistant and the report copilot run inside the production app. The pilot is expanding to all farms from May 2026. The company’s team can extend the graph after handover, because the method, not just the result, was delivered.

Under the hood
  • Knowledge graph on Neo4j, 35+ node types, multiple databases per domain. Text-to-Cypher with type-routed few-shot examples; execution is read-only in code (writes blocked, 30-row cap).
  • LLM extraction with conflict resolution calibrated on 381 human-resolved cases, plus an LLM-judge evaluation loop, deterministic merge and dedupe.
  • Report pipeline in LangGraph: transcript → fact extraction → topic selection → per-topic subgraph (fetch references → map facts → disambiguate only when the deterministic mapper flags ambiguity → write → render → verify). Every bullet’s evidence list is validated against transcript turn indices; a bullet with no surviving evidence is removed. A separate verification pass is demote-only: it can downgrade, never promote.
  • Document intelligence: one loader for PDF, DOCX, XLSX, PPTX, CSV with encoding fallbacks, and an HWPX (Korean office format) parser written from scratch.
  • Measured, then left off: an LLM term-correction step was re-measured on 84 real calls and shipped disabled; 85 catalog rows and one rule did the job. The pattern throughout is LLM proposes, code applies and guards.
  • Operations: one gateway is the only caller of the GPU plane (OpenAI-compatible interface, so Gemini, OpenAI or the in-house model swap by URL); nginx + TLS; LangSmith tracing in production; per-repo architecture notes so coding agents read the map before touching code.
The knowledge graph rendered in Neo4j
The knowledge graph itself — conditions, diseases, disorders, cultivation types, beneficial insects — as the assistant searches it
GraphRAG pipeline diagram
How knowledge becomes a graph, and how the assistant uses it
Crop-monitoring dashboard
Crop-monitoring dashboard delivered alongside the assistant
  • LangGraph
  • Neo4j
  • FastAPI
  • React
  • Gemini
  • vLLM
  • Chroma
  • Redis
  • Docker
Bar chart of character error rate on three kinds of speech, our engine against a commercial API and the open-source base model

Farmhannong (FarmsAll)LG subsidiary · agriculture2025–2026

Speech recognition that beats the commercial API on real farm calls

Problem
Generic speech recognition mangled crop, chemical and brand names on phone calls, so nobody trusted the transcripts.
What I built
A speech model fine-tuned on farm calls, a 56,000-term vocabulary, a review platform and a gateway every other system calls.
  • 10.45% error on real farming calls vs 16.39% for the leading commercial API
  • From 16.18% (open-source base) to 10.45% with ~1,300 h of curated audio + 7 h of the client's calls
  • Self-hosted on 4 × B200 · four serving modes · annotation platform in daily use
Full case studyClose

Trained on 1% of the client's own calls. The other 99% is the roadmap.

  • Farming calls10.45% CER · commercial API 16.39% · base model 16.18%
  • Everyday speech7.16% · commercial API 7.33% — nothing lost
  • Training data~1,300 h curated from ~14,000 h reviewed, + 471 of the client's own calls
  • Serving4 × B200, vLLM, batch / streaming / diarized

What they needed

Consultants speak with farmers by phone all day. Those calls hold the facts that should end up in the farm’s diary and the consultation report. But generic speech recognition mangled the crop names, chemicals, brand names and local terms, and users stopped trusting the transcripts. Off-the-shelf engines are trained to score well on news and studio conversation; nobody optimises them for farming calls on a mobile phone.

What I built

Four parts. A human-review platform where the team corrects transcripts and builds the training set. A fine-tuned speech model with speaker separation, served on the company’s own GPUs. A 56,000-entry farming vocabulary that feeds recognised terms back to the engine as hints. And a gateway that every other system calls — batch or streaming, with or without speaker labels. Downstream, the transcript feeds the report copilot.

What changed

Measured on 2,354 sentences from the client’s real calls, with the same scoring rules for every engine. Our engine: 10.45% character error rate. The leading commercial Korean speech API: 16.39%. The open-source model we started from: 16.18%. On everyday conversation the engine stays level with the commercial API (7.16% vs 7.33%), and on public mobile-phone recordings it is ahead (10.54% vs 11.56%). By the industry’s usual reading, around 20% is “usable but double-check” and around 10% is “trusted”. The client moved from the first band to the second — with only about 1% of their recorded calls labelled.

Under the hood
  • Data: eight public Korean speech corpora (~14,000 h) reviewed, ~11,000 h licensed and cleaned, ~1,300 h (~910k sentences) used for training with telephone-channel recordings raised to 35% of the mix — the mix, not the volume, moved the number. Telephone recordings were verified by channel fingerprint. Plus 471 of the client’s own calls (~9 h): ~7 h for training, 93 calls (~2 h) held out as the fixed judge set.
  • Acceptance rule: every candidate engine is judged on the same held-out set of the client’s calls; everyday-speech performance may not regress by more than 0.3 points.
  • Scaling evidence: with the same training volume, going from 99 to 347 client calls moved the error from 12.7% to 11.3% — roughly 0.7 points per doubling. About 1,000 h of client calls remain unlabelled; the projection is ~7% with all of them and 5% with vocabulary and post-processing work.
  • Model and serving: Qwen3-ASR fine-tuning and pyannote diarization on 4 × B200, served with vLLM behind an OpenAI-compatible /v1/audio/transcriptions surface; modes: batch, streaming, diarized, diarized-streaming.
  • Review platform: TypeScript monorepo (NestJS, Postgres, S3), 771 commits; a shared text-correction package consumed by the API, worker and web apps.
  • Measurement hygiene: same normalisation for all engines (punctuation removed, number spelling unified, spacing ignored). Inconsistent spacing in the judge set alone was worth 1.7 points, so the transcription guideline is fixed before mass labelling. A post-processing engine measured on 2,354 utterances added nothing and was documented as such.
  • Housekeeping: checkpoint purge reclaimed 9.7 TiB of GPU storage.
Character error rate on three kinds of speech — our engine, the commercial API, the open-source base
Measured 24 Sep 2026 on held-out audio, same normalisation for every engine
Voice pipeline diagram from recording to report
From recording to report — the pipeline
Audio labeler platform
The human-review platform where the training set is built
  • Qwen3-ASR
  • pyannote
  • vLLM
  • FastAPI
  • NestJS
  • PostgreSQL
  • S3
TestGlider product

TestGlider (Databank)EdTech startup · nearly 1M users2020–2023

Scoring a million learners' essays and speech

Problem
Essays and spoken answers needed instant, consistent scores at a scale no human grading or per-call vendor could support.
What I built
In-house essay, speaking and correction models plus self-hosted speech-to-text, run as four production services on AWS.
  • Nearly 1M users · supported a $5M+ raise
  • In-house model replaced the LLM correction pipeline
  • Self-hosted speech-to-text replaced a per-call vendor bill
Full case studyClose

The in-house model did not just match the API. It replaced it in production.

  • UsersNearly 1M
  • Funding supported$5M+ raise
  • Services in production4, with staging and production CI/CD
  • ResearcharXiv 2309.02740 — state of the art at publication

What they needed

TOEFL and IELTS learners want to know, right now, what their essay or spoken answer would score. Human graders do not scale, and paying a vendor for every answer does not either. The scores also had to hold up: a junk answer must score zero, and the score the user sees must be the score the model produced.

What I built

The company’s own essay-scoring model — the research behind it was state of the art when published — and a speaking-scoring model that listens to the audio and reads the transcript. Later, an in-house correction model that replaced the GPT-4 pipeline, and a self-hosted speech-to-text service that replaced the vendor bill. All four ran as production services on AWS with a test harness that checked the live models on every deploy.

What changed

The product scaled to nearly a million users on models the company owned, and the AI was a core part of a $5M+ raise.

Under the hood
  • Essay scoring: transformer regression with rubric-specific heads, adversarial training and augmentation (arXiv:2309.02740). Speaking: wav2vec audio features fused with text.
  • Serving: Flask/uwsgi with Celery + RabbitMQ workers, Docker → ECR → ECS, separate staging and production pipelines. Speech-to-text ported to ONNX Runtime for cost.
  • Harness: live-model checks on deploy — junk answers must score 0, served scores must match local inference.
  • Adjacent: a TypeScript RAG application for the same product line (NestJS, hybrid vector + keyword search, React).
Model serving architecture
How the models are trained, tested against live traffic and served
  • PyTorch
  • HuggingFace
  • wav2vec
  • ONNX Runtime
  • Flask
  • Celery
  • RabbitMQ
  • Docker
  • AWS ECS
Six slides from a generated classroom board-game deck, in the school's own template

MangoFactoryEdTech SaaS · paying teachers2025

Classroom-ready lesson decks, in the school's own template

Problem
Teachers rebuilt lesson decks, worksheets and quizzes by hand in the school's template, every evening.
What I built
The whole product: an engine that fills the school's own PowerPoint in place, retrieval over 120GB+ of curriculum, the app and the admin console.
  • Paying users; certified teachers use it in real classrooms
  • Built solo in five months, July–November 2025
  • Fills the client's own PPTX template — nothing is rasterised
Full case studyClose

Not a slide picture. A real deck a teacher can open and change.

  • Corpus120GB+ curriculum, retrieval-augmented
  • OutputsPPTX · DOCX · XLSX · quizzes
  • Verified deck59 slides, 11 layouts, fully editable
  • TeamOne person, ~230 commits

What they needed

Korean elementary teachers spend their evenings making lesson decks, worksheets and quizzes. The client wanted a product that produces those files — in the school’s own design, editable, aligned to the curriculum — from a short description of the lesson.

What I built

The whole product: the teacher-facing app, the admin console, the backend, and the generation agents. The core is a template engine that fills the client’s own PowerPoint in place, so the branding is the template itself and every text box stays editable. A retrieval layer over a 120GB+ corpus grounds the content; images are chosen by meaning, not keyword.

What changed

Teachers pay for it and use the output in class. The deck length follows the lesson duration; the template survives untouched.

Under the hood
  • In-place PPTX filling with python-pptx: rewrites existing text frames, runs and table cells, then copies run/paragraph/frame styles forward. Slide masters, layouts and theme are inherited from the source deck.
  • Layout fidelity: stable shape addressing across grouped shapes and tables, geometry save/restore, and binary-search font autofit with word-wrap simulation so text fits the original box.
  • Constrained generation: LLM output is bound to a Pydantic schema (slide, shape, content) with a deterministic fallback mapping.
  • Deck pipeline in LangGraph: computes target length from lesson duration, selects and reorders which template slides survive, swaps images by semantic vector search (Pinecone + embeddings) at the original placeholder geometry.
  • Ingestion: OCR and image auto-tagging; adjacent DOCX/XLSX/CSV/PDF generators in the same codebase.
  • Stack: Lovable/React apps, Supabase + FastAPI backend, S3, Redis, EC2/Vercel.
Six slides from a generated 29-slide classroom board-game deck
Six of 29 slides from one generated deck — a classroom board game for a 4th-grade Korean lesson, filled into the school's template (Korean text; every box is editable)
MangoFactory app — the teacher-facing product
The teacher-facing app — describe the lesson, receive the files
Generated quiz slide
A quiz slide from a second generated deck
  • LangGraph
  • python-pptx
  • Pinecone
  • FastAPI
  • Supabase
  • React
  • Lovable
  • S3
  • Redis
First page of a generated buyer research report

FederationLabsTrade-intelligence startup2025

Export-buyer research reports, written by a team of agents

Problem
Finding and vetting overseas buyers took an analyst days per product.
What I built
Eight crews of agents that classify the goods, search companies, pull trade records, check regulations, grade buyers and write the report.
  • Delivered March–June 2025, live at federationlabs.ai
  • Buyer profiles graded A/B on real trade records
  • A self-critique pass revises the report before it ships
Full case studyClose

The report grades each buyer, and the grade has a reason attached.

  • Agents8 crews, 10 agents
  • ResearchBrave search + Firecrawl
  • DeliveredIdea → production, 4 months
  • OutputStructured report + chat over it

What they needed

Finding the right overseas buyer for a product used to take an analyst days per product: classify the goods, search companies, pull trade data, check regulations, rank, write. The client wanted that as a product.

What I built

A multi-agent pipeline of eight crews. One extracts the trade classification, one searches companies, one pulls trade records, one researches regulations. One scores buyers, one writes, one checks quality and revises, and one lets the user chat over the finished report. Grades are based on real trade records, and the writer must survive its own critic.

What changed

The service went from an idea to a deployed product in four months and is live today.

Under the hood
  • CrewAI crews with typed handoffs; web research via Brave search and Firecrawl (the only managed scraper — no headless browsers).
  • Scoring from trade records, producing A/B buyer grades with the reasoning kept.
  • Quality loop: a self-critique revision step before the report is stored.
  • Platform: FastAPI, Supabase, Redis; reports persisted and served with a chat interface.
Report page two
Buyer profiles with grades
Report page three
Regulations and next steps
Multi-agent pipeline diagram
The crews, in order
  • CrewAI
  • Brave
  • Firecrawl
  • FastAPI
  • Supabase
  • Redis
Fact-checking pipeline diagram

Upwork client (confidential)Software company · Upwork engagement2024–2025

Catching an AI's made-up claims before the user does

Problem
Answers sometimes contained things the reference documents never said, and nobody could tell which sentence.
What I built
A pipeline that splits the answer into claims, checks each against the source and rewrites only what failed, shipped as an SDK.
  • 5★ review · $18K engagement
  • Shipped as a pip-installable SDK
  • Failures are per claim, so the fix is per claim
Full case studyClose

Five stars, eighteen thousand dollars, and a component you can drop into any pipeline.

  • EngagementNovember 2024 – March 2025
  • Rating5★, $18K
  • Delivered asPython SDK, wheel + sdist
  • EvaluationSynthetic hallucination injection, confusion-matrix scoring

What they needed

Their product answered questions from reference documents, and sometimes the answers contained things the documents never said. They needed to know which sentence was wrong, automatically, and to fix the answer rather than throw it away.

What I built

A pipeline that turns an answer into subject–predicate–object claims and checks each claim against the reference. When too many fail, it re-prompts with the exact failure map, so the rewrite fixes what was wrong and keeps what was right. Delivered as an installable SDK with baselines and an evaluation kit so the client can measure it on their own data.

What changed

The client left a five-star review. The component is transferable: it plugs into any question-answering system that has reference text.

Under the hood
  • Verdicts per claim: each answer triplet is checked against all reference triplets in one constrained LLM request whose output is forced to triplet_idx:result, then merged into a verdict map.
  • Challenge-and-revise: above a threshold, a reprompter regenerates from question + answer + references + the per-triplet failure map.
  • Swappable parts by CLI flag (answer generator, triplet extractor, checker) with non-LLM baselines — Stanford OpenIE, exact and partial match — for comparison.
  • Evaluation kit: synthetic hallucination injection and confusion-matrix scoring.
  • Packaging: LLMTripletValidator public API, built as wheel and sdist.
Rendered map of claims and verdicts
A rendered map of one answer's claims and their verdicts
  • Python
  • LLM APIs
  • Stanford OpenIE (baseline)
Agent orchestration diagram

STIAMarketing-technology company · active engagement2025–2026

Twenty specialist agents, two human sign-offs

Problem
A marketing platform had to draft strategy and outreach at scale without the AI ever publishing or contacting anyone on its own.
What I built
Twenty specialist agents with two human approval gates, resumable runs, and a lead pipeline synced with the CRM.
  • 20 specialist subagents with two sequential approval gates
  • Runs survive process death and resume mid-way
  • ~1,300 tests across the backend
Full case studyClose

The AI can propose. Only a person can release.

  • OrchestrationLangGraph, fan-out into 4 branches, fan-in
  • Human gatesExpert → client, per-gate authorisation
  • ModelsAnthropic, OpenAI, Gemini; Perplexity for research
  • Tests~1,300 test functions, 147 files

What they needed

A platform that produces marketing strategies and runs lead outreach at scale. The AI never publishes or contacts anyone on its own, and an interrupted run loses no work.

What I built

A strategy pipeline of twenty specialist agents that fan out, come back together, and stop at two gates: an expert reviews, then the client approves. A lead pipeline that qualifies leads by explicit rules (no AI guesswork on who to contact), enriches them and syncs with the CRM. It listens to LinkedIn outreach events and drafts replies that a human approves before sending. Every run keeps a decision trace and a cost line per agent.

What changed

The client has a system where approvals are enforced in code, runs are resumable, and every model call has a price tag. (This engagement is ongoing; no outcome figures are claimed.)

Under the hood
  • Strategy DAG as a LangGraph StateGraph; each subagent is its own class with its own prompt module and typed Pydantic input/output. A conditional edge skips the client gate when routing returns BLOCKED.
  • Gates via interrupt() with per-gate caller authorisation; every run records per-node human | logic sources and an immutable decision trace.
  • Durability: Firestore checkpointer; a run paused in a dead process is rebuilt by class name + thread id; mid-DAG entry for partial re-runs. Approvals are first-write-wins and idempotent because LangGraph re-runs a node on resume; an append-only ledger records every delivery attempt; reconciliation retries are capped, then dead-lettered.
  • Cost control: token tracker (model, tokens, cost per call per agent), usage quotas → 429, hard cap on concurrent campaigns.
  • Lead pipeline: rule-based three-tier qualification (auto / human review / exclude) with a written reason per lead; enrichment via Apollo and Apify; HubSpot two-way CRM sync built; LinkedIn outreach webhooks reference-counted in a Firestore transaction; reply drafts held in a review queue until a human releases them.
  • LangGraph
  • Firestore
  • FastAPI
  • Anthropic
  • OpenAI
  • Gemini
  • HubSpot API
From research model to production serving

Published workApplied AI research2016–2025

State-of-the-art AI model research, put into production

Problem
Essay scoring needed a model that was both state of the art and ready for a product.
What I built
The rubric-specific scoring model in the paper, plus prompt-selection, RLHF and hallucination-detection research since.
  • State-of-the-art result, arXiv 2309.02740 — planned, run and written by me
  • Model went straight from paper to a 1M-user product
Full case studyClose

An experiment is only finished when it runs in a product.

  • PaperRubric-Specific Approach to Automated Essay Scoring with Augmentation Training
  • ResultState of the art at publication
  • Other topicsPrompt selection with LLM feedback · RLHF · hallucination detection
  • EarlierSales forecasting · disease-surveillance clustering · smart farming

The published work

Rubric-Specific Approach to Automated Essay Scoring with Augmentation Training (arXiv:2309.02740). Planned, designed, executed and authored by me, from research question to published paper: a state-of-the-art result on automated essay scoring. The model it describes went into production and scored essays for nearly a million learners.

Other research

  • Optimised prompt selection using a task-specific model and LLM feedback (conference submission).
  • Reward modelling and reinforcement learning from human feedback (T5 reward model + PPO).
  • Hallucination detection for LLM answers (2025) — shipped as the fact-checking SDK.
  • Earlier applied work: SKU-level and agri-food sales prediction, spatial clustering of migratory birds for avian-flu surveillance, smart-livestock and soil-sensor modelling, and risk signals from PC log data.
Why this matters for a client

The research habit shows up in delivery as evaluation sets, judge models and before/after numbers — and in the willingness to switch a feature off when the measurement says it adds nothing.

Two more projects are under NDA.

They are described below without the client, the industry or any figures, so you can judge whether the capability fits your problem.

See the confidential work

Work I cannot name.

Confidential · 2026

A conversational agent with memory and guardrails

An AI companion that remembers the user across conversations and stays inside strict behavioural limits, including when parts of the system are unavailable: it degrades gracefully instead of failing.

  • LangGraph 1.0
  • long-term memory
  • guardrails
  • evaluation harness
  • tracing
How it was built
  • Long-term memory layered over the conversation graph; guardrails include an offline language check that needs no API call.
  • An evaluation harness and full tracing, so behaviour changes are measured, not felt.
  • Graceful degradation by capability tier: the agent keeps answering with less when a dependency is down.
  • Typed, linted and tested (mypy, ruff, pytest) and handed over with documentation.
Confidential · 2026 · in production

An AI service the platform team never has to open

The AI brain and the client's platform are separated by a written interface contract with a stub mode, so their engineers integrate against a stable surface while the AI evolves behind it.

  • FastAPI async
  • job queue
  • real-time events
  • PDF/DOCX export
  • payments
  • IaC + CI
How it was built
  • Service interface specification first; a stub implementation lets the platform team build and test before the AI is finished.
  • Async API with background jobs and real-time updates; document export; payment integration; error monitoring.
  • Infrastructure as code and CI from day one; running in production.

I design it, build it, and prove it with numbers.

  • 8yrsin AI research & development
  • SOTAworld-leading results, shown in a paper
  • 100%of code and documents handed over
  • The engineer behind these cases, directly

    No split between sales and engineering. The AI engineer with eight years of experience who built the cases above takes the first call, designs the system, builds it and runs it. The person you talk to is the person who writes the code.

  • All code and documents handed over

    Code, trained models, design documents, evaluation sets and runbooks all become your assets. Vendor parts are called "integrate", never hidden as "built".

  • Verified with measured numbers

    Instead of "it works well", I report performance as numbers measured on the same data, by the same standard. Features that measure zero gain are switched off and documented.

  • Operations after launch

    Monitoring, incident response, retraining and new features: I stay on after launch.

Frequently asked questions

Q.Why two-week cycles?

With AI projects, a lot is unknown until you actually use the thing. Seeing something that works every two weeks lets us correct course early, which means less wasted money.

Q.Do we get the source code and deliverables?

Yes. Source code, trained models, data-processing scripts and design and operations documents are all handed over, and what is handed over is written into the contract.

Q.How do I pay?

Fixed price per milestone. Each milestone comes with a written test you run yourself, and payment follows the test, not the calendar. Through Upwork or a direct contract with TwoWeeks.

Q.How do we know the AI is good enough?

I check it with numbers measured on real data. For speech recognition, for example, I ran the same 2,354 call sentences through my model and a commercial service and compared the errors side by side.

Q.What if our data is a mess?

That's fine. I start by turning scattered documents, recordings and spreadsheets into something usable.

Q.Are you a freelancer or a company?

Both. I work through my company, TwoWeeks, and also through Upwork. Either way you work with me directly.

In two weeks, I'll show you something that works.

Message me on Upwork with the problem you want solved. You leave the first call with the problem stated in one sentence and a view on whether AI is the right tool for it. And if not, what is.

Working in Korea? The same work is offered through my company, TwoWeeks. twoweeks-acu.pages.dev ↗

Include these for a faster reply
  1. Your company and service
  2. The problem to solve
  3. The data you have
  4. Timeline and budget