← All work

TestGlider (Databank)EdTech startup · nearly 1M users2020–2023AI researcher and engineer, model to production

Scoring a million learners' essays and speech

The in-house model did not just match the API. It replaced it in production.

An English-test practice service needed instant, reliable scores for essays and spoken answers — at a scale no human grading or per-call vendor bill could support.

  • UsersNearly 1M
  • Funding supported$5M+ raise
  • Services in production4, with staging and production CI/CD
  • ResearcharXiv 2309.02740 — state of the art at publication

What they needed

TOEFL and IELTS learners want to know, right now, what their essay or spoken answer would score. Human graders do not scale, and paying a vendor for every answer does not either. The scores also had to hold up: a junk answer must score zero, and the score the user sees must be the score the model produced.

What I built

The company’s own essay-scoring model — the research behind it was state of the art when published — and a speaking-scoring model that listens to the audio and reads the transcript. Later, an in-house correction model that replaced the GPT-4 pipeline, and a self-hosted speech-to-text service that replaced the vendor bill. All four ran as production services on AWS with a test harness that checked the live models on every deploy.

What changed

The product scaled to nearly a million users on models the company owned, and the AI was a core part of a $5M+ raise.

Under the hood
  • Essay scoring: transformer regression with rubric-specific heads, adversarial training and augmentation (arXiv:2309.02740). Speaking: wav2vec audio features fused with text.
  • Serving: Flask/uwsgi with Celery + RabbitMQ workers, Docker → ECR → ECS, separate staging and production pipelines. Speech-to-text ported to ONNX Runtime for cost.
  • Harness: live-model checks on deploy — junk answers must score 0, served scores must match local inference.
  • Adjacent: a TypeScript RAG application for the same product line (NestJS, hybrid vector + keyword search, React).
Model serving architecture
How the models are trained, tested against live traffic and served
  • PyTorch
  • HuggingFace
  • wav2vec
  • ONNX Runtime
  • Flask
  • Celery
  • RabbitMQ
  • Docker
  • AWS ECS