All work

Case study / Teamcast.ai / 2025 to present

Maya: live AI interviews a regulated buyer can audit.

A real-time voice interviewer and the scoring system behind it. The hard part was never making the model talk. It was making every automated hiring judgment trace back to something the candidate actually said.

concurrent live sessions sustained
250concurrent live sessions sustained
completed conversations
5,000+completed conversations
end users
6,000+end users
points on a held-out benchmark
+9.2points on a held-out benchmark

Production figures, 2026.

01

Who was the customer?

Enterprise hiring teams screening at volume, in workflows where an automated decision about a person can be challenged. That includes an agent-to-agent integration with a Fortune 100 partner, where I was the sole technical point of contact from the first ambiguous ask through a V1 to V2 API migration and the migration guide their engineers built against.

The buyer is not the recruiter alone. It is the recruiter, their legal team, and the candidate who may later ask why they were rejected.

02

What was broken?

First-round screens do not scale, and the obvious fix makes things worse. A model that reads a transcript and returns a score is fast, but it cannot say which sentence earned which point, it rewards candidates who talk longer, and it gives a compliance reviewer nothing to inspect.

The live conversation had a second, opposite constraint. It has to respond in well under a second, which leaves no room for careful, repeated scoring in the loop.

03

What did I build?

The production voice agent end to end, and all three pipelines around it. Speech in through Deepgram, Whisper, and Google Speech; speech out through Cartesia and ElevenLabs; inference on Groq and Google AI Studio, in one low-latency conversation loop. Around it sits a compliance layer: 45 evidence types mapped to 36 signals across 6 competency clusters, per-score confidence intervals, PII-safe transcript handling, and an audit trail behind every automated decision.

P1

Before the call

Generates a role-specific interview plan from a canonical ontology. A human approves it before any candidate is invited.

P2

During the call

Listens for evidence turn by turn and decides whether to probe or move on. Optimized for latency, never for scoring.

P3

After the call

Scores the full transcript offline, where it can be slow, repeated, and audited. A second human checkpoint gates the result.

04

How does it work?

The core decision was to decouple live routing from scoring. The interviewer only decides what to ask next. Scoring happens after the call, where it can be slow, run several times, and be audited. Conversations stay fast and scores stay defensible, and neither is traded for the other.

During the call, the choice to probe or move on is not left to the model. It runs through eight guardrails in a fixed order, first match wins:

  1. G1Session critical. Under two minutes left: advance, and mark unprobed signals SKIPPED.
  2. G2Time pressure. Remaining time per remaining topic falls below the floor: advance.
  3. G3Topic overtime. A topic has run past twice its fair share: advance.
  4. G4Minimum turns. Only one turn so far with plenty of time left: follow up, so one rich answer cannot close a topic.
  5. G5Coverage sufficient. Enough signals observed for this topic: advance, mark the rest MISSING.
  6. G6Deepening. Nothing left to probe but most of the session remains: ask one broader question.
  7. G7No gaps. Nothing left to probe: advance.
  8. G8Fallback. Gaps exist and time remains: follow up.

The difference between MISSING and SKIPPED is kept as an audit fact. MISSING means there was time and the signal never appeared. SKIPPED means the clock forced a move before it could be probed. Extended-time accommodations relax only the clock rules, never coverage or scoring: an accommodation changes the time, not the measurement.

After the call, the boundary between model and code is explicit:

DecisionOwnerHow
Whether to probe or advanceCodeEight ordered guardrails, first match wins
What the candidate demonstratedModelExtracts evidence units with a type and a quote
Which signal that evidence supportsCodeCanonical evidence-to-signal map
How strong a signal isCodeMaximum contribution, not the mean
Rubric componentsModelTemperature 0, structured output, no access to the transcript
The final scoreCodeA deterministic formula over those components
Whether a result shipsHumanThreshold gates, then human review

Signal strength takes the maximum contribution rather than the mean. A candidate is as capable as their strongest demonstration, and it makes padding useless: restating a weak point ten times cannot raise a score. That single aggregation choice is the system's main defense against verbosity.

05

How did I evaluate it?

I ran a benchmark sprint that fine-tuned the assessor model and measured it against a held-out set, lifting output quality by +9.2 points. It shipped to production behind threshold gates, with human-in-the-loop escalation for anything below them.

What I got wrong first. An earlier version pitted two model calls against each other, a validator and a resolver, and let the more pessimistic verdict win. It looked rigorous. In practice a single skeptical call could wipe out a strong, well-evidenced score, so the stage added variance while claiming to reduce it. I removed it and moved that judgment into the deterministic formula, where it can be inspected.

The same lesson applied to rating bands. Fixed, adaptive score bands drifted as the candidate pool changed, so they were replaced with cohort ranking.

06

What would I change before the next stage?

  • Version stamps enforced in CI. Every score carries the pipeline version that produced it. That contract only holds if a scoring change cannot merge without moving the stamp, so the check belongs in CI rather than in review.
  • A published calibration set. The held-out benchmark proves improvement internally. A buyer's auditor should be able to rerun a fixed, anonymized calibration set and get the same numbers.
  • Failure drills, not just failure handling. Session state persists after every change so a dropped call can resume and be partially scored. That path should be exercised on a schedule, not discovered in an incident.