Skip to content

Outcomes and checks

GET /v1/conversations/{id} returns the conversation as you sent it, plus what Midwater found. id can be Midwater’s ID or your external_id.

status
queued Stored, waiting to be scored.
evaluating Being scored now.
done Scored: outcome and results are final.
failed Scoring didn’t finish. Midwater retries automatically, up to five attempts in all; if it still fails, it stays failed and someone on your team can re-run it from the conversation in the app.

Scoring usually finishes within seconds. Poll until done or failed (the SDKs’ wait() does this), or subscribe to the conversation.evaluated webhook.

outcome says how the conversation ended for the caller. It drives the resolution rates on agent health.

outcome Shown in the app as
resolved Resolved The caller got what they called for.
unresolved Unresolved They didn’t, or they asked for a person and didn’t get one.
escalated Handed to a person A transfer or handoff happened, or a person was requested and provided.
not_real_inquiry Not a customer call Sales calls, spam, wrong numbers. Kept out of resolution rates.
null Still scoring, or no outcome check applied to this conversation.

Planned Two values get clearer names: escalated becomes handed_to_person, and not_real_inquiry becomes not_customer_call. Accept both until the change is announced in the changelog; the SDKs’ types already do.

results has one entry per check that ran:

{
"check_key": "appointment_not_completed",
"check_name": "Appointment request not completed",
"check_version": 1,
"check_status": "active",
"score": 0,
"verdict": "pass",
"choice": null,
"decided_by": "rule",
"reason": "A booking tool call succeeded.",
"evidence_turns": [],
"scorer_version": "2026-10-01.1",
"latency_ms": 0
}
Field
check_key The check’s stable key. Use it to send feedback.
check_name The check’s name as your team sees it.
check_version Increases when someone edits the check.
check_status active, or shadow for checks being tried out: scored, but they never alert.
verdict The check’s result. pass: no problem found. fail: Midwater found this problem. uncertain: unclear, a person should look. not_applicable: didn’t apply to this conversation. Gating questions, such as whether this was a real customer inquiry, answer met or not_met.
score Likelihood that the problem occurred, 0 to 1, or null.
choice The chosen option, for multiple-choice checks.
decided_by rule: decided from the events you sent. model: Midwater’s model read the whole conversation. llm_judge: a second review, for conversations that were unclear (planned rename: second_review; accept both). human: someone on your team.
reason Why, in a sentence or two.
evidence_turns The transcript turns the result points to.
scorer_version Midwater’s version of how the result was scored, for example 2026-10-06.3. It changes when the scoring changes.
latency_ms How long this check took.

Which checks run is set per project in the app, on the Checks page. A conversation whose events show a booking also gets the built-in booking checks.

When someone on your side reviews a conversation, send their answer for a check back to Midwater. Feedback names the check by its check_key, which you take from the conversation’s own results:

  1. Get the conversation: GET /v1/conversations/{id} (Midwater’s ID or your external_id).
  2. Pick the result your team reviewed from results, and take its check_key.
  3. Send the answer: POST /v1/conversations/{id}/feedback with that check_key and pass or fail.
const conversation = await midwater.conversations.get("call_8f2a91");
const result = conversation.results.find((r) => r.check_key === "appointment_not_completed") ?? conversation.results[0]!;
await midwater.conversations.feedback(conversation.id, { check_key: result.check_key, verdict: "fail", note: "The booking failed after the call." });

verdict is pass (no problem) or fail (the problem happened); note is optional, up to 1,000 characters. The answer is recorded against that check’s result with source api, and GET /v1/conversations/{id} keeps returning the result as Midwater scored it. Each request records a new answer (idempotency keys on this endpoint are planned), so don’t retry it blindly. A check_key with no result on that conversation answers 404.

There’s no endpoint to list a project’s checks or its conversations yet: both are planned, and they’ll be listed in the changelog when they ship. Until then, check_keys come from a conversation’s results, and the Checks page in the app lists every check.