Question Types

noul, choice and score — the exact criteria shape each one takes, the exact answer shape each one returns, and the two mistakes that cost everyone an afternoon.

Three question types cover everything the model answers. They differ in what criteria looks like on the way in and what the answer looks like on the way out.

TypecriteriaAnswerUse for
noulnonea probability in 0..1yes/no, "does this apply?"
choiceobject (label → description)one label plus a distributionrouting, classification, matching
scorearray (tier labels, low to high)a decimal position plus a distributionseverity, priority, quality

The mistake worth reading twice

choice takes an object and score takes an array. They are not interchangeable, and the raw model answers a swap with a bare 400 reading Invalid input: expected record, received array — which does not say which question, or which direction. We check both shapes before the request leaves our server and answer with the question id and the shape that was expected.

noul — A Probability

No criteria. The answer is a single probability that the statement holds.

{
  "spam": {
    "type": "noul",
    "instructions": "Is this message unsolicited promotional content?"
  }
}
{
  "spam": { "type": "noul", "noul": 0.93 }
}

noul carries no separate confidence field, and it does not need one: distance from 0.5 is the confidence. 0.93 is a confident yes, 0.51 is a coin flip that happens to lean yes. Thresholding at 0.5 throws that away — a good rule is usually two-sided, e.g. auto-act above 0.9, auto-ignore below 0.1, queue the middle for a human.

choice — One Label Out of a Set

criteria is an object mapping each label to a short description of what belongs under it. The descriptions are what the model reads, so write them as boundaries, not synonyms of the label.

{
  "queue": {
    "type": "choice",
    "instructions": "Route this ticket to the team that should own the first reply.",
    "criteria": {
      "billing": "Invoices, refunds, duplicate charges, plan changes, tax.",
      "technical": "API errors, integrations, outages, SDKs.",
      "sales": "Pre-purchase questions about plans, limits or trials."
    }
  }
}
{
  "queue": {
    "type": "choice",
    "choice": "billing",
    "probabilities": { "billing": 1, "technical": 0, "sales": 0 },
    "confidence": 1
  }
}
  • choice is always one of your labels.
  • probabilities is keyed by label and sums to ~1. Use it to detect the cases worth escalating: {"billing": 0.52, "technical": 0.48} is a real decision your router should not take alone.
  • Between 2 and 24 labels, each up to 64 characters, each description up to 400.

Do not send an array

"criteria": ["billing", "technical", "sales"]

This is the score shape. On a choice question it is rejected with a 400 whose detail names the question and the object shape it expected.

score — A Position on an Ordered Scale

criteria is an array of tier labels, ordered from lowest to highest. The index is the scale: four labels means a 0..3 scale.

{
  "anger": {
    "type": "score",
    "instructions": "Rate the customer's frustration.",
    "criteria": ["Calm", "Mildly annoyed", "Frustrated", "Angry"]
  }
}
{
  "anger": {
    "type": "score",
    "score": 1.89,
    "legend": { "0": "Calm", "1": "Mildly annoyed", "2": "Frustrated", "3": "Angry" },
    "probabilities": { "0": 0, "1": 0.11, "2": 0.89, "3": 0 },
    "confidence": 0.89
  }
}

`score` is a decimal

1.89 is a real answer on a four-tier scale. The value is the expected position under the probability distribution, not the index of the winning tier — which is exactly what makes it useful: you can sort a queue by it, or threshold at 2.5, in a way that 2 would not allow. Type it as a float. A schema that declares it an integer will reject valid responses.

To recover the winning tier, take the highest-probability key and look it up in legend:

const top = Object.entries(answer.probabilities).sort(
  (a, b) => b[1] - a[1]
)[0][0];
const label = answer.legend[top]; // "Frustrated"

legend and probabilities are both keyed by the tier index as a string ("0", "1", …), not as a number. Between 2 and 10 tiers, each label up to 400 characters.

Combining Them

Ask everything you need about one state in a single call. The state is read once and charged once no matter how many questions ride along, so the marginal cost of the fourth question is close to zero:

{
  "state": "…one support ticket…",
  "questions": {
    "queue":            { "type": "choice", "instructions": "…", "criteria": { "billing": "…", "technical": "…" } },
    "urgency":          { "type": "score",  "instructions": "…", "criteria": ["Whenever", "This week", "Today", "Now"] },
    "needs_human":      { "type": "noul",   "instructions": "…" },
    "refund_requested": { "type": "noul",   "instructions": "…" }
  }
}

Writing Instructions That Work

  • Describe the decision, not the output. There is no format to ask for.
  • Put the boundary cases in criteria, not in instructions. The label descriptions are where the model learns where one bucket ends.
  • Do not ask for reasoning or an explanation. This model answers with a distribution; there is no prose to produce and asking for it only spends input tokens.
  • Keep one decision per question. "Is it urgent and billing-related?" is two questions that will be cheaper and sharper apart.