# Evaluate state with a Model

> Use ctx.evaluate to answer typed Choice, Score, and Noul questions with a bound evaluation Model.

`ctx.evaluate()` judges one shared state against named questions and returns constrained answers. It is a journaled, priced Model invocation. It does not generate a conversational turn or consume the Run's turn allowance.

## Before you begin {#before-you-begin}

Use `@constal/sdk` 3.5.0 or newer. Bind an evaluation Model that exposes `evaluate`; the platform catalog provides `model/jev` for TypeSafe System One. The Agent's `model` declaration still names a bound Model. The example keeps a generation Model under `model` and selects the evaluation Model through `judge`.

An evaluation needs a string or JSON state and at least one named question. The Model and applicable Policy must allow the call, and the Run budget must cover its price. Agent code receives no provider Credential.

SDK 4.0 validates question structure without imposing question-count, criterion-count, or request-size quotas. State and instructions are preserved in full. The selected provider's limits still apply; its rejection is returned as an invocation error.

## Steps {#steps}

1. Add the evaluation Model to `constal.agent.json` alongside the Agent's default Model:

   ```json constal.agent.json
   {
     "schemaVersion": 2,
     "kind": "agent",
     "id": "triage",
     "namespace": "default",
     "version": "1.0.0",
     "entry": "src/index.ts",
     "mode": "script",
     "bindings": {
       "model": "crn:constal:production:platform:default:model/gpt-5.6-luna",
       "judge": "crn:constal:production:platform:default:model/jev"
     },
     "policies": [],
     "tools": [],
     "limits": { "maxRunMicroUsd": 100000, "maxTurns": 1 }
   }
   ```

2. Ask typed questions in the Agent. `model: "judge"` selects the manifest binding, not the upstream provider identifier:

   ```ts src/index.ts
   import { agent, noulDecision, scoreLevel } from "@constal/sdk";

   export default agent({
     id: "triage", version: "1.0.0", model: "model",
     async onMessage(input, ctx) {
       const message = typeof input === "string" ? input : JSON.stringify(input ?? null);
       const { answers, cost, model } = await ctx.evaluate({
         model: "judge",
         state: { message },
         questions: {
           team: { type: "choice", instructions: "Which team owns this request?",
             criteria: { billing: "Charges, invoices, or refunds", support: "Other support" } },
           severity: { type: "score", instructions: "How severe is the impact?",
             criteria: ["Low", "Degraded", "Customer is blocked"] },
           urgent: { type: "noul", instructions: "Does this require an immediate response?" },
         },
       });
       return { team: answers.team.choice, teamConfidence: answers.team.confidence,
         severity: answers.severity.score, nearestSeverityLevel: scoreLevel(answers.severity),
         urgent: noulDecision(answers.urgent, 0.8), urgencyProbability: answers.urgent.noul,
         model, cost };
     },
   });
   ```

3. Interpret each answer according to its question type:

   | Question | Answer | How to use it |
   | --- | --- | --- |
   | `choice` | `choice`, option `probabilities`, `confidence` | Route by an option key; use confidence to decide whether to request review. |
   | `score` | Fractional `score`, indexed `legend` and `probabilities`, `confidence` | Compare the weighted position with a threshold or call `scoreLevel()` for the nearest level. A three-level scale runs from 0 to 2. |
   | `noul` | `noul` probability from 0 to 1 | Choose an application threshold with `noulDecision()`; there is no separate confidence field. |

   `state`, instructions, and descriptions may contain JSON structure. A Choice needs at least two options; a Score needs at least two ordered levels. The SDK adds no question-count or request-size ceiling. The selected provider enforces its context window; see [Jev's current model limits](https://docs.typesafe.ai/models). Generation options and Tools do not belong in an evaluation request.

To select a model through a bound Gateway, use `gateway: "inference"` and set `model` to an identifier accepted by that Gateway, such as `"typesafe-ai/jev"`. The Gateway must declare the `constal.model-evaluation` capability and expose `evaluate`. The same Run Policy and accounting boundaries apply to Model and Gateway bindings. See [Gateways and Models](/docs/resources/gateways-and-models.md#per-turn-selection).

## Verify {#verify}

Deploy the package and start a Run with a known support message. Inspect the Run journal for one `evaluate` invocation, its selected Model, answers, and settled cost. The evaluation should not add a `turn` entry. Try examples near the thresholds used by your application; probability estimates and routing decisions are separate concerns.

## Next steps {#next-steps}

Use [Build Agents with the SDK](/docs/agents/sdk.md) for the Agent contract, [Add and manage Models](/docs/resources/models.md) for Model bindings, and [Start a Run](/docs/runs/start.md) to invoke the Agent. Evaluation Suites and Scorers assess Agent behavior across cases; see [Evals](/docs/evals.md) for that workflow.
