Evaluate state with a Model
Use ctx.evaluate to answer typed Choice, Score, and Noul questions with a bound evaluation Model.
ctx.evaluate() judges one shared state against named questions and returns constrained answers. It is a journaled, priced Model invocation. It does not generate a conversational turn or consume the Run's turn allowance.
Before you begin
Use @constal/sdk 3.5.0 or newer. Bind an evaluation Model that exposes evaluate; the platform catalog provides model/jev for TypeSafe System One. The Agent's model declaration still names a bound Model. The example keeps a generation Model under model and selects the evaluation Model through judge.
An evaluation needs a string or JSON state and at least one named question. The Model and applicable Policy must allow the call, and the Run budget must cover its price. Agent code receives no provider Credential.
SDK 4.0 validates question structure without imposing question-count, criterion-count, or request-size quotas. State and instructions are preserved in full. The selected provider's limits still apply; its rejection is returned as an invocation error.
Steps
- Add the evaluation Model to
constal.agent.jsonalongside the Agent's default Model:
{
"schemaVersion": 2,
"kind": "agent",
"id": "triage",
"namespace": "default",
"version": "1.0.0",
"entry": "src/index.ts",
"mode": "script",
"bindings": {
"model": "crn:constal:production:platform:default:model/gpt-5.6-luna",
"judge": "crn:constal:production:platform:default:model/jev"
},
"policies": [],
"tools": [],
"limits": { "maxRunMicroUsd": 100000, "maxTurns": 1 }
}- Ask typed questions in the Agent.
model: "judge"selects the manifest binding, not the upstream provider identifier:
import { agent, noulDecision, scoreLevel } from "@constal/sdk";
export default agent({
id: "triage", version: "1.0.0", model: "model",
async onMessage(input, ctx) {
const message = typeof input === "string" ? input : JSON.stringify(input ?? null);
const { answers, cost, model } = await ctx.evaluate({
model: "judge",
state: { message },
questions: {
team: { type: "choice", instructions: "Which team owns this request?",
criteria: { billing: "Charges, invoices, or refunds", support: "Other support" } },
severity: { type: "score", instructions: "How severe is the impact?",
criteria: ["Low", "Degraded", "Customer is blocked"] },
urgent: { type: "noul", instructions: "Does this require an immediate response?" },
},
});
return { team: answers.team.choice, teamConfidence: answers.team.confidence,
severity: answers.severity.score, nearestSeverityLevel: scoreLevel(answers.severity),
urgent: noulDecision(answers.urgent, 0.8), urgencyProbability: answers.urgent.noul,
model, cost };
},
});- Interpret each answer according to its question type:
| Question | Answer | How to use it |
|---|---|---|
choice | choice, option probabilities, confidence | Route by an option key; use confidence to decide whether to request review. |
score | Fractional score, indexed legend and probabilities, confidence | Compare the weighted position with a threshold or call scoreLevel() for the nearest level. A three-level scale runs from 0 to 2. |
noul | noul probability from 0 to 1 | Choose an application threshold with noulDecision(); there is no separate confidence field. |
state, instructions, and descriptions may contain JSON structure. A Choice needs at least two options; a Score needs at least two ordered levels. The SDK adds no question-count or request-size ceiling. The selected provider enforces its context window; see Jev's current model limits. Generation options and Tools do not belong in an evaluation request.
To select a model through a bound Gateway, use gateway: "inference" and set model to an identifier accepted by that Gateway, such as "typesafe-ai/jev". The Gateway must declare the constal.model-evaluation capability and expose evaluate. The same Run Policy and accounting boundaries apply to Model and Gateway bindings. See Gateways and Models.
Verify
Deploy the package and start a Run with a known support message. Inspect the Run journal for one evaluate invocation, its selected Model, answers, and settled cost. The evaluation should not add a turn entry. Try examples near the thresholds used by your application; probability estimates and routing decisions are separate concerns.
Next steps
Use Build Agents with the SDK for the Agent contract, Add and manage Models for Model bindings, and Start a Run to invoke the Agent. Evaluation Suites and Scorers assess Agent behavior across cases; see Evals for that workflow.