Evals and training API

Create versioned evaluation inputs, operate durable suites, compare baselines, and start metered SFT or RL jobs.

All Evals routes are namespace-scoped and use normal Platform authentication, authorization, request IDs, and JSON error responses. Dataset and Scorer references are immutable triples of crn, version, and hash. Suite and training IDs are caller-stable idempotency identities.

Eval door Policy

text
GET /v1/namespaces/{namespace}/evals/policy
PUT /v1/namespaces/{namespace}/evals/policy

The door Policy owns live sampling rules and daily Judge budgets, the Suite concurrency and spend ceilings, verdict retention, and an optional evalResourcePolicy CRN. A Suite resolves that CRN once, pins the exact Policy hash into its Resource snapshot, and applies it as an additional boundary to every Tool call, Resource invocation, and Credential use. This is where operators prohibit production side effects or allow only evaluation doubles without creating an eval-specific execution path.

PUT accepts the Policy directly or under policy, plus expectedHash for compare-and-set updates. evalResourcePolicy is either null or a deployed Policy in the same tenant and namespace. The returned snapshot has a monotonically increasing revision and content hash.

Datasets and Scorers

text
POST   /v1/namespaces/{namespace}/evals/datasets/{dataset}/draft
GET    /v1/namespaces/{namespace}/evals/datasets/{dataset}/draft
GET    /v1/namespaces/{namespace}/evals/datasets/{dataset}/cases
POST   /v1/namespaces/{namespace}/evals/datasets/{dataset}/cases
DELETE /v1/namespaces/{namespace}/evals/datasets/{dataset}/cases/{case}
POST   /v1/namespaces/{namespace}/evals/datasets/{dataset}/capture
POST   /v1/namespaces/{namespace}/evals/datasets/{dataset}/publish
POST   /v1/namespaces/{namespace}/evals/scorers

Published Datasets and Scorers are ordinary Resources and are read through the namespace Resource routes. Draft cases support bounded keyset pagination with after and limit.

The capture body has one of two forms. Use source for one exact Run locator, or export for an immutable raw analytics export:

json
{ "source": { "agent": "support", "session": "customer-42", "run": "RUN_ID" }, "tags": { "source": "incident" } }
json
{ "export": "ANALYTICS_EXPORT_SHA256", "maximumRuns": 1000, "tags": { "source": "failed-production-run" } }

Bulk capture deduplicates locators and returns truncated when either the export or requested maximum is incomplete. Captures are idempotent for the Dataset, export, and exact Run identity. Dataset publication remains a separate request.

Suites, comparison, and baselines

text
GET  /v1/namespaces/{namespace}/evals/suites?after={cursor}&limit={count}
POST /v1/namespaces/{namespace}/evals/suites
GET  /v1/namespaces/{namespace}/evals/suites/{suite}
POST /v1/namespaces/{namespace}/evals/suites/{suite}/cancel
POST /v1/namespaces/{namespace}/evals/compare
POST /v1/namespaces/{namespace}/evals/baseline

Starting a Suite returns 202 with its durable identity unless it has already completed. Read the job route for authoritative progress and result. Comparison and baseline bodies use content-addressed Suite result hashes, not the mutable job row.

The start body sets mode to fresh, replay, or shadow. Replay and Shadow cases must have been captured from completed Runs. onDivergence may be recorded (default) or live; it applies only to Replay and the live option applies only to changed Model requests. A Replay result includes divergencesRef, an immutable manifest of request hashes, bounded diffs, effect classes, and resolutions. Divergent effectful non-Model operations fail closed.

A Shadow Suite validates each production source, forks the candidate at the captured pre-run head, and scores both outputs with the same pinned Scorers. Its compare value points to the immutable production-side Suite artifact and contains paired per-Scorer deltas and significance flags. Mixed production subject identities are rejected rather than averaged together.

When calibration names one of the pinned Judge Scorers, every Dataset case must have a human expected.score and may include expected.pass. The Suite result includes exact and within-tolerance agreement, mean absolute error, correlation, and Cohen's kappa. Missing labels or Judge verdicts fail the Suite rather than silently shrinking the sample.

Continuous schedules

text
GET    /v1/namespaces/{namespace}/evals/schedules
GET    /v1/namespaces/{namespace}/evals/schedules/{schedule}
PUT    /v1/namespaces/{namespace}/evals/schedules/{schedule}
DELETE /v1/namespaces/{namespace}/evals/schedules/{schedule}
POST   /v1/namespaces/{namespace}/evals/schedules/{schedule}/run

The PUT body contains state, intervalMs, a Suite request without suiteId, training, and expectedHash. Set training to null for evaluation only. Otherwise it contains a complete pinned training recipe, an output Model prefix, and a gate with minimum completed cases, maximum errors, and optional aggregate Scorer thresholds. Use a null expectedHash only to create a new Schedule; use the returned hash for later compare-and-set changes. DELETE accepts expectedHash in JSON or If-Match. The run endpoint accepts eventId, also available through Idempotency-Key, and starts one deterministic immediate Suite without shifting the cadence.

A qualifying successful Suite starts at most one training job per exact Schedule revision. lastTraining exposes the qualifying Suite fact, derived job and output IDs, Agent run, attempts, state, and skip or failure reason. Launch retries keep the same deterministic identity. Training uses the normal eval Agent, training Resource, Policy authorization, budget reservation, measured usage settlement, and billing path. A successful job produces candidate weights only; the API never promotes or binds them implicitly.

Training

text
GET  /v1/namespaces/{namespace}/evals/training-providers
GET  /v1/namespaces/{namespace}/evals/training-models?provider={crn}&hash={hash}
GET  /v1/namespaces/{namespace}/evals/training-jobs?after={cursor}&limit={count}
POST /v1/namespaces/{namespace}/evals/training-jobs
GET  /v1/namespaces/{namespace}/evals/training-jobs/{job}
POST /v1/namespaces/{namespace}/evals/training-jobs/{job}/cancel
GET  /v1/namespaces/{namespace}/evals/training-jobs/{job}/checkpoint-gateways?step={step}&ref={ref}
POST /v1/namespaces/{namespace}/evals/training-jobs/{job}/models

The provider collection lists installed Connection Resources that implement the exact Training Provider capability. Select one by CRN and immutable hash before reading its model catalog. The catalog includes provider model identifiers, context windows, supported methods and checkpoint behavior, and the per-million-token price snapshot used for settlement. A training start body carries the same provider {crn, hash} pin.

A training status includes the exact provider identity, phase, method, base and output Model IDs, step and epoch progress, exact usage units, settled micro-USD cost, retry state, terminal error, checkpoint format and creation time, and final fact when available.

After success, the checkpoint Gateway route returns only installed exact Gateway revisions compatible with the selected checkpoint format. The Model creation body is:

json
{
  "id": "support-model-v1",
  "displayName": "Support Model v1",
  "temperature": 0.2,
  "gateway": { "crn": "GATEWAY_CRN", "hash": "GATEWAY_HASH" },
  "checkpoint": { "step": 400, "ref": "CHECKPOINT_CONTENT_HASH" }
}

Creation may return 202 while the Gateway import reconciles. Retrying the same body resolves the same deterministic deployment identity. On completion, the platform registers an ordinary Model whose immutable source records the provider, job, base model, and checkpoint. It never mutates an Agent binding.

Request discipline

Send the same job ID and body when retrying an uncertain start request. Do not create a second ID after a timeout. Treat queued, running, and evaluating as non-terminal. Bound list requests and follow the returned continuation cursor. Keep the API key outside request bodies and logs; every referenced Resource still passes the caller's Policy checks.