Gateways and Models
Understand how connection Resources, logical Models, Credentials, scoped bindings, and pricing work together.
A Gateway is not a separate execution system. It is an ordinary invokable Resource marked as a connection. That Resource owns the external endpoint, Credential bindings, network rules, lifecycle, and the capabilities it implements. It continues to use the normal Resource deployment, Policy, pinning, invocation, and journal machinery.
A Model is a reusable configuration for inference. It selects a provider model ID, default parameters, technical limits, accounting references, and either a fixed or scoped Gateway. Use a Model binding for a named preset, or bind a Gateway directly when the Agent should choose its model while working. Neither choice exposes the Gateway's Credentials to the Agent.
Preset: Agent → Model → Gateway → provider
Per turn: Agent → Gateway + selected model → providerWhy the split exists
One Gateway can expose several Models, and one logical Model can resolve to different Gateways for different customers. A Model preset provides a stable configuration name; a direct Gateway binding leaves model selection to the Agent. Journals retain the called Resource, selected model, and usage in either case.
Gateways retain their concrete Resource kind because the kind describes how they operate. The connection role describes how people use them. This avoids a parallel Gateway registry or a second invocation path. A normal service Resource is not a Gateway unless it explicitly has the connection role.
The platform-provided Constal Gateway
Constal Gateway is a platform-catalog, platform-managed, platform-paid connection projected into every tenant. Tenants do not provide a Credential for this Gateway and Constal does not create or update one copy per tenant. Use the platform Models made available through it when Constal should operate upstream inference and apply platform pricing and limits through Policy.
The Gateway and default Model documents keep platform catalog CRNs under the reserved platform ownership boundary. Tenant authorization, accepted Run snapshots, usage, and billing remain tenant-scoped when that shared catalog capability is invoked. Tenants may read the projected entries but cannot mutate or delete platform-managed Resources.
Tenant-managed Gateways remain the bring-your-own alternative. They use the same Resource abstraction, but the tenant supplies the endpoint and Credential and is normally responsible for upstream charges.
Fixed and scoped Gateways
| Gateway binding | Use it when | Resolution |
|---|---|---|
| Fixed | Every invocation of the Model uses the same managed or tenant-owned connection | The Model pins one exact Gateway CRN, document hash, and capability contract |
| Tenant scoped | Each tenant needs a different connection, with a tenant default | Admission resolves the authenticated tenant before a Run starts |
| Customer scoped | A shared Agent serves downstream customers that bring their own provider account | Admission resolves the verified customer identity supplied by the Channel |
| Principal scoped | Individual users authorize their own provider account | Admission resolves the verified subject identity supplied by the Channel |
A scoped binding is a lookup rule, not a late-bound secret. Admission resolves it to one exact Gateway and Credential set, records the assignment revision and owner, and pins that evidence into the Run snapshot. A missing required assignment fails before provider dispatch. An existing Run keeps its accepted snapshot if an assignment changes later.
The Model requires a stable capability contract such as “model completion v1.” Any Gateway used by the binding must declare that it implements the contract. Provider-specific operation catalogs may differ internally, but they cannot change what the Model promises to the Agent.
Models use Gateways
Install a Gateway first. You can then add a logical Model as a reusable preset, or bind the Gateway directly. Start with Add and manage Gateways.
Choose a preset or select a model per turn
With Model bindings, model names a binding in the Agent's deployment:
await ctx.turn({
system: "You investigate production incidents.",
objective: incident,
model: "investigator",
maxOutputTokens: 8192,
temperature: 0.2,
effort: "high",
});With a Gateway binding, gateway names the binding and model is the identifier accepted by that Gateway. No separate Model installation is required:
await ctx.turn({
system: "You investigate production incidents.",
objective: incident,
gateway: "inference",
model: "openai/gpt-5.6-luna",
effort: "high",
});The Agent can choose those values from repository configuration or the current task. Both forms use the same authorization, Credential resolution, durable journal, usage accounting, and recovery. A tenant-paid Gateway records usage and available provider cost without adding a platform inference charge.
maxOutputTokens, temperature, and effort are optional per-turn choices. Model presets and applicable Policy can limit generation. Omitted effort is not sent to the provider. Supported effort values depend on the selected model; Constal does not silently replace an unsupported effort with another value.
Apply Policy to model use
Attach executable Policy to the Gateway to cover every model reached through it, including Models that use that Gateway. The Policy receives the authenticated principal, Agent identity, target Resource, selected model, and invocation arguments. A Policy attached only to one Model does not restrict other Models or other Gateways.
import { parseResourceName, policy } from "@constal/sdk";
export default policy({
id: "model-access",
version: "1",
evaluate(input) {
const invocation = input.invocation;
if (input.action !== "resource:invoke" || invocation?.operation.op !== "complete") {
return { kind: "allow" };
}
const args = invocation.arguments as { model?: string; effort?: string };
if (!input.principal.roles.includes("researcher") && args.effort === "max") {
return { kind: "deny", code: "ResearchRoleRequired" };
}
const agent = parseResourceName(invocation.run.agent).path;
if (agent === "triage" && args.model !== "openai/gpt-5.6-luna") {
return { kind: "deny", code: "ModelNotAllowed" };
}
return { kind: "allow" };
},
});Use the existing modelControls() and modelLimit() helpers for token, temperature, request-rate, and concurrent-use limits. Their optional second argument is the target's pinned governance contract reference: omit it for a Model, or pass the Gateway's contract for a directly bound Gateway. Limits can use the existing tenant, customer, subject, Run, Session, Resource, or binding scopes. No separate model authorization system is involved.
Credentials and ownership
The Gateway references Credentials; the Model does not contain provider secrets. At invocation time Constal authorizes the Gateway and, when used, the Model preset. The integration receives the authorized Credential material. Agent code receives the result, never reusable secret material.
A platform-managed Gateway normally uses a platform-paid provider account. A tenant-managed Gateway is the bring-your-own path and normally uses tenant-paid Credentials. The Gateway records who manages the connection and who pays the upstream provider so the runtime can reject contradictory pricing configuration.
Use a Credential Provider to import, mint, renew, or authorize the secret material. Use a scoped binding when the correct Gateway or Credential depends on tenant, customer, or principal identity.
Pricing and rate limits
Model accounting keeps three values separate:
| Accounting lane | Meaning | Authority |
|---|---|---|
| Tracked cost | The estimated or reported upstream cost used for cost and margin analysis | Provider or other trusted cost source |
| Platform charge | What Constal charges the tenant and what consumes the Run budget | Platform Policy |
| Customer charge | What the tenant may charge its downstream customer | Tenant Policy |
The rates themselves are accepted through runtime Policy, not trusted merely because a Model document mentions them. A Model pins immutable table identities so the runtime can prove it applied the intended rates. The same logical model may therefore have a different platform charge or limit policy for each tenant without changing Agent code or dispatch behavior.
For a platform-paid Gateway, Policy must provide a platform charge. For a tenant-paid Gateway, the platform charge may be absent, so the invocation consumes no Constal model budget while tracked cost or a tenant-owned customer charge can still be recorded. Policy can independently cap requests, input tokens, output tokens, and concurrent requests for the Model.
Invocation sequence
- Admission resolves the Agent's Model or Gateway bindings.
- For a Model preset, it resolves that Model's fixed or scoped Gateway.
- It verifies the Gateway implements model completion and resolves Credential bindings under the authenticated identity.
- It pins Resource configuration, assignments, Policy, and pricing into the Run snapshot.
- Each turn selects its model and generation options. Policy evaluates that call against the Gateway and any Model preset involved.
- The Gateway's integration performs the request. The journal records the called Resource, selected model, and options.
- Usage settles into tracked cost, platform charge, and customer charge independently.
This preserves one Resource abstraction end to end. Gateways use the same identity, Policy, binding, invocation, and accounting contracts as every other Resource.
What to verify
In Resources → Gateways, confirm the Credential relationship, upstream billing owner, and exact CRN. For a Model preset, also check its model ID and Gateway. Start a Run and inspect its bindings and invocation journal: a preset calls the Model and records its resolved Gateway; a direct selection calls the Gateway and records the chosen model in the invocation arguments.