@intelligo-dev/chat
The AI-SDK-native chat transport: one Route Handler with auth, rate limits, feature gates, credits and persistence.
Install
pnpm add @intelligo-dev/chat aiai (the Vercel AI SDK) is a peer: the transport takes its tools, its models
and its stop conditions natively, and never wraps them.
Use
// app/api/chat/route.ts
import { createChatHandler } from "@intelligo-dev/chat";
import { chatServerConfig } from "@/lib/chat-server-config";
export const { POST, DELETE } = createChatHandler(chatServerConfig);// lib/chat-server-config.ts
import type { ChatServerConfig } from "@intelligo-dev/chat";
import { composeIntelligo, executions } from "@/lib/intelligo";
export const chatServerConfig: ChatServerConfig = {
executions,
onRequest: composeIntelligo,
model: { defaultId: "google/gemini-2.5-flash", resolve: getChatModel },
agent: { id: "assistant", systemPrompt: "You are a helpful assistant." },
};That is a working chat. Every turn goes through auth, the plan’s rate limit,
the feature gate, conversation persistence (@intelligo-dev/core/conversations)
and the execution boundary (executions.begin() decides entitlement against
the model that is about to run, and settles exactly once).
The seams
Each optional field is something a real product needed and used to fork the route to get:
| Field | What it decides |
|---|---|
resolveAgent(turn) |
Which agent runs — from the body, the row, a table; prompt, tools, model |
streamTurn(turn, prepared) |
Another runtime than streamText — a Mastra agent, an eve session — returning the AI SDK’s chunks and the run’s usage; everything else stays the transport’s |
models |
The models a request may pick; anything else is FEATURE_GATED (model_not_allowed) |
prepareMessages(turn, msgs) |
What the model is shown — windowing, summaries, injected context |
agent.generation |
How the model samples — temperature, a token ceiling, a tool choice, a seed: an allowlist of the streamText options that do not touch settlement. maxOutputTokens defaults to the registered model’s, the figure admission held |
attachments |
Which file parts are accepted; mode: "stored" uploads them through the storage port and signs URLs for the model only |
reasoning, sources |
Whether reasoning and source parts stream to the client |
messageMetadata |
{ modelId, usage, finishedAt } on the reply (default on) |
cors |
Origins an embedded widget may call from |
deriveTitle |
The conversation’s title, sync or model-written |
persist |
Where the turn’s messages go |
onTurn |
Telemetry: start, complete, fail, refuse, approval, feedback |
messages(request) |
Refusal copy in the caller’s locale |
authenticate, rateLimit |
The defaults are the framework’s; a worker or a test can replace them |
None carry product vocabulary; all close over the caller’s tenancy, so the model is never told which workspace it is in.
resolveAgent, agent.tools(turn) and deriveTitle receive a ChatTurnContext; the later
seams — prepareMessages, streamTurn, persist — receive a ChatTurn, which
is the same context plus the resolved agent and history(). The context
carries workspaceId and userId, the request, the conversationId and its
row (conversation, null on the first turn), the trigger, and body: the
fields the client sent beyond the AI SDK’s own, such as ChatPanel’s body
prop, which is how a page tells a tool which record it is about. body is the
caller’s input, so a tool reads a record through the turn’s workspace, never
by the id alone. Beside them sit write, updateMetadata, state and
addUsage:
agent: {
tools: (turn) => ({
summarize: tool({
inputSchema: z.object({ section: z.string() }),
execute: async ({ section }) => {
// Scoped to the caller's workspace: an id from `body` is untrusted.
const record = await getRecord(turn.workspaceId, String(turn.body.recordId));
if (!record) return { error: "not found" };
const { text, usage } = await summarize(record, section);
turn.addUsage(usage, { model: "google/gemini-2.5-flash" });
return text;
},
}),
}),
},A tool reaches the client mid-turn through the turn: turn.write() sends a
data-chat-* part (a status line, a plan), and createArtifactWriter(turn, { kind, title }) streams a document into the chat’s canvas — append
deltas, finish({ documentId }). A tool that calls a model itself — a
grounded search, a sub-agent, an embedding for retrieval — hands its tokens
to turn.addUsage(usage, { model }). Tokens on the turn’s own model (or
with no model) settle with the turn’s; tokens on another registered model
are summed per model and recorded as that model’s own execution
(capability "chat.embedding" for an embedding model, parentExecutionId
in its metadata) when the turn settles. An unregistered model throws
UnknownModelError at the call rather than bill at another model’s price.
turn.state is a Map that lives for one request and is shared by every
seam that receives the turn — resolveAgent, prepareMessages, tools,
streamTurn, persist, the onTurn events — so what one of them reads
(a profile, a retrieval result) the next can use without a second query.
sanitizeForShare strips a transcript
for a public page; recordChatFeedback records a vote and tells the hook.
Stored attachments mount two more handlers, createChatUploadHandler and
createChatAttachmentHandler.
Refusals
Every non-2xx answer is { error, code, reasonCode? } with one status per
code (CHAT_ERROR_CODES, parseChatError). Admission’s refusals split by
who can fix them:
| Admission code | Answer | Whose problem |
|---|---|---|
insufficient_credits, allowance_depleted |
402 QUOTA_EXCEEDED |
The workspace’s: upgrade or top up |
billing_not_configured |
503 BILLING_NOT_CONFIGURED |
The deployment’s: no product or plans |
unknown_model |
503 MODEL_UNAVAILABLE |
The deployment’s: the model id has no price |
The two 503s are logged at error level — unknown_model with the model id
— and MODEL_UNAVAILABLE answers with neutral copy (modelUnavailable)
rather than the engine’s reason, which names the registry. reasonCode
carries the admission code either way.
getChatQuotaState({ modelId }) is the same decision as a read, for a
page that wants to say so before a turn is spent; getChatQuotaStates({ modelIds }) reads it once per model a picker offers.
Subpaths
@intelligo-dev/chat/client— error codes, the quota-state shape,parseChatError, thedata-chat-*parts vocabulary (ChatUIMessage,isChatDataPart) andChatModelOption, for a client bundle. Imports nothing at runtime.@intelligo-dev/chat/testing—createStubLanguageModel, a deterministic model that streams with no API key.
Entry points
@intelligo-dev/chat@intelligo-dev/chat/client@intelligo-dev/chat/testing