Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,32 @@

Every version, newest first. Each entry is the description of the pull request that made the change, which is written as a change note when the change is made.

## v1.63 (2026-09-23)

### Feature previews is its own settings section (#343)

The voice preview switch sat at the foot of **External access**, below the retrieval endpoint and the MCP configuration. A deployment reported the preview missing while its own bundle carried the control; it was findability, not a deploy gap.

Feature previews is now its own section in the settings rail, between Assistant tools and Extensions. The rail is extracted as `settingsSections()`, a pure function with a test, so a new section has to be placed deliberately. The voice text now also states that the Unmute deployment must be reachable from the browser in use, since a closed network needs its own rather than one running elsewhere.

## v1.62 (2026-09-15)

### The model listing answers from a cache and refreshes behind it (#342)

Every load of the model list fetched the gateway catalog and pinged each provider family (about a second on dev, several seconds where the gateway and providers are further away and each user's credential probes on its own), and the client asked three times per mount.

**Server:** the computed listing is cached per viewer and credential (the listing is shaped by the viewer, since shared providers are labeled with their owner). A load inside the 60 s fresh window is answered from cache; an older one is answered from cache and refreshed in the background (deduplicated), so the next load is current; only the first load per credential waits. `?refresh=1` drops the probes and recomputes before answering, and the key test in Settings clears every listing. The last-call-failed decoration is applied after the cache so it stays live.

**Client:** one deduplicated request (`fetchModels` in `api.ts`) serves the picker, the banner, the credential notice, and the fleet page, with a 15 s reuse. Freshness without a reload: refetch when the tab becomes visible and every five minutes while visible; the picker is remounted only when the set of marked models changes, never mid-reply.

Measured on dev: first load 1.2 s, next three under 10 ms, forced refresh 0.4 s. Tests cover cached loads, refresh seeing a changed catalog, and the live decoration. 187 server and 12 web tests pass.

## v1.61 (2026-09-15)

### Neutral provider names in the probe fixtures (#341)

The two new probe cases used a real provider prefix and its model ids as fixture names. Fixtures use neutral names like the rest of the suite.

## v1.60 (2026-09-15)

### A change log and release notes generated from pull request descriptions (#339)
Expand Down
77 changes: 15 additions & 62 deletions pnpm-lock.yaml

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

30 changes: 29 additions & 1 deletion server/src/chat/gateway.ts
Original file line number Diff line number Diff line change
Expand Up @@ -210,6 +210,18 @@ export function invalidateProviderProbes(): void { providerProbes.clear() }

export function probeableProvider(prefix: string): boolean { return PROBE_PATTERN.test(prefix) }

/** The provider's own sentence from an error body, for showing a person. */
export function providerReason(text: string): string {
let t = text
for (let i = 0; i < 3; i++) {
const m = /"message"\s*:\s*"((?:[^"\\]|\\.)*)"/.exec(t)
if (!m) break
t = m[1].replace(/\\"/g, '"').replace(/\\\\/g, '\\')
}
t = t.replace(/^received error while streaming:\s*/i, '').replace(/\s+/g, ' ').trim()
return t.length > 180 ? t.slice(0, 177) + '...' : t
}

export function extractUnlockUrl(text: string): string | null {
return /unlock_url[\\":\s]*(https?:\/\/[^"\\\s]+)/.exec(text)?.[1] ?? null
}
Expand Down Expand Up @@ -426,8 +438,24 @@ export async function streamTurn(
const reader = res.body.getReader()
const decoder = new TextDecoder()
let buf = ''
// A provider can accept a request, open the stream, and send nothing.
// Only the client's own abort ended that, so a reply could sit on its
// spinner indefinitely. Each read now has a limit, reset by every chunk
// and generous enough for providers that buffer a whole answer (about a
// minute is common) or a model that thinks before its first token.
const idleMs = Number(process.env.STUDIO_STREAM_IDLE_MS) || 300_000
const readWithin = async () => {
let timer: NodeJS.Timeout | undefined
const idle = new Promise<never>((_, reject) => {
timer = setTimeout(() => {
reader.cancel().catch(() => {})
reject(new Error(`The model sent nothing for ${Math.round(idleMs / 1000)} seconds, so the reply was stopped. Try again, or pick another model.`))
}, idleMs)
})
try { return await Promise.race([reader.read(), idle]) } finally { clearTimeout(timer) }
}
for (;;) {
const { done, value } = await reader.read()
const { done, value } = await readWithin()
if (done) break
buf += decoder.decode(value, { stream: true })
let nl
Expand Down
13 changes: 7 additions & 6 deletions server/src/chat/routes.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ import { suggestNext } from './suggest.js'
import { getConversation } from '../conversations.js'
import { listRuns } from '../runs.js'
import { GATEWAY_BASE, MAX_TOOL_ITERATIONS, SIDECAR_ENDPOINT, INDEX_BASE } from '../config.js'
import { providerReason } from './gateway.js'
import { aiHealth, gatewayConfigured, isSidecarModel, listModels, listSidecarModels, modelFailure, probeProvider, probeableProvider, extractUnlockUrl, invalidateProviderProbes, sidecarConfigured, sidecarTarget, StreamedTurnError, streamTurn, WireMessage, WireToolCall, type TokenUsage } from './gateway.js'
import { TOOL_CALLS, TOOL_SPECS, activeToolSpecs, activeToolSpecsWithRemote, commandFor, customToolSpecs, executeTool, expandSlashCommand, skillToolSpec } from './tools.js'
import { systemPrompt } from './context.js'
Expand Down Expand Up @@ -300,15 +301,15 @@ export function stripAvailabilityMark(id: string): string {
*/
export function markImpaired<M extends { id: string; name?: string }>(
models: M[],
verdicts: Map<string, { ok: boolean; kind: 'locked' | 'unavailable' | null; unlockUrl: string | null }>,
): { models: (M | (M & { callable: false; locked: boolean; unavailable: true; unlock_url: string | null }))[]; impaired: { id: string; locked: boolean; unlock_url: string | null }[] } {
const impaired: { id: string; locked: boolean; unlock_url: string | null }[] = []
verdicts: Map<string, { ok: boolean; kind: 'locked' | 'unavailable' | null; unlockUrl: string | null; message?: string }>,
): { models: (M | (M & { callable: false; locked: boolean; unavailable: true; unlock_url: string | null }))[]; impaired: { id: string; locked: boolean; reason?: string; unlock_url: string | null }[] } {
const impaired: { id: string; locked: boolean; reason?: string; unlock_url: string | null }[] = []
const out = models.map(m => {
const id = String(m.id)
const v = verdicts.get(id.slice(0, Math.max(id.indexOf('/'), 0)))
if (!v || v.ok) return m
const tag = v.kind === 'locked' ? 'locked' : 'unavailable'
impaired.push({ id, locked: v.kind === 'locked', unlock_url: v.unlockUrl })
impaired.push({ id, locked: v.kind === 'locked', ...(v.message ? { reason: providerReason(v.message) } : {}), unlock_url: v.unlockUrl })
const name = typeof m.name === 'string' && m.name.trim() ? `${m.name.trim()} [${tag}]` : m.name
return { ...m, id: `${id} [${tag}]`, name, callable: false as const, locked: v.kind === 'locked', unavailable: true as const, unlock_url: v.unlockUrl }
})
Expand Down Expand Up @@ -422,7 +423,7 @@ export async function chatRoutes(app: FastifyInstance): Promise<void> {
// "look again": it drops the probes and recomputes before answering.
// Per-request decorations (a model whose last call failed) are applied
// after the cache, so they stay live.
type Listing = { models: any[]; impaired: { id: string; locked: boolean; unlock_url: string | null }[]; unreachableSessions: unknown[] }
type Listing = { models: any[]; impaired: { id: string; locked: boolean; reason?: string; unlock_url: string | null }[]; unreachableSessions: unknown[] }
const LISTING_FRESH_MS = 60_000
const LISTING_KEEP_MS = 30 * 60_000
const listings = new Map<string, { at: number; v: Listing }>()
Expand Down Expand Up @@ -554,7 +555,7 @@ export async function chatRoutes(app: FastifyInstance): Promise<void> {
if (probeableProvider(prefix) && !prefixes.has(prefix)) prefixes.set(prefix, id)
}
}
const verdicts = new Map<string, { ok: boolean; kind: 'locked' | 'unavailable' | null; unlockUrl: string | null }>()
const verdicts = new Map<string, { ok: boolean; kind: 'locked' | 'unavailable' | null; unlockUrl: string | null; message?: string }>()
if (prefixes.size) {
await Promise.race([
Promise.allSettled([...prefixes.entries()].map(async ([prefix, sample]) => {
Expand Down
15 changes: 15 additions & 0 deletions server/test/providerProbe.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -100,3 +100,18 @@ test('a healthy stream is read only to its first frame', async () => {
return { ok: r.status < 400, status: r.status, text: async () => r.text }
}
})

test('the provider reason is its own sentence, unwrapped from nested error bodies', async () => {
const { providerReason } = await import('../dist/chat/gateway.js')
const body = '{"error":{"message":"received error while streaming: {\\"message\\": \\"Requested model is not available and no compliant same-model variant was found.\\", \\"type\\": \\"invalid_request_error\\"}","type":"error"}}'
assert.equal(providerReason(body), 'Requested model is not available and no compliant same-model variant was found.')
assert.equal(providerReason('plain text'), 'plain text')
})

test('a marked model carries the reason to the client', async () => {
const { markImpaired } = await import('../dist/chat/routes.js')
const verdicts = new Map([['me:vega', { ok: false, kind: 'unavailable', unlockUrl: null, message: '{"error":{"message":"model not served","type":"error"}}' }]])
const r = markImpaired([{ id: 'me:vega/model-a' }], verdicts)
assert.equal(r.impaired[0].reason, 'model not served')
assert.equal(r.impaired[0].locked, false)
})
4 changes: 2 additions & 2 deletions web/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,8 @@
"dependencies": {
"@fontsource-variable/geist": "^5.3.0",
"@fontsource-variable/geist-mono": "^5.3.0",
"@parallelworks/ai-chat": "0.4.5",
"@parallelworks/ui": "^0.5.0",
"@parallelworks/ai-chat": "0.5.0",
"@parallelworks/ui": "^0.16.0",
"react": "^19.2.8",
"react-dom": "^19.2.8",
"streamdown": "^2.5.0",
Expand Down
10 changes: 8 additions & 2 deletions web/src/adapter.ts
Original file line number Diff line number Diff line change
Expand Up @@ -180,10 +180,16 @@ export function createStudioAdapter(): ChatAdapter {
location.hash = `#open=file:${encodeURIComponent(rel).replace(/%2F/gi, '/')}`
} catch { /* malformed id: ignore */ }
},
list: async ({ limit, offset }: { limit: number; offset: number }) => {
// ai-chat 0.5 pages by an opaque cursor and asks whether there is
// more; the server pages by offset, so the cursor is the next offset.
list: async ({ limit, cursor }: { limit: number; cursor?: string }) => {
const offset = Number(cursor) || 0
const res = await fetch(`/api/chat/attachments?limit=${limit}&offset=${offset}`)
if (!res.ok) throw new Error(`attachments: ${res.status}`)
return res.json()
const page = await res.json() as { attachments: unknown[]; total: number }
const next = offset + (page.attachments?.length ?? 0)
const hasMore = next < (page.total ?? 0)
return { ...page, hasMore, ...(hasMore ? { nextCursor: String(next) } : {}) }
},
upload: async (file, _conversationId) => {
const fd = new FormData()
Expand Down
2 changes: 1 addition & 1 deletion web/src/api.ts
Original file line number Diff line number Diff line change
Expand Up @@ -229,7 +229,7 @@ export interface IndexJob {
// client asking three times for one answer.
export interface ModelsResponse {
models: { id: string; name?: string; callable?: boolean; [k: string]: unknown }[]
impaired?: { id: string; locked: boolean; unlock_url: string | null }[]
impaired?: { id: string; locked: boolean; reason?: string; unlock_url: string | null }[]
unreachableSessions?: unknown[]
error?: string
credential?: string
Expand Down
Loading
Loading