The requirement
Lambda Feedback requires SSE streaming, but /evaluate and /chat only document 200 → application/json. ChatCapabilities.supportsStreaming already exists on /chat/health, and LLMConfiguration.stream exists on the request side, so streaming is acknowledged as in scope, but the spec doesn't define what a streaming response actually looks like. Without that, every implementation invents its own frame format and clients have nothing to code against.
(A separate PR already extends callbackUrl/202 Accepted to /chat, so this issue treats async delivery and SSE as parallel, complementary transport options rather than something to resolve here.)
Why this needs to be in the spec
/chat and /evaluate need to stream identically, and future endpoints shouldn't have to reinvent it.
supportsStreaming and LLMConfiguration.stream are meaningless without a documented contract behind them.
- A streaming contract is largely generic across operations. The endpoint-specific parts are the terminal payload shape (
Feedback[] for evaluate, ChatResponse for chat) and any progress-stage vocabulary; everything else (the opt-in mechanism, keep-alives, error delivery) can be defined once.
Requirements
- Opt-in via
Accept: text/event-stream. Any operation may support it. When set, the response is a sequence of SseProgressStep frames, optional keep-alive comments, and exactly one terminal frame. X-Request-Id and X-Api-Version are set once, as SSE headers, at stream open, since there's no per-frame HTTP header mechanism.
- Shared progress schema. One
SseProgressStep type (stage, message, timestamp), reused by both endpoints. Stage values are informative, not a strict enum, so implementations can add stages without a spec change; each service's /health response advertises its own supported stage values, the same way supportedArtefactProfiles advertises formats.
- Terminal frame reuses the existing success/error shape. On success, the terminal frame's payload is exactly the operation's normal
200 body (Feedback[] for /evaluate, ChatResponse for /chat), composed via allOf against each operation's existing 200 schema rather than duplicated. On failure, the payload is the existing ErrorResponse schema, so SSE errors carry the same title/code/details shape clients already handle for 400/500/etc.
- HTTP status stays
200 for the life of the stream. Failures are always delivered as a terminal ErrorResponse frame, not an HTTP error status.
configuration.llm.stream is distinct from Accept: text/event-stream. The former, if kept, governs token-level streaming from the LLM provider; the latter governs step-level progress streaming at the API layer. The spec should say so explicitly to avoid the two being conflated.
supportsStreaming moves to a shared capability shape. Currently only ChatCapabilities declares it; EvaluateCapabilities should gain the same field so /evaluate/health can advertise streaming support too.
- Per operation, add
text/event-stream to the existing 200 response with a $ref to the relevant terminal frame. New endpoints get streaming for free.
The requirement
Lambda Feedback requires SSE streaming, but
/evaluateand/chatonly document200 → application/json.ChatCapabilities.supportsStreamingalready exists on/chat/health, andLLMConfiguration.streamexists on the request side, so streaming is acknowledged as in scope, but the spec doesn't define what a streaming response actually looks like. Without that, every implementation invents its own frame format and clients have nothing to code against.(A separate PR already extends
callbackUrl/202 Acceptedto/chat, so this issue treats async delivery and SSE as parallel, complementary transport options rather than something to resolve here.)Why this needs to be in the spec
/chatand/evaluateneed to stream identically, and future endpoints shouldn't have to reinvent it.supportsStreamingandLLMConfiguration.streamare meaningless without a documented contract behind them.Feedback[]for evaluate,ChatResponsefor chat) and any progress-stage vocabulary; everything else (the opt-in mechanism, keep-alives, error delivery) can be defined once.Requirements
Accept: text/event-stream. Any operation may support it. When set, the response is a sequence ofSseProgressStepframes, optional keep-alive comments, and exactly one terminal frame.X-Request-IdandX-Api-Versionare set once, as SSE headers, at stream open, since there's no per-frame HTTP header mechanism.SseProgressSteptype (stage, message, timestamp), reused by both endpoints. Stage values are informative, not a strict enum, so implementations can add stages without a spec change; each service's/healthresponse advertises its own supported stage values, the same waysupportedArtefactProfilesadvertises formats.200body (Feedback[]for/evaluate,ChatResponsefor/chat), composed viaallOfagainst each operation's existing200schema rather than duplicated. On failure, the payload is the existingErrorResponseschema, so SSE errors carry the sametitle/code/detailsshape clients already handle for400/500/etc.200for the life of the stream. Failures are always delivered as a terminalErrorResponseframe, not an HTTP error status.configuration.llm.streamis distinct fromAccept: text/event-stream. The former, if kept, governs token-level streaming from the LLM provider; the latter governs step-level progress streaming at the API layer. The spec should say so explicitly to avoid the two being conflated.supportsStreamingmoves to a shared capability shape. Currently onlyChatCapabilitiesdeclares it;EvaluateCapabilitiesshould gain the same field so/evaluate/healthcan advertise streaming support too.text/event-streamto the existing200response with a$refto the relevant terminal frame. New endpoints get streaming for free.