Luigit
repositories / smith

smith

There are many coding harnesses - but this one is fast

owned by admin

.system/research/SMH-RESEARCH-AIAGD001-provider-stream-normalization-design/index.md

Raw
Rendered preview

id: SMH-RESEARCH-AIAGD001 type: research title: "Provider Stream Normalization Design"

Provider Stream Normalization Design

Question

Starting from one OpenAI-compatible adapter, can Anthropic-style and Google adapters be added later without rework, and what must be designed in from the beginning?

Facts

Anthropic Messages streaming and OpenAI chat-completions streaming differ structurally:

dimension OpenAI-compatible Anthropic normalization burden
framing untyped data: JSON chunks, data: [DONE] terminator typed events: message_start, content_block_delta, message_delta, message_stop, ping per-vendor state machine over one output vocabulary
text deltas choices[].delta.content text_delta on a block index trivial: both map to TextDelta
tool arguments arguments JSON string fragmented across chunks, keyed by tool_calls[].index input_json_delta partial JSON per content block, block start carries id and name adapter-local accumulator per index, emits ToolUse complete
thinking not standardized across the compat surface (reasoning_content in some servers) first-class thinking_delta maps to ThinkingDelta, absent or present
stop reason finish_reason on a chunk (stop, length, tool_calls, content_filter) stop_reason in message_delta (end_turn, max_tokens, stop_sequence, tool_use, refusal, pause_turn) exhaustive map into shared StopReason
usage optional final-chunk usage, requires stream_options.include_usage cumulative usage in message_delta maps to Usage at Stop
errors {error:{message,type}} plus HTTP codes typed error events, overloaded_error typed fault categories

Google streamGenerateContent delivers function calls complete in one part, a third argument-delivery shape.

The shared vocabulary in smith/src/stream.rs (TextDelta, ThinkingDelta, complete ToolUse, terminal Stop or Error) already expresses every row of the table; the differences live entirely inside per-vendor state machines.

Adding a provider later becomes hard only if vendor shapes cross the adapter boundary: fragmented argument strings, choice or block indices, [DONE] semantics, or usage-at-end assumptions in the agent loop.

Tradeoffs

  • Normalized-vocabulary-first costs nothing extra for adapter one and makes later adapters additive rather than rework.
  • Designing all three vendors in now would front-load wire details that the conformance fixtures will correct anyway.
  • Typed provider faults cost one enum now versus editing every adapter twice later, mirroring the FrameFault decision in SMH-PLAN-CORE0002.
  • New StopReason variants (Refusal, ContentFilter) are breaking only pre-release; silence-mapping them to EndTurn loses required stop information.

Candidates

  • Add StreamFn to smith before the first adapter so it compiles against neutral request and stream types.
  • Build plan item 6's fixture harness as a provider-agnostic conformance kit; adapters two and three plug into it unchanged.
  • Exhaustive stop-reason maps per adapter: unmapped vendor reason is a compile or explicit runtime fault, never a silent EndTurn.
  • Defer: Anthropic-only mechanics (message_start usage, ping, cache events), mid-stream usage deltas, server-side tools and pause_turn.

Sources

  • Anthropic Messages streaming documentation, event types and error taxonomy.
  • OpenAI chat-completions streaming behavior, chunk shape, finish_reason values, stream_options.include_usage.
  • Google Gemini streamGenerateContent function-call delivery.
  • This repository: smith/src/stream.rs, SMH-PLAN-CORE0002 decisions on typed faults.
---
id: SMH-RESEARCH-AIAGD001
type: research
title: "Provider Stream Normalization Design"
---

# Provider Stream Normalization Design

## Question

Starting from one OpenAI-compatible adapter, can Anthropic-style and Google adapters be added later without rework, and what must be designed in from the beginning?

## Facts

Anthropic Messages streaming and OpenAI chat-completions streaming differ structurally:

| dimension | OpenAI-compatible | Anthropic | normalization burden |
| --- | --- | --- | --- |
| framing | untyped `data:` JSON chunks, `data: [DONE]` terminator | typed events: `message_start`, `content_block_delta`, `message_delta`, `message_stop`, `ping` | per-vendor state machine over one output vocabulary |
| text deltas | `choices[].delta.content` | `text_delta` on a block index | trivial: both map to `TextDelta` |
| tool arguments | `arguments` JSON string fragmented across chunks, keyed by `tool_calls[].index` | `input_json_delta` partial JSON per content block, block start carries `id` and `name` | adapter-local accumulator per index, emits `ToolUse` complete |
| thinking | not standardized across the compat surface (`reasoning_content` in some servers) | first-class `thinking_delta` | maps to `ThinkingDelta`, absent or present |
| stop reason | `finish_reason` on a chunk (`stop`, `length`, `tool_calls`, `content_filter`) | `stop_reason` in `message_delta` (`end_turn`, `max_tokens`, `stop_sequence`, `tool_use`, `refusal`, `pause_turn`) | exhaustive map into shared `StopReason` |
| usage | optional final-chunk usage, requires `stream_options.include_usage` | cumulative usage in `message_delta` | maps to `Usage` at `Stop` |
| errors | `{error:{message,type}}` plus HTTP codes | typed error events, `overloaded_error` | typed fault categories |

Google `streamGenerateContent` delivers function calls complete in one part, a third argument-delivery shape.

The shared vocabulary in `smith/src/stream.rs` (`TextDelta`, `ThinkingDelta`, complete `ToolUse`, terminal `Stop` or `Error`) already expresses every row of the table; the differences live entirely inside per-vendor state machines.

Adding a provider later becomes hard only if vendor shapes cross the adapter boundary: fragmented argument strings, choice or block indices, `[DONE]` semantics, or usage-at-end assumptions in the agent loop.

## Tradeoffs

- Normalized-vocabulary-first costs nothing extra for adapter one and makes later adapters additive rather than rework.
- Designing all three vendors in now would front-load wire details that the conformance fixtures will correct anyway.
- Typed provider faults cost one enum now versus editing every adapter twice later, mirroring the `FrameFault` decision in SMH-PLAN-CORE0002.
- New `StopReason` variants (`Refusal`, `ContentFilter`) are breaking only pre-release; silence-mapping them to `EndTurn` loses required stop information.

## Candidates

- Add `StreamFn` to `smith` before the first adapter so it compiles against neutral request and stream types.
- Build plan item 6's fixture harness as a provider-agnostic conformance kit; adapters two and three plug into it unchanged.
- Exhaustive stop-reason maps per adapter: unmapped vendor reason is a compile or explicit runtime fault, never a silent `EndTurn`.
- Defer: Anthropic-only mechanics (`message_start` usage, `ping`, cache events), mid-stream usage deltas, server-side tools and `pause_turn`.

## Sources

- Anthropic Messages streaming documentation, event types and error taxonomy.
- OpenAI chat-completions streaming behavior, chunk shape, `finish_reason` values, `stream_options.include_usage`.
- Google Gemini `streamGenerateContent` function-call delivery.
- This repository: `smith/src/stream.rs`, `SMH-PLAN-CORE0002` decisions on typed faults.