--- id: SMH-RESEARCH-RSCH0007 type: research title: "Provider Normalization Research" --- # Provider Normalization Research Purpose: check whether Pi generated provider data, catwalk provider configs, and live provider model APIs are similar enough to normalize for Smith provider suggestions. ## Sources Checked - Pi generated registry clone: `https://github.com/earendil-works/pi-mono.git` at `3eb0027`, file `packages/ai/src/models.generated.ts`. - Catwalk clone: `https://github.com/charmbracelet/catwalk.git` at `2fe397e`, files `internal/providers/configs/zai.json` and `internal/providers/configs/minimax.json`. - Live Z.AI API probes using `ZAI_API_KEY` from environment. Key was not printed. - Live MiniMax API probes using `MINIMAX_API_KEY` from environment. Key was not printed. ## Shape Comparison Observed common normalized fields: - provider id, - provider name, - API/base URL, - API family (`openai-completions` / OpenAI-compatible, `anthropic-messages` / Anthropic-compatible), - model id, - model display name, - context window, - max output/default max tokens, - input modality / attachment support, - reasoning support, - per-token or per-million-token costs, - cache read/write costs when known. Observed shape differences: | Concept | Pi generated registry | catwalk config | |---|---|---| | API family | `api` values like `openai-completions`, `anthropic-messages` | `type` values like `openai-compat`, `anthropic` | | Base URL | per model `baseUrl` | provider `api_endpoint` | | Context | `contextWindow` | `context_window` | | Max output | `maxTokens` | `default_max_tokens` | | Reasoning | `reasoning` | `can_reason` | | Modalities | `input: ["text", "image"]` | `supports_attachments` | | Cost | `cost.input/output/cacheRead/cacheWrite` | `cost_per_1m_*` | | Compat quirks | `compat` object | mostly absent | | Defaults | not primary in model entries | `default_large_model_id`, `default_small_model_id` | ## Z.AI Observations Pi direct `zai` entries observed: - `glm-4.5-air` - `glm-4.7` - `glm-5-turbo` - `glm-5.1` - `glm-5v-turbo` Catwalk `zai` entries observed: - `glm-5.1` - `glm-5-turbo` - `glm-5` - `glm-4.7` - `glm-4.7-flash` - `glm-4.6` - `glm-4.6v` - `glm-4.5` - `glm-4.5-air` - `glm-4.5v` Live API probes: ```text zai-coding-models: 200 → glm-4.5, glm-4.5-air, glm-4.6, glm-4.7, glm-5, glm-5-turbo, glm-5.1 zai-general-models: 200 → glm-4.5, glm-4.5-air, glm-4.6, glm-4.7, glm-5, glm-5-turbo, glm-5.1 zai-coding-chat-smoke: 200 ``` Observed drift: - Pi has `glm-5v-turbo`; live model listing did not show it. - Catwalk has `glm-4.7-flash`; live model listing did not show it. - Catwalk has pricing for Z.AI; Pi direct Z.AI costs are zero in generated data. - Pi has provider-specific `compat.thinkingFormat = "zai"` and `zaiToolStream` flags; catwalk does not. ## MiniMax Observations Pi direct `minimax` entries observed: - `MiniMax-M2.7` - `MiniMax-M2.7-highspeed` Catwalk `minimax` entries observed: - `MiniMax-M2.7` - `MiniMax-M2.7-highspeed` - `MiniMax-M2.5` - `MiniMax-M2.5-highspeed` - `MiniMax-M2.1` - `MiniMax-M2.1-highspeed` - `MiniMax-M2` Live API probes: ```text minimax-openai-models: 200 → MiniMax-M2.7, MiniMax-M2.7-highspeed, MiniMax-M2.5, MiniMax-M2.5-highspeed, MiniMax-M2.1, MiniMax-M2.1-highspeed, MiniMax-M2 minimax-anthropic-models: 200 → MiniMax-M2.7, MiniMax-M2.7-highspeed, MiniMax-M2.5, MiniMax-M2.5-highspeed, MiniMax-M2.1, MiniMax-M2.1-highspeed, MiniMax-M2 minimax-anthropic-smoke: 200 ``` Observed drift: - Catwalk coverage matches live model listing better than Pi direct MiniMax. - Pi has Anthropic-compatible base URL and cache/cost fields for direct MiniMax. - Catwalk context values are rounded (`200000`) while Pi uses `204800`. - Live model listings confirm model IDs but not costs/context/compat behavior. ## Candidate Conclusion Pi generated data and catwalk configs are similar enough to normalize into a Smith suggestion schema for common model metadata. They are not similar enough for silent replacement. Required review points: - model coverage drift, - context/max-token conflicts, - pricing conflicts or missing pricing, - provider-specific compat quirks missing from catwalk, - live model endpoints listing IDs but not enough metadata for correctness. Best fit: Pi as primary, catwalk as gap-fill, live provider APIs as optional ID-availability evidence only. Current checked-in `providers.json` remains the runtime authority.