release aidev SIG 4/5

Oh My Pi 17.3.5: xAI onboarded to Responses API, GLM-5.3 support, Grok-4.5 default, broad provider reliability overhaul

A substantial cross-package release. The paid xAI provider (XAI_API_KEY) moves from Chat Completions to the OpenAI Responses API, aligning with SuperGrok. Default model moves to grok-4.5 on both xai and xai-oauth. GLM-5.3 is now available with full reasoning-effort ladder and 1M context on z.AI. Dozens of fixes improve retry handling across Anthropic, DeepSeek, Kimi Code, Alibaba, and xAI providers, plus fixes to usage reporting, session recovery, replay of thinking content, and more.

PUBLISHED2026-08-16
OBSERVED2026-08-21
AGE9d
SOURCES1

@oh-my-pi/pi-agent-core

  • Added automatic retry for transient provider failures during one-shot completions (compaction, handoff, branch summarization, manual /compact)

@oh-my-pi/pi-ai

  • retryTransientCompletion support for non-agent LLM calls — retries Anthropic overload/rate-limit errors, HTTP 429/500/502/503/529
  • Fixed xAI availability detection: paid-key-only setups now default to xai/grok-4.5 instead of the free SuperGrok catalog
  • Fixed xAI Requests sending unsupported parameters (reasoning summary, presence/frequency penalties)
  • Fixed Umans usage reporting: corrected raw-request-count → weighted-usage, added soft-cap warning and hard-exhaustion limit
  • Fixed omp usage invalidate not fully clearing stale data
  • Fixed Cursor HTTP/2 connection errors being treated as session-ending instead of transient
  • Fixed OpenAI-compatible streams (DeepSeek) cut off mid-generation being silently accepted
  • Fixed DeepSeek resource-exhaustion interruptions not retrying
  • Fixed tool-call IDs lost during same-model replay
  • Fixed Kimi Code multi-account routing to prefer higher-quota accounts
  • Fixed Anthropic custom signing-proxy conversations losing tool-search results and thinking during replay
  • Fixed runaway response loops failing gracefully instead of repeating indefinitely
  • Fixed xAI rejecting turns due to certain MCP tool schema shapes
  • Fixed Alibaba DashScope/Bailian per-minute rate limits misclassified as quota exhaustion
  • Fixed Anthropic-compatible streams dropping thinking content
  • Updated Alibaba Coding Plan China login flow URL

@oh-my-pi/pi-catalog

  • Added GLM-5.3 on z.AI: unified reasoning-effort ladder, mandatory thinking, 1M context, default-model status
  • Paid xAI (XAI_API_KEY) switched from Chat Completions to OpenAI Responses API — matches SuperGrok for prompt-cache affinity, reasoning-effort handling, encrypted-reasoning replay
  • Default model for xai → grok-4.5; default model for xai-oauth → grok-4.5
  • xAI models now request and replay encrypted reasoning content across multi-turn Responses API calls
  • Fixed Codex Daybreak Blue/Red showing zero token prices (incorrectly labeled free)
  • Fixed Baseten Kimi-K3 catalog metadata for thinking levels
  • Fixed opencode-go/deepseek-v4-flash Responses sending forced tool_choice selectors rejected during thinking

@oh-my-pi/pi-coding-agent

  • Added Extensions tab group to settings schema
  • Routed paid xAI models through Responses API (aligned with SuperGrok), including encrypted reasoning replay
  • Default model for XAI_API_KEY → grok-4.5
  • Stopped sending presence/frequency penalties and stop sequences to xAI reasoning models
  • Fixed hub job/wait lists hiding stale running subagent registrations
  • Fixed external thinking scratchpads conflicting with native xAI Grok 4 reasoning
  • Fixed llama.cpp model discovery missing /v1 prefix for non-Qwen models (404 errors)
  • Fixed prompt caching on open-weight providers (DeepSeek, Qwen, GLM) across directory changes and midnight rollovers
  • Fixed omp --fork omitting artifact directory
  • Fixed long ask option labels being hard-truncated at terminal width
  • Fixed toggling display.showTokenUsage leaving stale rows
  • Fixed mid-run auto-compaction blocking live loop while waiting on extension handlers

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.