release aidev BREAKING SIG 4/5

Oh My Pi 17.4.0: model-scoped Tokenizer replaces global token counting (breaking)

Oh My Pi 17.4.0 removes the global token-counting functions in @oh-my-pi/pi-agent-core and replaces them with immutable per-model Tokenizer instances, which several context-management functions now require explicitly. On the agent side it adds /cleanse, omp ps, speculative compaction, and an extendedContext setting controlling whether premium long-context windows are used.

PUBLISHED2026-08-20
OBSERVED2026-08-21
AGE6d
SOURCES1

Breaking (@oh-my-pi/pi-agent-core)

  • countTokens, countTokensConservatively, setTokenizerModel, and estimateTokens removed; use model-scoped agent.tokenizer with countTokens(text, mode?), countMessage, countMessages.
  • findCutPoint, prepareBranchEntries, collectShakeRegions, pruneToolOutputs, pruneSupersededToolResults, and trimRemoteCompactionInputToContextWindow now require an explicit Tokenizer.
  • createCompactionSummaryMessage takes an options object after (summary, tokensBefore, timestamp).

Tokenizer / catalog

  • Exact embedded token counting added for Claude, Qwen 3.5+, DeepSeek V3/V4/R1, Kimi K2/K3, and GLM-5+; Tokenizer constructs from a resolved catalog Model.
  • New Tokenizer.checkTokenBudget(text, budget) with fast byte-bound pre-check.
  • Provider-anchored transcript estimation: findTranscriptUsageAnchor, isTranscriptUsageAnchor, estimateTranscriptTokens.
  • Models gain an optional tokenizer family field; overridable per model and via modelOverrides for proxy models.
  • Long-context cost tiers (cost.longContext) added to subscription Codex GPT-5.6 models (Sol, Terra, Luna), matching first-party API pricing above 272K input tokens.

Coding agent

  • /cleanse and omp cleanse run the checker/repair loop in-session with a live status board.
  • omp ps interactive monitor for daemon-supervised background processes.
  • extendedContext setting and /extended-context toggle decide whether premium 272K/1M windows are used or the session compacts early to stay on standard pricing.
  • Speculative compaction via compaction.asyncEnabled; compaction.methodOrder replaces compaction.strategy / compaction.remoteEnabled.
  • /handoff now compacts in place instead of forking a new session.
  • Composer layouts (composer.shape), statusLine.contextLine gauge, backgroundable Python eval cells, revamped todo HUD, compaction divider naming the method and before→after sizes.

Fixes

  • Tool-argument repair no longer applies lossy transforms when validating anyOf/oneOf union schemas.
  • Reasoning-effort fallback fixed for local OpenAI-compatible servers returning 400 on chat_template_kwargs.reasoning_effort; qwenTemplateReasoningEffort added for strict local servers.
  • DeepSeek-family models on hosts like Fireworks no longer lose reasoning when tools are offered (redundant tool_choice: "auto" omitted).
  • Fixed tool-call turn failures for opencode-go/muse-spark-1.2 and related variants.

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.