Source of truth: CHANGELOG.md on GitHub
Changelog
All notable changes to OwlCoda public releases are documented here.
[0.15.30] — 2026-07-18
Native runtime reliability release for explainable approvals, resumable session
truth, bounded context hygiene, and clearer terminal evidence.
Added
- Added structured approval risk tiers, primary reasons, persistent-approval
- Added App Server recovery and runtime-rail contracts that preserve thread,
- Added persisted usage totals and model-bound context denominators, including
demotion, and durable decision audit records.
turn, session, and interrupted-work identity across reconnects.
explicit approximate status when the model limit is inferred.
Changed
- Made interactive continuation opt-in by default and prevented synthetic
- Made rolling context hygiene use 90/75 hysteresis, protect six recent tool
- Kept collapsed tool summaries bounded and deduplicated while retaining the
runtime or system turns from creating false unanswered work on resume.
results, aggregate notices, and provide concrete Read recovery hints.
real command arguments needed for audit.
Fixed
- Preserved Markdown bold state across line boundaries, accurate new-file diff
- Returned a small truthful stub for repeated identical reads after a real
- Kept resumed narration, exchange counts, transcript filtering, and runtime
hunk counts, and the original case of displayed CWD and branch values.
filesystem read, without destabilizing read counters.
status aligned with persisted session evidence.
[0.15.29] — 2026-07-11
Harness reliability release for unattended execution, evidence-backed
completion, bounded provider lifecycles, and resumable task receipts.
Added
- Added atomic RunWorkspace artifact registration and completion gates that
- Added contract-aware headless approval for declared safe work while keeping
- Added first-token, idle, and total provider timeouts with structured fallback
- Added interruption receipts, run-scoped write sets, scratch/final path
- Added content-fingerprint-aware verification retries and exact small-sample
require readable artifacts, completed checkpoints, and passing verification.
undeclared paths and unsafe operations fail-closed.
classification, plus cursor-based incremental PTY transcripts.
persistence, and explainable TaskContract permission diffs.
McNemar policy with fail-closed asymptotic guards.
Fixed
- Prevented artifact registration, verification, loop, or repair failures from
- Filtered terminal redraw control sequences from machine-readable transcript
- Returned restart guidance when filesystem roots require startup-time
being reported as completed tasks.
output while preserving raw evidence for review.
configuration instead of repeating ineffective approval attempts.
[0.15.28] — 2026-07-10
Execution-economics patch for bounded structured-output tasks and state-aware
verification retries.
Added
- Added persistent task-level limits for provider calls, input/output tokens,
- Added explicit primary/rerun idempotency reservations and independent
- Added
doctor --jsonrelease identity and deterministic schema bundle hash.
elapsed time, and caller-priced cost, with typed stop checkpoints.
provider, parse, repair, salvage, and rerun counters.
Fixed
- Aborted real upstream requests when the remaining task elapsed budget
- Allowed a failed verification command to run again after tracked worktree
expires, and serialized shared budget/idempotency reservations across
runtime processes.
state changes while still intercepting unchanged no-progress retries.
[0.15.27] — 2026-07-10
Structured-output correctness patch for contract identity, provider-native
schema transport, and typed output-budget failures.
Added
- Added provider/model contract matrices with inspectable provenance for the
- Added complete nested JSON Schema transport for explicitly supported
- Added typed
output_budget_exhaustedevidence across responses, attempts,
effective locale, token budget, and reasoning controls.
OpenAI-compatible and Anthropic structured-output routes.
persisted artifacts, runtime events, and OpenAPI.
Fixed
- Made hybrid built-in/custom structured-output contracts fail closed instead
- Prevented built-in and custom preset/schema identity forgery, including
- Preserved prompt-and-parse fallback for providers whose native JSON Schema
of silently combining mismatched identities, schemas, and defaults.
mutation of shared canonical contract values.
capability is unknown, while keeping successful repair and salvage results
usable when a provider reports max_tokens.
[0.15.26] — 2026-07-10
Workspace safety and state consistency release for destructive-write recovery,
managed worktrees, task verification truth, browser evidence, and long-running
conversation hygiene.
Added
- Added destructive overwrite protection with raw-byte recovery snapshots for
- Added managed worktree lifecycle ledgers, safe resume, fail-closed cleanup,
- Added static preflight for literal cross-workspace
ln -sdependency links - Added run-scoped touched-file receipts that distinguish created, modified,
Write, Edit, and Bash operations.
and separate authorization before deleting commits created after entry.
across common shell stages and wrappers.
unknown, and pre-existing dirty paths.
Changed
- Made TaskVerify the atomic source of verification evidence and task
- Added auditable
pythontopython3fallback and module-type safeguards for - Kept BrowserJob evidence out of project roots, inherited active RunWorkspace
- Executed final task bookkeeping before preserving final reports, continued
completion truth, including failure downgrade and later recovery.
generated scripts.
references, and preserved usable partial evidence on timeout.
non-destructive repair loops autonomously, and reduced repeated long output.
Fixed
- Fixed
git diff --no-index --checksemantic exit handling without masking
actual whitespace diagnostics or compound-command failures.
[0.15.25] — 2026-07-07
Harness reliability patch for local tool availability, task verification, MCP
argument safety, probe evidence, and runtime replay diagnostics.
Added
- Added Codex bundled
rgdiscovery for Bash execution environments, including - Added TaskVerify path diagnostics for relative-path checks, including task
- Added runtime transcript replay metadata derived from persisted runtime
/Applications/Codex.app/Contents/Resources/rg, so local harness checks can
use ripgrep even when the shell path is sparse.
cwd, resolved path, existence, and stat errors when verification fails.
events, exposing timeline, associations, reconnect strategy, and incomplete
replay diagnostics.
Changed
- Reused the read-only verification command profile in TaskCreate, allowing
- Broadened TaskVerify safe read-only validation for common local commands such
- Updated the TUI welcome marker so
CWD/BRANCHlabels stay uppercase
local validation commands such as npx --no-install tsc --version without
falling back to ad-hoc Bash usage.
as node -e, npx tsc --noEmit, and npx vitest run, while still rejecting
mutating flags such as --fix and dangerous commands.
without uppercasing the actual filesystem path; existing cwd paths are
resolved through realpath when possible.
Fixed
- Rejected placeholder MCP arguments such as
server_name,uri,..., and - Prevented ProbePlan
mustContainchecks from being satisfied by ordinary - Preserved partial BrowserJob artifacts on timeout, keeping recovery evidence
- Cleaned temporary desktop preview compile directories after smoke runs.
… before ReadMcpResource dispatches to an MCP server.
echoed stdout; content probes now need file or artifact-backed evidence.
available instead of losing the incomplete run state.
[0.15.24] — 2026-07-07
Runtime rate-limit and task-boundary reliability release.
Added
- Added provider rate-limit handling that stops automatic short-interval
- Added agent provider-rate-limit classification with
- Added
EnterWorktreepreflight checks for high-risk untracked dependency and - Added a structured-output
dataalias that points to the same payload as
retries on direct upstream 429 responses while keeping explicit /retry
available for user-confirmed retries.
providerRateLimited=true, so throttled sub-agent failures stay isolated and
do not become parent-task terminal failures.
source files, with an explicit allow_untracked=true override.
artifact, making HTTP consumers that expect data easier to integrate.
Fixed
- Stopped REPL auto-continue from resending immediately after provider
- Avoided treating CWD banners, terminal transcript snippets, tool-output file
- Resolved short
TaskUpdateIDs such ast-1to canonicaltask-1, and - Classified local
command not foundfailures such as missingrgas - Preserved Bash background-timeout metadata without mislabeling detached work
rate-limit failures; the prompt now guides users toward /model, cooldown,
or explicit /retry.
listings, edit-history lines, or model labels as task-approved write paths.
returned TaskList / TaskCreate recovery guidance for genuinely missing
tasks.
tool:command_not_found, rather than confusing them with remote semantic
failures or missing project data.
as killed.
[0.15.23] — 2026-07-04
Runtime defect follow-up candidate for real workload regressions found in the
sieracMes-AI test environment.
Added
- Added
TaskUpdate({ completePrevious: true })for atomic active-step handoff: - Added raw tool-output artifact preservation for oversized retained tool
- Added per-turn input token budget warnings for oversized provider requests;
- Added gateway auth failure source attribution, grouping 401/403 records by
- Added
/v1/messagesaudit attribution for the requested model, so
when moving a new step to in_progress, OwlCoda can complete the previous
active step first if that previous step is legally completable.
results, with artifactRef, local artifact path, and sha256 in the
truncated transcript marker.
middleware.perTurnInputTokenBudget can tune the threshold or disable it.
API-key fingerprint, user-agent, remote address, and optional client id.
/v1/audit no longer records the endpoint path as the model name.
Changed
- Standardized recoverable public WebFetch
403blocks as - Aligned BrowserJob tool schema with runtime behavior: default browser
- Updated model-routing comments/tests to match the current gateway contract:
- Suppressed duplicate TUI tool completion render events when the same runtime
- Preserved substantial final-answer text when the model appends redundant
remote:blocked_source, keeping them non-terminal and preserving guidance to
use BrowserJob, documented APIs, or blocked-source evidence instead of
treating anti-bot pages as credential failures.
artifacts live under ~/.owlcoda/browser-jobs unless a runRef or explicit
artifactDir is provided.
explicit unknown model requests fail fast instead of silently serving the
configured default.
tool identity is completed more than once.
TodoWrite bookkeeping; OwlCoda now drops the redundant bookkeeping tool
request instead of turning an already-rendered answer into a tool_loop.
Notes
- This candidate still does not add a full history-level rehydration system for
saved tool-output artifacts; it preserves raw evidence and exposes budget
pressure so long sessions no longer have to rely only on transcript text.
[0.15.22] — 2026-07-04
Runtime reliability and observability repair release.
Added
- Added gateway audit data to
/v1/perf, including auth failure counts, status - Added dashboard and
/dashboardslash command warnings for gateway and - Added
ToolDisplayLifecycleso TUI tool start/end rendering pairs by runtime - Added structured
active_step_conflictrepair hints for recoverable
counts, gateway success rate, usable output rate, and zero/thin/slow output
counters.
usable-output health problems.
tool IDs, with same-name FIFO fallback when IDs are unavailable.
TaskUpdate step-state conflicts.
Changed
POST /v1/messagesnow fails fast with404 not_found_errorwhen a request- BrowserJob default artifacts now live under
~/.owlcoda/browser-jobsinstead owlcoda stop --forcenow prints the live REPL client and session details it- Plan mode tools now keep the shared
operatingModeState.modeand legacy
names an explicit unknown model, instead of silently falling back to another
configured model.
of the project root .owlcoda-browser-jobs directory.
is about to detach.
PlanModeState.inPlanMode synchronized under OWLCODA_MODES.
Fixed
- Redacted Bash stdout, stderr, progress lines, provider thinking fields,
- Stored long Bash output as a redacted artifact and returned an
artifactRef - Truncated oversized Grep lines and total Grep output with metadata, and capped
- Kept recoverable
web-fetch:http-403failures from escalating into terminal - Rejected
owlcoda serve --port ... &launched from REPL Bash with guidance to - Avoided misclassifying macOS
/home/...paths mapped through the data volume
bearer/API tokens, URL tokens, long hex tokens, and common cloud/provider
secret shapes before tool output reaches the transcript.
instead of flooding context.
full-file Read output for large files with guidance to continue by range.
semantic failures, and avoided treating successful JSON quota wording as a
terminal failure.
use lifecycle commands instead.
as sensitive /System paths.
Notes
- This release is a CLI/runtime reliability fix. It does not include
RunKit/Desktop, Mem, demo-lab, OwlFootball business logic, or private
execution prompts in the npm package.
[0.15.21] — 2026-07-04
Auditable instruction chain loading release.
Added
- Added layered built-in, user, and project instruction loading with an
- Added
owlcoda instructions inspect --jsonand human-readable instruction - Added
sources,skipped, andlimitsaudit surfaces so operators can see - Added package coverage for root
AGENTS.mdand
auditable source chain.
inspection output for release and runtime audits.
which instruction files were read, skipped, capped, or failed.
docs/INSTRUCTION_CHAIN.md.
Fixed
- Empty
~/.owlcoda/AGENTS.mdnow explicitly blocks Codex fallback instead of - Broken instruction symlinks are reported as
read-errorrather than being AGENTS.override.mdand path-scoped.claude/rulesfiles now record skip
silently reviving built-in defaults.
invisible.
reasons in the audit chain.
Notes
- This release does not include RunKit, Mem, OwlFootball business logic, or
private execution prompts in the npm package.
[0.15.20] — 2026-07-04
Provider streaming delta and agent working guidelines release.
Added
- Added provider-level streaming delta support for
POST /v1/structured-output - Added activity-aware structured-output timeout handling so usable text/content
- Added streaming attempt metadata such as provider SSE mode and delta source,
- Added root
AGENTS.mdwith OwlCoda repository-level working guidelines for
when the resolved model declares streaming support and the request provides a
positive idleTimeoutMs.
deltas refresh the idle timer while heartbeat-only or thinking-only streams do
not count as usable output.
while keeping non-streaming structured-output calls on the existing JSON
compatibility path.
dirty checkout handling, release truth, lane boundaries, verification
discipline, runtime truth, and safety expectations.
Changed
- Extended model capability metadata with declared streaming support so
- Updated CLI harness governance receipts and runtime evidence surfaces for
structured-output routing can decide whether to use provider SSE or remain on
non-streaming transport.
structured-output, workflow, and browser-job recovery paths.
Fixed
- Preserved partial structured-output text when a provider stream interrupts
- Kept thinking-only provider output as non-usable for completion, while still
after usable output has started.
preserving thinking text for diagnostics and fallback artifacts.
Notes
- This release does not include RunKit, Mem, OwlFootball business logic, or
private execution prompts in the npm package.
[0.15.19] — 2026-07-02
Workflow consumer harness read surface.
Added
- Added
WorkflowConsumerManifest v1, a read-only manifest surface for - Added
owlcoda workflow list --jsonand - Added App Server read methods
workflowRun/listandworkflowRun/read, plus - Added workflow outcome facts to scorecard and trajectory surfaces so consumer
workflow runs that lets consumers inspect runtime truth without parsing
natural-language transcripts.
owlcoda workflow inspect --run-id <id> --json for scriptable workflow run
discovery and inspection.
typed client helpers for desktop or external consumers.
layers can reason over execution outcomes as structured runtime facts.
Fixed
- Tightened final-report gating around workflow receipts, required-step
failures, missing artifacts, skipped steps without reasons, and structured
output failed-fallback artifacts.
Notes
- This is a generic runtime harness capability. It is not OwlFootball-specific
and does not include RunKit, Mem, or OwlFootball business logic.
[0.15.18] — 2026-06-27
Release-blocking reliability fixes.
Fixed
TaskVerifynow separates retryable check failures from unsatisfiable orTaskCreateandTaskUpdatenow reject verification policies that wouldTodoWriteandTaskUpdatepreserveblockedandskippedtask states,WebFetch403 responses are recoverable evidence failures instead ofBrowserJobpreserves partial artifacts for selector misses and capture- Bash and long-task timeout reporting now surfaces incomplete snapshots rather
ReadMcpResourceroutesfile://and absolute filesystem paths through the- Structured-output capability handling no longer treats fallback
verdict-blocked failures, and returns structured repair checkpoints with
next-action guidance instead of pushing models into blind workarounds.
require unsafe verification commands before the task is accepted.
require a reason for skipped work, and no longer count skipped or blocked
steps as completed.
terminal research dead ends.
failures, so follow-up recovery can inspect the evidence that did land.
than letting watchdog timeouts be summarized as completed work.
file reader, reducing MCP-tool misuse dead ends.
maxOutputTokens as a hard cap over an explicit caller maxTokens; only
declared or manual model limits cap the request.
Notes
- This release closes the currently exposed release-blocking reliability defects
around verification, recovery evidence, and structured-output provider
controls. It does not claim all deeper platform issues are permanently solved.
[0.15.17] — 2026-06-27
Resumable workflow runner release.
Added
- Added
WorkflowRun, a native resumable workflow runner for multi-step plans - Registered
WorkflowRunacross the native tool registry, CLI command
with persisted workflow state, step outcomes, and recovery metadata.
surfaces, completions, tool risk classification, and public tool docs in the
same release slice.
Changed
- Multi-step workflow execution can now move through the runtime/tool layer
instead of relying on transcript-only task memory.
Notes
- This release keeps the
0.15.16runtime harness consolidation boundary and
adds the reviewed WorkflowRun P0 slice. It does not claim complete resolution
of long-running degradation.
[0.15.16] — 2026-06-27
Runtime harness consolidation release for the Phase D entry point.
Added
- Added the
owlcoda/desktoppackage export for desktop-shell consumers. - Added desktop product-shell view models, live event adapters, smoke probes,
- Added provider evaluation and scorecard surfaces for RL-ready run accounting:
- Extended the structured-output harness with artifact persistence, app-server
runtime facts drilldown, and capability gating for App Server driven shells.
default provider selection, headless audit/runner, report generation,
persistent eval records, and scorecard adapters.
access, role-level rerun support, provider matrix checks, and desktop-facing
artifact APIs.
Changed
- Release validation now treats the runtime harness as the product boundary:
execution, artifacts, provider capability, scorecard, and desktop surfaces are
verified together before public packaging.
Fixed
- Hardened long-task recovery, runtime event accounting, model capability
routing, multimodal image message handling, job supervision, markdown
normalization, and review-center partial apply paths covered by the expanded
release gate.
[0.15.7] — 2026-06-14
Packaging-hygiene release. No runtime change from 0.15.6.
Fixed
- The published tarball is now built from a clean
dist/. Aprebuildstep
removes dist/ before every build, so compiled output whose source has been
deleted can no longer linger. This removes six stale dead-code files from the
removed stub tools (repl, schedule-cron, send-message) that had been
shipping — unregistered and unreachable — since 0.15.5. 0.15.6 remains
published but carries those phantom files; 0.15.7 supersedes it with a clean
tree.
[0.15.6] — 2026-06-14
Command-surface refinement release: fewer, clearer slash commands and keyboard
mode switching.
Added
- **Shift+Tab cycles the operating mode** (
normal → auto → plan;yolostays
explicit), with the mode rail updating live as you cycle.
Changed
- Consolidated the slash-command surface.
/confignow absorbs everything /resetnow combines both former reset commands, and observability output
/settings showed (approval mode, theme, persistent always-allow), and five
pure-duplicate commands were removed: /settings, /color, /tokens,
/reset-circuits, /reset-budgets. The old names still work — they print a
friendly "use X instead" redirect rather than erroring.
folds into /status. /mode with no arguments explains every mode.
Fixed
- Slash-command fuzzy search matches the command **name** first, so typing a few
letters of a command finds it instead of being buried under description
matches.
Notes
- Ships FIFA Phase 2 for the
worldcup-predictordemo to the public source
mirror: deterministic post-match data backfill that feeds an honest tactical
prior into the pre-match brief (combines with, never overrides, pre-match
evidence). The demo lives under demo/ and is excluded from the npm package.
[0.15.5] — 2026-06-14
Safety release: closes the gates through which a code-executing command could run
without the approval the active operating mode promised.
Security
- **HIGH — unattended headless
--mode yoloran dangerous bash without the - **Task sub-agents no longer bypass the approval gate by delegation.** A parent
- **Wrong-case tool names no longer slip past the risk and mode gates.** A model
deny-gate.** In headless runs, mode auto-approve was overriding the headless
safety deny-gate, so a destructive command could execute with no human present.
The deny-gate now survives mode auto-approve: dangerous bash is blocked even
under --mode yolo, while read-only and workspace-test commands still pass.
could hand a dangerous command to a spawned sub-agent, which ran with no
approval callback at all. Dangerous bash in a sub-agent is now gated like the
parent's, closing the delegation bypass.
emitting Bash (instead of canonical bash) sidestepped the destructive-
command gate. Tool names are now canonicalized in the risk, mode, headless,
write-scope, intent, and TUI gates, and in the persistent "Always allow" store
(a wrong-case grant previously never stuck).
Fixed
- Startup
--mode yolono longer desyncs the auto-approve mirror, and/mode /modecopy advertisesyolo, and the/planhint is accurate and gated on
and /plan clear that mirror so you can switch back out of yolo.
whether modes are enabled.
Notes
- Ships the
worldcup-predictordaily auto-review (self-grading loop) v1 demo to
the public source mirror. The demo lives under demo/ and is excluded from the
npm package; the runtime change in this npm release is the safety hardening
above.
[0.15.4] — 2026-06-13
Dogfood-driven harness fixes, operating-mode consolidation, and a sub-agent model override.
Added
- Sub-agents can run on a different model than their parent: the
Agenttool - Admin "test connection" now shows the endpoint's real reported model version.
takes an optional model, resolved as input.model > OWLCODA_SUBAGENT_MODEL
> parent — so an orchestration sub-agent need not share the parent's backend.
Changed
- Unified the operating-mode surface:
/mode,/yolo,/approve, and/plan - Removed three non-functional stub tools (
repl,schedule-cron,
all write one shared mode state that the permission gate reads.
send-message).
Fixed
- Tool dispatch is case-insensitive, so a model emitting
Bashinstead of the - A bare
cdis classified as a safe read-only command instead of being gated. - Tightened bash risk classification so commands that execute code are not
ReadMcpResourcerecovers from common mistakes: it coerces parameter aliases,- The loop guard now catches cross-turn accumulation of the same failure class
- "Step not found" errors from the task tools now list the available step ids.
canonical bash no longer fails with unknown tool.
mislabeled read-only: python -m pytest … -v (pytest's *verbose* flag, not a
version check) and env VAR=val <cmd> now classify by what they actually run.
lists the keys it received, and redirects a file:// URI to Read.
(keyed by tool + failure category), and TaskVerify flags unsatisfiable
checks so the guard stops futile re-verification.
[0.15.3] — 2026-06-12
Interrupt-recovery and rendering hotfix.
Fixed
- Interrupting a tool loop and submitting a new message no longer poisons the
- IME pre-edit no longer renders one row above the composer: the declared
- Ordered-list numbering resets at headings, code fences, and tables, and
session with deterministic 400s: merged consecutive user turns now keep
tool_result blocks first, and an outbound wire guard re-orders any
non-conforming body as a last line of defense.
cursor origin is clamped to the terminal's physical bottom row.
honors an explicit start number — long CJK documents no longer continue
a previous list's counter.
Added
- Render-incident capture: on a render-path throw, the raw-chunk ring buffer
is persisted to a dump for diagnosis.
[0.15.2] — 2026-06-11
Transcript chrome, compaction-resilience, and protocol-hygiene release.
Added
- Transcript chrome S1–S3: collapsed tool results with a unified ok/err shape
and an /expand toggle; narration ● gutter with merged action+result
groups and hanging-indent wrapping; one-line notices, a merged turn footer,
and shared key-value slash panels (including /cost).
Fixed
- Emergency heap-pressure compaction no longer erases task context: task
- Orphaned
tool_use/tool_resultpairs are stripped at the send chokepoint, - Long CJK ordered-list items are no longer split mid-item and renumbered by
- Headless runs fail loudly when a model emits tool-call markers that never
- Unknown models now default to a 200k context window instead of 32768, and
anchors stay pinned, an ineffective-cut breaker stops repeated zero-value
cuts, and a heap-significance gate skips conversations too small to matter,
with pressure diagnostics for each decision.
preventing deterministic 400 loops after interruptions; the daemon now dumps
4xx request shapes for diagnosis.
the fallback sentence splitter.
executed, instead of reporting silent success.
mimo-v2.5 models are recognized at 1M.
[0.15.1] — 2026-06-11
First npm package release on the GPL source line.
Changed
- Moved the public npm install line from the historical
0.14.xstream to - Added the post-source-open runtime fixes already shipped through
0.14.64 - Synced the bilingual README and Admin model screenshot into the npm package
- Updated package metadata, lockfile metadata, Admin display version, and
0.15.x.
to the public source line, including mode visibility, submission recovery,
terminal width hardening, headless exports, third-party skill hardening, and
streaming usage accounting.
surface.
corresponding-source wording for 0.15.1.
Notes
- Paired with public source tag
v0.15.1.
[0.15.0] — 2026-06-04
License boundary and public source availability.
Changed
- Relicensed the OwlCoda core package from
Apache-2.0to - Added
SOURCE.mdto make the corresponding-source requirement explicit for - Updated package metadata, lockfile metadata, OpenAPI license metadata,
GPL-3.0-or-later starting with the 0.15.0 boundary.
npm packages that ship compiled dist/.
README distribution posture, product truth, NOTICE, CONTRIBUTING, and
SECURITY docs for the GPL source line.
Notes
- Commercial, OEM, or embedded distribution is handled through a separate
- Historical published versions remain under the license terms that accompanied
maintainer license path.
those versions when they were published.
Older Releases
Historical package versions remain under the license terms that accompanied those versions when they were published.