---
title: Release Notes
description: SIE release history.
canonical_url: https://superlinked.com/docs/reference/release-notes
last_updated: 2026-08-25
---

{/* AUTOGENERATED by scripts/fetch-release-notes.mjs. Do not edit by hand. */}

Latest version: **v0.7.1** (2026-08-09).

## v0.7.1 (2026-08-09)

### Highlights

- **New capabilities:** serve docling OCR with the English recogniser
- **Reliability and operations:** meter executed GLiNER token window

### Features

* **server:** serve docling OCR with the English recogniser

### Bug Fixes

* **server:** meter executed GLiNER token window

## v0.7.0 (2026-08-08)

### Highlights

- **Breaking change:** aws-control-plane no longer accepts component_contract_sha256 or injects its environment marker.
- **New capabilities:** add fail-closed estate-bootstrap saved-plan inspection; add frozen assembly apply seam; add minimal promotion eligibility guard; admit estates by declared capability and add the sandbox desired state; allow benchmark campaign self-review; assert apply-free and same-scope second convergence runs
- **Reliability and operations:** align catalog smoke timeout budget; apply the checkpoint whitelist to record, harden the git guard; bind the published price table to the active billing authority; harden document field fixture validation; harden generation response handling

### ⚠ BREAKING CHANGES

* **terraform:** aws-control-plane no longer accepts component_contract_sha256 or injects its environment marker.

### Features

* **cloud:** add fail-closed estate-bootstrap saved-plan inspection
* **cloud:** add frozen assembly apply seam
* **cloud:** add minimal promotion eligibility guard
* **cloud:** admit estates by declared capability and add the sandbox desired state
* **cloud:** allow benchmark campaign self-review
* **cloud:** assert apply-free and same-scope second convergence runs
* **cloud:** batch the priced rate-grade cells behind one approval
* **cloud:** bill audio_ms at a 1,150 ms minimum duration
* **cloud:** bind the inspection verdict to the exact plan digest
* **cloud:** declare the terraform backend kind per estate cell
* **cloud:** deepen the read-only estate bootstrap preflight
* **cloud:** default tier points to min(floor x 2.0, ceiling x 0.85)
* **cloud:** freeze infrastructure values in assemblies
* **cloud:** hydrate terraform outputs for remote-backend cells
* **cloud:** implement the #2441 pricing-policy decisions (PR #3024)
* **cloud:** local-only session auth for the L3 verification stack
* **cloud:** make the sandbox cell serving
* **cloud:** permanently allow benchmark campaign self-review
* **cloud:** persist deliverable metadata in s3
* **cloud:** publish the live rate book to a public price artifact
* **cloud:** record and verify estate bootstrap approvals
* **cloud:** record convergence evidence and compare two runs
* **cloud:** warn loudly when a --live preflight run proves nothing
* **control-plane:** serve-image capability manifest and operator-opened tasks
* **model-ops:** add the docling English OCR artifact publish tool

### Bug Fixes

* **ci:** refuse an unusable assets_path before the gate
* **ci:** replace the batch ledger atomically and unwrap the summary prose
* **clients:** address terminal response review
* **cloud:** accept padded whisper transcripts
* **cloud:** accept real names, escape paths, close the commondir fail-open
* **cloud:** address deployment review findings
* **cloud:** admit provisional task measurement pairs
* **cloud:** admit same-apply IAM unknowns by provenance, not by fiat
* **cloud:** align catalog canary model contracts
* **cloud:** align catalog smoke timeout budget
* **cloud:** align document field canary semantics
* **cloud:** align phase 3 exit criteria and make skip tests prove their cause
* **cloud:** align PII acceptance taxonomy
* **cloud:** allow task candidates across routes
* **cloud:** allowlist read-only init by its full argument set
* **cloud:** apply the checkpoint whitelist to record, harden the git guard
* **cloud:** assert the full run built something and order the pair
* **cloud:** bind every resource to the verified provider, not only data sources
* **cloud:** bind read-only terraform access to the resolved cell's root
* **cloud:** bind the published price table to the active billing authority
* **cloud:** bound Modal image build cache
* **cloud:** bound provenance to policy CONTENT, not just mentioned resources
* **cloud:** close checkpoint leak, env-independent repo guard, strict schema
* **cloud:** close fail-open branch chains in the estate preflight
* **cloud:** close fail-open paths in estate plan inspection
* **cloud:** close preflight fail-opens found by the mandated self-review
* **cloud:** close the fail-open paths found in review threads
* **cloud:** close the fail-opens the third self-review found
* **cloud:** close the holes the fourth self-review found in this round's own permissive paths
* **cloud:** close the plan-inspection bypasses from the adversarial review
* **cloud:** close the two backend-classifier evasions from verification
* **cloud:** commit job results before settlement
* **cloud:** confine catalog subsets to sandbox
* **cloud:** constrain checkpoint field values, not just field names
* **cloud:** constrain screenshot bbox axes
* **cloud:** correct the apply-check rationale and scope-check claim
* **cloud:** correct the offload-smoke minted tuple annotation and notice an ignored --account
* **cloud:** correct the provider-key rule against ground truth, never echo after values
* **cloud:** declare the revoke 404 contract and pin its no-oracle body
* **cloud:** decouple availability from measurements
* **cloud:** defer gateway replica probe until serve
* **cloud:** derive the declared commercial margin from the price multiplier
* **cloud:** enforce the D1 boundary on the terraform primitive
* **cloud:** escape the last raw path, refuse an unreadable HEAD
* **cloud:** fail closed on an unreadable desired-state spec
* **cloud:** freeze deployment check results
* **cloud:** fullmatch the rate-book version before it becomes a filename
* **cloud:** give every checkpoint field a format, not a length cap
* **cloud:** guard the path parameters, pin refusals, drop the bare except
* **cloud:** harden document field fixture validation
* **cloud:** harden generation response handling
* **cloud:** harden the estate preflight probes per security review
* **cloud:** harden the point default and the artifact publish tool
* **cloud:** harden the preflight per the CodeRabbit review threads
* **cloud:** hold the D1 remote-backend boundary in live preflight too
* **cloud:** hold the priced generation window's prefill mix to a review
* **cloud:** isolate deployment plan inputs
* **cloud:** keep production catalog topology strict
* **cloud:** keep the pricing-authority resolver import-safe off a checkout
* **cloud:** make shape handling total and stop rejecting the honest plan
* **cloud:** make the unpriced-storage zero loud
* **cloud:** parse s3 backend blocks brace-aware so nested blocks are read
* **cloud:** pin the AWS-edge coupling and refuse malformed input everywhere
* **cloud:** read not-owned terraform roots through a narrow allowlist
* **cloud:** read trust_remote_code from both adapter scopes everywhere
* **cloud:** recover externally cancelled live gateway candidate
* **cloud:** refuse impossible instants and stop the scope string overclaiming
* **cloud:** register migration 0079 and name the book on a merged sweep
* **cloud:** reject a run paired with itself and fix the stale exit bullet
* **cloud:** reject lineage cycles and parented full runs
* **cloud:** reject unusable SSM pagination tokens; pin cross-estate isolation
* **cloud:** repair chat catalog canary
* **cloud:** repair document field canary semantics
* **cloud:** require exact screenshot labels
* **cloud:** require statement constancy per security-relevant field
* **cloud:** restore cross-estate rejection on excluded entries, refuse composition
* **cloud:** restrict the instant pattern to ASCII digits
* **cloud:** retry transient job reads
* **cloud:** scan every leaf for foreign-estate references, not just name keys
* **cloud:** scope key mutations to an account the caller names
* **cloud:** snapshot untracked development roots
* **cloud:** stabilize catalog contract canary
* **cloud:** stabilize screenshot acceptance fixture
* **cloud:** stabilize screenshot contract canary
* **cloud:** stabilize whisper catalog canary
* **cloud:** stop arming gateway proxy auth without a public edge
* **cloud:** stop publishing SIE's markup to customers
* **cloud:** tighten promotion evidence validation
* **cloud:** tolerate EC2 worker visibility lag
* **cloud:** track %\{\} template-directive spans in the backend classifier
* **cloud:** track template interpolations in the backend classifier
* **cloud:** unpin gateway path dependency
* **cloud:** validate mode, harden provider-key parsing, make two tests honest
* **cloud:** validate read entries and close self-review fail-opens
* **cloud:** validate the estate id where it enters, tighten approver
* **cloud:** verify materialized candidate bytes
* **deps:** provide the dependencies the bundle and the eval already declare
* **model-ops:** pin trusted digests and refuse unrecorded provenance
* **release:** unpin sidecar path dependency
* **sdk:** refresh stale job result refs
* **server:** echo queued extract item ids
* **server:** normalize Grounding DINO instructions
* **server:** preserve pinned models in serve config
* **server:** validate pinned model selection
* **tools:** allow SDK 0.7 releases

### Code Refactoring

* **terraform:** drop control-plane contract marker

## v0.6.30 (2026-08-07)

### Highlights

- **New capabilities:** compile the deployment matrix from the estate registry
- **Reliability and operations:** drop the invalid queue key from the staging concurrency blocks
- **Performance:** accelerate Qwen thinking inference

### Features

* **cloud:** compile the deployment matrix from the estate registry

### Bug Fixes

* **ci:** drop the invalid queue key from the staging concurrency blocks

### Performance Improvements

* **models:** accelerate Qwen thinking inference

## v0.6.29 (2026-08-07)

### Highlights

- **New capabilities:** bind evidence and acceptance to estate identity; optimize Gemma thinking profiles
- **Reliability and operations:** preserve release tag during public sync verification; verify the acceptance OIDC identity token

### Features

* **cloud:** bind evidence and acceptance to estate identity
* **server:** optimize Gemma thinking profiles

### Bug Fixes

* **ci:** preserve release tag during public sync verification
* **cloud:** verify the acceptance OIDC identity token

## v0.6.28 (2026-08-07)

### Highlights

- **New capabilities:** derive resource names and paths from estate coordinates; gate public edge and vendor identity on estate capability; split managed gates into tier and estate capability; optimize non-thinking generation profiles
- **Reliability and operations:** classify entity canary failures; gate Stripe preflight before any provider request; keep vendor provider access tier-gated; refresh siglip diagnostic projection; remove stale console gateway URL and reject tampered cells

### Features

* **cloud:** derive resource names and paths from estate coordinates
* **cloud:** gate public edge and vendor identity on estate capability
* **cloud:** split managed gates into tier and estate capability
* **server:** optimize non-thinking generation profiles

### Bug Fixes

* **cloud:** classify entity canary failures
* **cloud:** gate Stripe preflight before any provider request
* **cloud:** keep vendor provider access tier-gated
* **cloud:** refresh siglip diagnostic projection
* **cloud:** remove stale console gateway URL and reject tampered cells
* **cloud:** scope Modal denial floor to serving estates
* **cloud:** update gateway profile fixtures
* **quality:** extend PubTabNet task budget
* **release:** extract refresh artifact by id
* **release:** keep diagnostic suites out of git
* **release:** preserve refresh artifact across retries
* **server:** pin speculative draft launches
* **server:** stage drafts for local models

## v0.6.27 (2026-08-06)

### Highlights

- **New capabilities:** add staging comparison diagnostics; pin benchmark campaigns to a verified main revision; add connector plan controls; expose recovery-bound connector repair; activate atomic connector run; activate explicit connector execute
- **Reliability and operations:** authenticate quality gate attempt markers; authenticate req14 quality gate markers; pin void authorization Python; bind retained recovery to candidate tasks; bind vendor mutations to estate authority

### Features

* **bench:** add staging comparison diagnostics
* **ci:** pin benchmark campaigns to a verified main revision
* **clients:** add connector plan controls
* **clients:** expose recovery-bound connector repair
* **cloud:** activate atomic connector run
* **cloud:** activate explicit connector execute
* **cloud:** activate governed postgres connectors
* **cloud:** activate recovery-bound connector repair
* **cloud:** add atomic connector run admission
* **cloud:** add bounded connector recovery probes
* **cloud:** add bounded postgres source planning
* **cloud:** add component convergence contracts
* **cloud:** add connector execution journal client
* **cloud:** add connector plan guardrails
* **cloud:** add connector plan schema policy
* **cloud:** add crash-safe connector outer runner
* **cloud:** add direct connector inference
* **cloud:** admit executable connector plans
* **cloud:** attest connector workers before payload
* **cloud:** band the size-class tier book on (operation, parameters)
* **cloud:** band the tier book on (operation, parameters)
* **cloud:** bind connector executor dispatch
* **cloud:** bind connector plans to worker identity
* **cloud:** bind direct connector executions
* **cloud:** bound postgres connector execution context
* **cloud:** bridge durable connector execution dispatch
* **cloud:** commit the control-plane OpenAPI contract with a diff gate
* **cloud:** compile estate vendor coordinates
* **cloud:** compile optional estate vendor authority
* **cloud:** compose connector recovery journal
* **cloud:** compose dormant connector execution
* **cloud:** enable native console previews
* **cloud:** enable selective component deployment
* **cloud:** expose connector execution transitions
* **cloud:** expose durable connector plans
* **cloud:** fail closed on unconfigured vendor authority
* **cloud:** fence postgres connector publication
* **cloud:** govern connector plan manifests
* **cloud:** isolate SigLIP rate-grade diagnostics
* **cloud:** journal connector recovery phases
* **cloud:** land governed Postgres connector runtime
* **cloud:** persist connector batch lane authority
* **cloud:** persist connector execution authority
* **cloud:** persist durable connector plans
* **cloud:** preflight connector worker discovery
* **cloud:** prepare and fence connector execution
* **cloud:** prove executable postgres preflight
* **cloud:** provision connector runtime authorities
* **cloud:** provision isolated console previews
* **cloud:** reconcile connector publication recovery
* **cloud:** record the owner's acceptance of the declared exposure
* **cloud:** replace catalog DSL with task implementations
* **cloud:** replay governed postgres manifests
* **cloud:** resolve estate coordinates at the provisioner boundary
* **cloud:** resolve terraform coordinates from the estate registry
* **cloud:** route generation through local ingest
* **cloud:** route managed generation through local ingest
* **cloud:** scaffold additional staging estate
* **cloud:** scaffold sandbox foundation
* **cloud:** settle connector staging before publication
* **cloud:** unify Modal environment authority in the estate registry
* **gateway:** expose immutable execution evidence
* **generation:** add 256k profile candidates
* **generation:** add hardware-specific context profiles
* **generation:** add strict Responses API parity
* **model-ops:** allowlist lightonai and OpenSearch-AI as trusted publishers
* **model-ops:** capture nested multimodal config scalars in onboarding evidence
* **model-ops:** enable multimodal image-text and document-retrieval discovery families
* **model-ops:** enable sparse-retrieval discovery family
* **model-ops:** re-task late-interaction discovery onto SCIDOCS retrieval evidence
* **model-ops:** trust JinaAI discovery publisher
* **models:** onboard lightonai/mLateOn
* **models:** onboard lightonai/mLateOn (scaffolded)
* **models:** seed mLateOn quality floors from governed measurement
* **server:** add bounded local-ingest generation client
* **server:** complete generative model serving
* **sidecar:** stream generation over local ingest
* **terraform:** register shared Scalr modules

### Bug Fixes

* **bench:** validate release evidence ingress early
* **build:** relock every standalone project that depends on sie-server
* **ci:** align standalone validation artifacts
* **ci:** authenticate quality gate attempt markers
* **ci:** authenticate req14 quality gate markers
* **ci:** clarify invalid model gate error
* **ci:** exempt registry library modules from the estate root list
* **ci:** fail closed for unlisted registry module roots
* **ci:** format campaign source tests and state the lowercase SHA rule
* **ci:** isolate connector verifier dependencies
* **ci:** pin void authorization Python
* **ci:** trust quality gate attempt cap source
* **ci:** validate quality gate void reasons
* **clients:** send connector repair idempotency header
* **cloud:** accept launch catalog schema v2
* **cloud:** accept operation-valid usage supersets
* **cloud:** address connector runtime review findings
* **cloud:** address Gateway identity review
* **cloud:** address staging rollback review
* **cloud:** align estate module contracts
* **cloud:** align preview posture lint config
* **cloud:** allow console preview bootstrap
* **cloud:** allow verified gateway final reuse
* **cloud:** batch the token-capped encode lanes and price two more launch cells
* **cloud:** bind connector intent to canonical worker
* **cloud:** bind every priced-window factor to non-caller evidence
* **cloud:** bind generation lanes to catalog hash
* **cloud:** bind retained benchmark owners
* **cloud:** bind retained recovery to candidate tasks
* **cloud:** bind task evidence to reviewed cases
* **cloud:** bind the billed report period to the priced window
* **cloud:** bind vendor mutations to estate authority
* **cloud:** bound Modal history reads separately
* **cloud:** bound Modal inventory reads separately
* **cloud:** close connector activation review gaps
* **cloud:** close connector review findings
* **cloud:** close connector review gaps
* **cloud:** close two vacuous guards in the tier banding logic
* **cloud:** declare the below-cost classes instead of hiding them
* **cloud:** default estate_id inside resolve_estate
* **cloud:** deny unused staging scheduler authority
* **cloud:** derive allocation shares from measured work, not fiat
* **cloud:** derive offered load from measured lane capacity
* **cloud:** disable staging cold-start reporting
* **cloud:** drop the exporter's path workaround, annotate the document type
* **cloud:** enforce console preview posture
* **cloud:** fence selective Gateway dispatch by contract
* **cloud:** finish connector protocol review sweep
* **cloud:** fold the 2026-08-04 lane sweeps into the offered-load plan
* **cloud:** forward selective deployment mode
* **cloud:** gate connector activation on desired state
* **cloud:** handle local generation stream failures
* **cloud:** harden connector activation boundaries
* **cloud:** harden deployment mode forwarding
* **cloud:** harden edge backend ordering, cell parity, bucket names
* **cloud:** harden managed generation streaming
* **cloud:** harden Modal inventory validation
* **cloud:** harden preview OIDC validation
* **cloud:** harden registry identity validation and lock estate identity
* **cloud:** harden selective deployment reuse
* **cloud:** harden staging estate isolation
* **cloud:** harden staging foundation identity
* **cloud:** ignore local source caches
* **cloud:** isolate console preview git config
* **cloud:** isolate sidecar subprocess secrets
* **cloud:** isolate staging estate authorities
* **cloud:** make Vercel readback authoritative
* **cloud:** measure the enlarged fixtures and price 18 more launch cells
* **cloud:** merge main and add band-6 rerank capacities
* **cloud:** normalize component source modes
* **cloud:** omit read-only webhook rules
* **cloud:** omit response-only webhook rules
* **cloud:** order config image local layers
* **cloud:** pin gateway plan evidence
* **cloud:** pin sandbox Scalr modules at their first healthy versions
* **cloud:** preserve connector execution topology
* **cloud:** preserve positive sandbox activation target
* **cloud:** preserve retained worker provenance in release gates
* **cloud:** preserve valid rate book evidence
* **cloud:** price the measured window, not the billed report hour
* **cloud:** re-render siglip diagnostics lock after server source change
* **cloud:** reconcile connector activation with main
* **cloud:** record the measured decode length on the generation cells
* **cloud:** recover native production gateway
* **cloud:** recover retained candidates after newer merges
* **cloud:** refresh combined topology evidence
* **cloud:** refresh siglip diagnostic lock
* **cloud:** refresh task contract test fixtures
* **cloud:** refuse an unpublished class name instead of raising KeyError
* **cloud:** reject empty connector inference text
* **cloud:** reject JSON Terraform escapes
* **cloud:** remove unused type suppression
* **cloud:** report rejected reuse evidence
* **cloud:** require integer estate schema version
* **cloud:** reserve governed page ceilings
* **cloud:** restrict promotion authority to full estates
* **cloud:** retain legacy standalone tag gate
* **cloud:** scope promotion authority to estate cells
* **cloud:** stabilize component binary identity
* **cloud:** stabilize named-entity canary fixture
* **cloud:** state the audio_ms clip-length assumption in the quote
* **cloud:** tighten connector manifest cleanup
* **cloud:** tighten task contract evidence gates
* **cloud:** unblock reproducible Modal convergence
* **cloud:** use explicit connector protocol stubs
* **cloud:** use non-ambiguous visual distractor
* **cloud:** use representative vision smoke fixture
* **cloud:** validate retained migration without activation
* **cloud:** validate reused worker runtime CAS
* **cloud:** validate selected gateway in catalog smoke
* **cloud:** validate settled acceptance usage
* **cloud:** verify Modal service-user access
* **colbert:** honor transformers-5 rope_parameters when serving on the 4.x bundle
* **generation:** align direct backend output limits
* **generation:** defer resolved profile output limits
* **generation:** fit Gemma 256k KV cache on H100
* **generation:** handle family-specific reasoning boundaries
* **generation:** hide split Gemma thinking boundary
* **generation:** pair long-context qwen image bounds
* **generation:** preserve Gemma reasoning boundaries
* **generation:** preserve qwen vision bounds
* **generation:** reject overflowing runtime numerics
* **generation:** reserve Gemma long-context workspace
* **generation:** unblock cold model startup
* **generation:** use compatible Qwen 27B grammar backend
* **generation:** use compatible Qwen 35B grammar backend
* **generation:** validate high-context thinking profiles
* **modal:** isolate overlay FlashInfer cache
* **modal:** keep image yaml loading local
* **modal:** support strict GPU allocation
* **modal:** validate strict GPU selectors
* **model-ops:** an unmeasured expected_dim never confirms a multivector smoke
* **model-ops:** batch discovery evidence revision listing via one pinned tree fetch
* **model-ops:** bind trusted terminal outcomes
* **model-ops:** bound authority listing by model directories, not raw entries
* **model-ops:** calibrate late-interaction and doc-retrieval scanner blocks
* **model-ops:** close scaffold authorization gaps
* **model-ops:** complete scaffolds before Stage 4
* **model-ops:** exempt tooling stops from attempt cap
* **model-ops:** fail closed on duplicate search errors
* **model-ops:** harden scaffold validation boundaries
* **model-ops:** ignore fenced tooling markers
* **model-ops:** insert matrix rows once per models node so alias blocks ride their anchor
* **model-ops:** isolate scaffold candidate tests
* **model-ops:** namespace the config projection_dim claim and cap consumer-side nested ints
* **model-ops:** order attempt comments deterministically
* **model-ops:** preserve adapter auto authorization
* **model-ops:** scope duplicate onboarding guard
* **model-ops:** teach the Stage-4 smoke to probe multivector models
* **models:** complete mLateOn serving recipe
* **quality-eval:** harden runner repair
* **quality-eval:** quarantine unregistered runners
* **quality-eval:** shorten wake grace and map sglang bundles
* **quality-eval:** shorten wake grace and support sglang logical bundles
* **quality-eval:** use prepared scorer environment
* **release:** bound commit history batches
* **release:** deduplicate generated changelog notes
* **release:** finalize source before diagnostics
* **release:** refresh topology-bound diagnostics
* **release:** reject repeated release headings
* **sdk:** retry pre-execution SSE capacity errors
* **server:** align Gemma xgrammar dependency
* **server:** await standalone smoke teardown
* **server:** bound Qwen visual context
* **server:** constrain SGLang compatibility paths
* **server:** enforce Qwen fast image bounds
* **server:** keep Qwen3-VL reranker batches working on transformers 5.3
* **server:** pin docling to its artifact revision and stop OCR gating load
* **server:** preprocess native extract images
* **server:** preserve LightOn SGLang compatibility
* **server:** replace stale root playground with redirect to Swagger docs
* **sidecar:** harden local generation cancellation
* **sidecar:** reserve cancel request bytes
* **terraform:** address internal module review
* **terraform:** bound internal module lifecycle
* **terraform:** exclude worktree caches from remote config uploads
* **terraform:** harden Scalr module releases
* **terraform:** package internal Scalr module archive root
* **terraform:** package private Scalr registry modules
* **terraform:** package Scalr module archive envelopes
* **terraform:** register Scalr modules from plain monorepo roots
* **terraform:** reject null module registry provider
* **terraform:** relocate Scalr registry modules under deploy tree
* **terraform:** validate sandbox module with OpenTofu

## v0.6.26 (2026-08-02)

### Highlights

- **New capabilities:** add H100 FP8 generation profiles; governed first-baseline mode + 16 vanilla floors for Qwen3-Embedding-8B and nomic v2 MoE
- **Reliability and operations:** add protected retained recovery dispatch; bind modal env before retained recovery; bind retained gateway recovery to failed release; bound retained recovery runtime; clarify retained gateway recovery

### Features

* **model-catalog:** add H100 FP8 generation profiles
* **quality-eval:** governed first-baseline mode + 16 vanilla floors for Qwen3-Embedding-8B and nomic v2 MoE

### Bug Fixes

* **bench:** initialize buildx for worker image
* **bench:** quiesce gc during timed campaigns
* **bench:** validate non-root worker runtime
* **cloud:** add protected retained recovery dispatch
* **cloud:** address provision review findings
* **cloud:** align rollback preflight diagnostic
* **cloud:** allow full cold dispatch preflight budget
* **cloud:** audit extracted provision secrets
* **cloud:** bind modal env before retained recovery
* **cloud:** bind retained gateway recovery to failed release
* **cloud:** bind retained gateway to failed release
* **cloud:** bound retained recovery runtime
* **cloud:** clarify retained gateway recovery
* **cloud:** close provision review gaps
* **cloud:** converge config before cold gateway
* **cloud:** fence initial gateway absence recovery
* **cloud:** govern reconciliation cadence in releases
* **cloud:** guard bootstrap and benchmark worker runtime
* **cloud:** guard current-schema bootstrap seed
* **cloud:** harden connector smoke recovery
* **cloud:** harden extracted provision edges
* **cloud:** harden retained recovery diagnostics
* **cloud:** honor generation smoke cold starts
* **cloud:** isolate config app runtime import
* **cloud:** isolate connector smoke environments
* **cloud:** isolate sibling smoke environments
* **cloud:** keep connector diagnostic out of MVP release gate
* **cloud:** keep prod generation warm and scope spans
* **cloud:** make generation smoke single shot
* **cloud:** narrow connector loopback scope
* **cloud:** preserve accepted bundle promotion compatibility
* **cloud:** preserve recovery fence diagnostics
* **cloud:** recover first gateway to absence
* **cloud:** recover retained legacy gateway
* **cloud:** recover retained staging gateway to absence
* **cloud:** recover staging gateway to absence
* **cloud:** reject incomplete cold gateway plans
* **cloud:** retain production deployment evidence
* **cloud:** retain zero-floor generation attestations
* **gateway:** preserve grammar-safe FP8 profiles
* **loadtest:** prepare pinned perf-lab source before AWS
* **model-ops:** bind scaffold AUTO to trusted evidence
* **model-ops:** mirror task-equivalent standard blocks
* preserve cold-start reporter membership history
* **quality-eval:** preserve missing metric diagnostics
* **quality:** run comparison through locked project
* **release:** identify expired Modal attestation
* **release:** keep Modal diagnostic trim ASCII-only
* **release:** reconcile lagging Modal attestations
* **release:** tolerate completed Modal exec targets
* **release:** wait for refreshed PR head
* validate initial reporter lease atomically

## v0.6.25 (2026-07-30)

### Highlights

- **New capabilities:** activate bootstrap-v2 pricing in prod; add --skip for recorded vendor-convergence skips; add cold-start live acceptance proof; add the mechanical launch-rate candidate promotion pipeline; add the size-class tier commercial policy to the rate book; admit rate-grade evidence in the launch catalog and coverage audit
- **Reliability and operations:** admit an image-witnessed authoritative zero at settlement; align runner authorization with smoke protocol; async-surface reservation parity — video guard, byte cap, tokenizer allowance; bound gateway cold-start recovery; cap identity convergence retry seam
- **Performance:** index the rotation-successor probe on api_keys

### Features

* **cloud:** activate bootstrap-v2 pricing in prod
* **cloud:** add --skip for recorded vendor-convergence skips
* **cloud:** add cold-start live acceptance proof
* **cloud:** add the mechanical launch-rate candidate promotion pipeline
* **cloud:** add the size-class tier commercial policy to the rate book
* **cloud:** admit rate-grade evidence in the launch catalog and coverage audit
* **cloud:** batch/jobs generation via buffered single-terminal settlement
* **cloud:** bound and measure the sealed cold-start subsidy
* **cloud:** carry an authoritative-dimension claim set on UnitCounts
* **cloud:** carry the settled charge into realtime usage blocks
* **cloud:** commit reviewed launch allocations
* **cloud:** editable per-key spend limits + monthly auto-recharge cap
* **cloud:** editable per-key spend limits + monthly recharge cap
* **cloud:** GA audio on /v1/extract — book-driven whisper pricing
* **cloud:** gate production on a degraded staging acceptance
* **cloud:** generation input-token reservation bound + settle-contract pins
* **cloud:** gpu_second sealed billing regime — custom models become sellable
* **cloud:** install digest-bound generation routes
* **cloud:** install digest-bound governed generation routes
* **cloud:** lease dynamic cold-start reporters
* **cloud:** low-balance threshold email outbox + SES seam
* **cloud:** low-balance threshold emails via SES
* **cloud:** metering GA — waves 1–3
* **cloud:** move the license exclusion verdict to the registry resolution seam
* **cloud:** org-scoped usage export API — CSV + JSON
* **cloud:** persist cold-start subsidy evidence
* **cloud:** POST /v1/estimate dry-run cost endpoint
* **cloud:** POST /v1/estimate dry-run cost endpoint + SDKs
* **cloud:** production-bootstrap-v2 — prefill closed, sealed lanes priced
* **cloud:** promote launch-rate candidates to rate-grade evidence
* **cloud:** rate-grade size-class tier book assembly + operator-gated activation
* **cloud:** response-final settlement for streaming generation
* **cloud:** storage GiB-day accrual + book-driven register quote
* **cloud:** surface credits_charged + rate_book_version in usage
* **cloud:** unfence OpenAI-compat /v1/audio/transcriptions
* **cloud:** unify the settled-charge name across jobs, batches, and SDKs
* **cloud:** weekly/daily reconciliation with paging alerts and operator-gated LIVE mode
* **cloud:** weekly/daily reconciliation, attribution categories and the operator LIVE gate
* **cloud:** widen whole-job settlement to every contract dimension
* **gateway:** add optional governed generation routing seam
* **model-ops:** add ASR discovery family
* **model-ops:** add fail-closed discovery family registry
* **model-ops:** add multimodal discovery families
* **model-ops:** add object detection discovery
* **model-ops:** add OCR discovery families
* **model-ops:** add report-only generation discovery
* **model-ops:** add VQA and KIE discovery
* **model-ops:** discover guardrail candidates
* **model-ops:** discover sparse retrievers
* **model-ops:** discover structured extraction candidates
* **model-ops:** discover text classification candidates
* **model-ops:** discover text rerankers
* **model-ops:** read the executor's trust clearance from the onboarding artifact
* **model-ops:** register guardrail discovery family
* **model-ops:** register structured extraction discovery
* **model-ops:** register text classification discovery
* **model-ops:** scan guardrail model candidates
* **model-ops:** scan structured extraction candidates
* **model-ops:** scan text classification candidates
* **model-ops:** support report-only family scans
* **model-ops:** wire the trust_remote_code scan into the AUTO gate
* **models:** onboard Qwen/Qwen3-Embedding-8B (scaffolded)
* **sdk:** estimate() on both SDKs for the /v1/estimate dry run
* **server:** bill sampled video frames as images
* **telemetry:** export the sealed cold-start metric remotely and chart it

### Bug Fixes

* **benchmark:** validate sorted deployment evidence
* **catalog:** bound GLiNER sequence lengths
* **ci:** reseal staging catalog evidence
* **cloud:** accept client-scoped WorkOS sessions
* **cloud:** accept Terraform canonical false in bootstrap guard
* **cloud:** accept WorkOS environment issuers
* **cloud:** address billing bootstrap review
* **cloud:** address CodeRabbit review — SES client bounds, send pacing, re-arm audit
* **cloud:** address CodeRabbit review on the /v1/estimate dry run
* **cloud:** address CodeRabbit review on the key-limit + recharge-cap PR
* **cloud:** address CodeRabbit review on the settled-charge surfaces
* **cloud:** address CodeRabbit review on the usage export
* **cloud:** address prerequisite gate review
* **cloud:** address the CodeRabbit review on the batch/jobs generation PR
* **cloud:** address the CodeRabbit round on the metering GA integration
* **cloud:** address the second CodeRabbit round on /v1/estimate
* **cloud:** admit an image-witnessed authoritative zero at settlement
* **cloud:** admit exact task-role tag bootstrap
* **cloud:** admit observed Modal US locations
* **cloud:** align benchmark smoke with rated profile
* **cloud:** align runner authorization with smoke protocol
* **cloud:** allow nested ECR repositories
* **cloud:** async-surface reservation parity — video guard, byte cap, tokenizer allowance
* **cloud:** attest benchmark runtime identity
* **cloud:** bind a rate_grade key to the price and hardware its envelope measured
* **cloud:** bind nightly health checks to staging
* **cloud:** bind paused forwarder candidate before prod cutover
* **cloud:** bind the charge-slot backstop to the stream, not the disconnect grace
* **cloud:** block the generation in/out split on a named measurement
* **cloud:** bound gateway cold-start recovery
* **cloud:** bound nightly evidence execution
* **cloud:** bound the sealed resolver cache and cap proxy response bodies
* **cloud:** cap identity convergence retry seam
* **cloud:** claim the hold before releasing it, and fold case like the registry
* **cloud:** classify benchmark evidence failures
* **cloud:** close cold recovery review gaps
* **cloud:** close cold-start evidence failure modes
* **cloud:** close final metering review faults
* **cloud:** close meter-only review gaps
* **cloud:** close metering GA review findings
* **cloud:** close rate evidence review gaps
* **cloud:** close reporter lease invariants
* **cloud:** close six gate findings on the /v1/estimate dry run
* **cloud:** close the batch-2a gate findings on editable limits + recharge cap
* **cloud:** close the CodeRabbit round on the candidate-promotion branch
* **cloud:** close the CodeRabbit round on the integration head
* **cloud:** close the CodeRabbit round on the reconciliation branch
* **cloud:** close the CodeRabbit round on the waves-1-3 integration head
* **cloud:** close the final CodeRabbit round on /v1/estimate
* **cloud:** close the gate findings on the resolution-seam verdict
* **cloud:** close the promotion gate's operator-assertion escapes
* **cloud:** close the second CodeRabbit round on the credit surfaces
* **cloud:** CodeRabbit review — extract params must be explicit per line
* **cloud:** CodeRabbit review fixes — bounded sealed-class cache, doc reconciliation
* **cloud:** CodeRabbit review fixes — quote authority degradation, single-model projection
* **cloud:** CodeRabbit review fixes — sealed posture floor, generator provenance
* **cloud:** CodeRabbit round 2 — evidence doc, strict decode, single query walk
* **cloud:** CodeRabbit round 3 — exhaustive ceiling match, honest parse row
* **cloud:** CodeRabbit round on the control-plane prerequisite gate
* **cloud:** CodeRabbit round-2 — redact resolver errors, tighten class guard
* **cloud:** CodeRabbit round-6 — quote survives a dead pool, pin e2e clock
* **cloud:** CodeRabbit round-7 — atomic cache admission, sweep resilience
* **cloud:** compile exact generation customer routes
* **cloud:** compile exact generation routes
* **cloud:** correct console wallet path to /console/usage — verified against sie-web
* **cloud:** correct the race() return annotation in the snapshot test
* **cloud:** cover regular fleet cold starts
* **cloud:** cover regular Modal fleet cold starts
* **cloud:** decide license exclusion at the registry resolution seam
* **cloud:** declare generation input_tokens in the launch catalog
* **cloud:** default customer charging to disabled
* **cloud:** disable recharge schedule without teardown
* **cloud:** distinguish authoritative zero from missing count at settlement
* **cloud:** enforce the image witness on both sides of the rating boundary
* **cloud:** expose collector test configs to non-root
* **cloud:** fail a tier book whose price point has no exact Metronome decimal
* **cloud:** fail closed on incomplete preflight reads
* **cloud:** fail closed when a declared skip cannot be recorded
* **cloud:** fail loudly when test prerequisites vanish
* **cloud:** gate batch usage on settlement evidence, not a running counter
* **cloud:** gate estimates on resolved model identity
* **cloud:** gate fixes for the usage export — day-sliced scan, pool isolation, uniform masking
* **cloud:** harden cold-start acceptance proof
* **cloud:** harden dynamic cold-start membership
* **cloud:** harden staging release verification
* **cloud:** harden the multipart licensing gate and lock the route wiring
* **cloud:** index count-gated SES resources via one()/splat
* **cloud:** isolate benchmark evidence runtime
* **cloud:** log truncated usage-export streams; pin legacy-row null passthrough
* **cloud:** make extract canonical for Docling billing
* **cloud:** make generation rate evidence prefix-cache safe
* **cloud:** make test prerequisites and tracing deterministic
* **cloud:** make the jobs loopback estimate allowance-aware and prove the guard end to end
* **cloud:** make the jobs planner uphold the invariant settlement now enforces
* **cloud:** make the reconciliation bars statements the cost basis supports
* **cloud:** make the response cap operative, unbilled, and honest
* **cloud:** make the witnessed zero reach the wire and cover both rails
* **cloud:** merge candidate promotion and close #2628 review
* **cloud:** mirror the widened model canonicalization in the control plane
* **cloud:** observe the sealed cold-start metric; scope it to the prometheus path
* **cloud:** parse the multipart model under the route's real body limit
* **cloud:** pin benchmark credential versions
* **cloud:** pin the sealed tariff evidence; harden the split and the posture tests
* **cloud:** pin the storage eligibility cutoff to UTC
* **cloud:** preflight input token rate bounds
* **cloud:** preserve cold-start evidence gaps
* **cloud:** preserve inactive prod rollback tasks
* **cloud:** price sealed-lane memory at the Sandbox rate, not standard
* **cloud:** propagate connector request identity
* **cloud:** publish a committed charge even when settlement then faults
* **cloud:** pull the authoritative invoice inside the single-flight lock
* **cloud:** recipient chain, re-arm edge, staleness bound, per-row commit
* **cloud:** reconcile what the batchgen merge hid from git
* **cloud:** record an admission outcome on every licensing rejection
* **cloud:** redact benchmark operator preflight
* **cloud:** reject a negative storage-accrual lookback
* **cloud:** reject an explicit storage-accrual day that is not closed
* **cloud:** reject expired recovery evidence
* **cloud:** reject null WorkOS client claims
* **cloud:** release the hold when the gateway cannot deliver the result
* **cloud:** renumber monthly-cap migration to 0040 — 0039 taken by sibling W3 branch
* **cloud:** renumber rotation-successor index migration to 0043 — 0042 taken by #2580
* **cloud:** report diagnostic cleanup failure
* **cloud:** report every live-preflight failure in one pass
* **cloud:** report malformed release pointers
* **cloud:** reseal release catalog dependents
* **cloud:** restore the spend-limit error contract on the W3 base
* **cloud:** retain exact benchmark deployment evidence
* **cloud:** retain exact benchmark deployment evidence\n\nRefs #2573
* **cloud:** retry benchmark identity convergence
* **cloud:** round-8 minors — env var name, authority restore, day-sweep guard
* **cloud:** route expected no-charge stream releases off the fault alerts
* **cloud:** route readiness failures through the gate without widening the opt-out
* **cloud:** route the async reservation through the shared video and token seams
* **cloud:** scope authoritative zero evidence
* **cloud:** scope rate preflight to exact authority
* **cloud:** scope the generation input bound to books that price input
* **cloud:** seal generation campaign artifact paths
* **cloud:** tell the truth about the video byte rail, and bound the tests that prove it
* **cloud:** tighten console origin validation, pace off the previous row
* **cloud:** tighten evidence diagnostic contract
* **cloud:** treat an explicit null usage as an empty block, not a refusal
* **cloud:** treat an explicitly null video as absent, not as a video
* **cloud:** type-gate the chat_template_kwargs allowlist
* **cloud:** unblock billing bootstrap migrations
* **cloud:** unseal console-free release subsets and offload-smoke
* **cloud:** unseal sealed release subsets, report every preflight failure, and gate degraded acceptance
* **cloud:** valid IAM tag charset + keep bootstrap secret-gate count at 2
* **cloud:** validate tier scheme inputs canonically
* **cloud:** wire the meter-only operator gate
* **gateway:** cap embeddings requests at 256 inputs
* **gateway:** harden governed generation dispatch
* **gateway:** preserve self-host grammar routing
* **loadtest:** use job token for GHCR pulls
* **model-ops:** accept GLiNER discovery models
* **model-ops:** address authorization review feedback
* **model-ops:** address evidence review feedback
* **model-ops:** authenticate the evidence producer and bind the identity to execution
* **model-ops:** bind CodeRabbit review completion
* **model-ops:** bind crash fallback repository
* **model-ops:** bind executor authorization
* **model-ops:** bind trusted completion status
* **model-ops:** close final trust gate review gaps
* **model-ops:** close malformed trust evidence gap
* **model-ops:** close three bypasses in the trust_remote_code gate
* **model-ops:** dispatch the Stage 4-5 gate completion explicitly
* **model-ops:** distinguish disabled discovery families
* **model-ops:** exclude OCR-only KIE models
* **model-ops:** explicitly dispatch discovery handoffs
* **model-ops:** gate readiness on bot reviews
* **model-ops:** harden bot review handoff
* **model-ops:** harden cold Modal image builds
* **model-ops:** harden completion status fallbacks
* **model-ops:** harden report-only scanning
* **model-ops:** harden trust gate provenance
* **model-ops:** keep a failed gate origin recoverable
* **model-ops:** narrow document OCR compatibility
* **model-ops:** never log gh's stderr, which carries GH_TOKEN by construction
* **model-ops:** preserve completed review mutations
* **model-ops:** re-derive the clearance in matrix-block, bind evidence to the revision
* **model-ops:** re-grant contents:read to the dispatch job
* **model-ops:** recognize Copilot bot aliases
* **model-ops:** reconcile gate recovery reviews
* **model-ops:** refuse an unset RUNNER_TEMP instead of falling back into the workspace
* **model-ops:** refuse instead of crashing when the evidence fetch faults
* **model-ops:** refuse unusable evidence instead of falling back to blockers
* **model-ops:** reject ambiguous family matches
* **model-ops:** reject failed Copilot reviews
* **model-ops:** repair gate recovery replay
* **model-ops:** retry transient Modal build downloads
* **model-ops:** scan all loader-visible code
* **model-ops:** scan zero-shot classification models
* **model-ops:** suppress dynamic artifact failure logging
* **model-ops:** tolerate pending status retries
* **model-ops:** validate gate head before status writes
* **model-ops:** validate review completion evidence
* **model-ops:** wake the adapter executor from an agentless dispatch job
* **model-ops:** wake the adapter-exec executor by explicit dispatch
* **quality-eval:** retain models pending vanilla floors
* **quality-eval:** surface unfloored matrix rows
* **release:** bootstrap a fresh production database
* **release:** compare fresh proof types exactly
* **release:** retry malformed Modal attestations
* **release:** retry transient Modal attestation output
* **release:** verify empty replay target digest
* **server:** address CodeRabbit review on the video meter
* **server:** bill SigLIP padded text work
* **server:** bound isolation fan-out and close the NaN decode escapes

### Performance Improvements

* **cloud:** index the rotation-successor probe on api_keys

## v0.6.24 (2026-07-26)

### Highlights

- **New capabilities:** add agent shell GPU fallbacks; add agent shell placement controls; add launch catalog contract canary; add Modal-native placement primitives; add native placement primitives; allow snapshotless native lanes
- **Reliability and operations:** align rate-book recovery checks; harden dispatch preflight gate; harden final generation gates; harden retained release recovery; preserve capacity diagnostics on timeout races
- **Performance:** cache MUVERA projection state

### Features

* **cloud:** add agent shell GPU fallbacks
* **cloud:** add agent shell placement controls
* **cloud:** add launch catalog contract canary
* **cloud:** add Modal-native placement primitives
* **cloud:** add native placement primitives
* **cloud:** allow snapshotless native lanes
* **cloud:** attest realized placement identity
* **cloud:** establish broad-native release baseline
* **cloud:** govern prod legacy billing cutover
* **cloud:** preflight offline tokenizer dependencies
* **server:** pinnable LoRA adapter revisions in lora_paths

### Bug Fixes

* **cloud:** accept scheduled orphan-hold sweeps
* **cloud:** address native placement review
* **cloud:** address release review feedback
* **cloud:** align rate-book recovery checks
* **cloud:** align tokenizer snapshot dependencies
* **cloud:** allow absent generation in final smoke
* **cloud:** allow bounded prod platform bootstrap
* **cloud:** attest final gateway replacement
* **cloud:** bound snapshot lifecycle boots
* **cloud:** bound worker identity collection
* **cloud:** canonicalize generation GPU profiles
* **cloud:** close bounded prod bootstrap plan
* **cloud:** close launch review findings
* **cloud:** declare bootstrap alert container
* **cloud:** default orphan hold sweep ttl
* **cloud:** derive exact rate topology from broad release
* **cloud:** derive gateway gpu admission from plan
* **cloud:** diagnose bounded worker capacity waits
* **cloud:** drain legacy generation revisions
* **cloud:** enforce global canary request identity
* **cloud:** fail closed on tokenizer pin drift
* **cloud:** finalize gateway after generation deploy
* **cloud:** frame reusable candidate outputs
* **cloud:** gate billable release commit safely
* **cloud:** gate every gateway replica on dispatch readiness
* **cloud:** guard broad placement pending realized identity
* **cloud:** harden dispatch preflight gate
* **cloud:** harden final generation gates
* **cloud:** harden retained release recovery
* **cloud:** integrate launch remediation release gates
* **cloud:** keep slow batch leaders warm
* **cloud:** make Vercel deploy creation transport-safe
* **cloud:** package excluded gateway target
* **cloud:** package gateway from crate target
* **cloud:** preflight gateway generation dispatch
* **cloud:** preflight route and worker hash parity
* **cloud:** preserve capacity diagnostics on timeout races
* **cloud:** preserve launch route errors
* **cloud:** preserve snapshot cancellation cleanup
* **cloud:** reconcile launch remediation release gates
* **cloud:** recover ambiguous Vercel deploy creates
* **cloud:** recover bounded prod bootstrap outputs
* **cloud:** recover rate-book releases drain-first
* **cloud:** redact readiness diagnostics
* **cloud:** refresh staging catalog evidence digests
* **cloud:** rehearse immediate launch accounts
* **cloud:** reject partial readiness placement
* **cloud:** reserve token inputs from governed bounds
* **cloud:** respect Better Stack token scope
* **cloud:** restore authenticated gateway rollback health
* **cloud:** restrict rehearsal recovery fields
* **cloud:** serialize bootstrap state reads
* **eval:** make run-eval-group.sh and its tests portable to macOS
* **eval:** reject symlink escapes from the /tmp guard; reap the requirements temp dir
* **loadtest:** use valid AWS IAM purpose tag
* **models:** declare tokenizer dependencies
* **quality:** address remaining review findings
* **quality:** close nightly model gaps and add exact quarantine
* **quality:** gate bounded NVIDIA MUVERA checks
* **quality:** isolate parameterized FiNER evals
* **quality:** preserve floor metric contracts
* **quality:** record bounded concurrency-one targets
* **quality:** record faithful CMedQA floors
* **quality:** record faithful quality floors
* **quality:** reject non-finite qrel scores
* **quality:** restore bounded FiNER coverage
* **quality:** restore Mixedbread query expansion
* **quality:** stabilize scoped PR quality runs
* **release:** close prod bootstrap output contract
* **release:** redact bootstrap state read failures
* **smoke:** restore blocking-probe assertions and cold-touch retry coverage
* **worker:** copy sie_telemetry crate into worker image builds

### Performance Improvements

* **server:** cache MUVERA projection state

## v0.6.23 (2026-07-24)

### Highlights

- **New capabilities:** accept native multimodal generation; support routed media measurement campaigns; add production bootstrap rate book; add pre-public task measurement candidates; bind launch catalog to commercial rate book; bind launch rates to commercial book
- **Reliability and operations:** harden generation quality sharding; harden multimodal batch validation; harden queue readiness for generation defaults; harden shard failure capture; restore governed retrieval floors
- **Performance:** bind bounded extraction evidence; bind current nonvision launch evidence; recover granite guardian queue evidence; reuse siglip preprocessing

### Features

* **api:** accept native multimodal generation
* **bench:** support routed media measurement campaigns
* **billing:** add production bootstrap rate book
* **catalog:** add pre-public task measurement candidates
* **catalog:** bind launch catalog to commercial rate book
* **catalog:** bind launch rates to commercial book
* **catalog:** prepare launch serving evidence
* **catalog:** prepare non-public launch catalog checkpoint
* **catalog:** record current nonvision evidence
* **catalog:** record Snowflake and R3 evidence
* **catalog:** validate native multimodal recipes
* **catalog:** verify document vision recipes
* **catalog:** verify existing launch recipes
* **catalog:** verify GLiNER2 large launch evidence
* **catalog:** verify R3 pipeline on L4
* **catalog:** verify redaction task evidence
* **catalog:** verify Snowflake on preferred L4
* **ci:** complete req14 gates after lane runs
* **ci:** remove runner-held waits from official-recipe gates
* **cloud:** activate accounts and route launch events
* **cloud:** add bounded prod platform bootstrap
* **cloud:** add exact launch rate campaign suite
* **cloud:** add immutable managed release pipeline
* **cloud:** add launch account credit flow
* **cloud:** add launch rate campaign suite
* **cloud:** add Qwen task evidence campaign
* **cloud:** add secure staging catalog evidence dispatch
* **cloud:** add secure staging evidence dispatch
* **cloud:** archive Modal measurement releases
* **cloud:** assemble commercial rates from exact evidence
* **cloud:** authenticate Modal rate evidence
* **cloud:** bracket generation rate window
* **cloud:** compile isolated measurement topology
* **cloud:** derive commercial rates from exact resource seconds
* **cloud:** derive commercial rates from Modal billing
* **cloud:** execute document vision recipes
* **cloud:** execute remaining reviewed recipes
* **cloud:** route signup notifications directly to Slack
* **cloud:** verify BGE and GTE launch evidence
* **deploy:** add scalr-bootstrap root for internal terraform state
* **deploy:** run scalr-bootstrap under opentofu with live scalr backend
* **deploy:** scalr-held terraform state for internal roots
* **deploy:** split scalr workspaces into cli- and vcs-driven
* **generation:** type structured output contracts
* **model-ops:** split feasibility PROCEED-WITH-ADAPTER-AUTO vs human
* **model-ops:** trust_remote_code security pre-gate
* **quality-eval:** in-pipeline wall-clock stop with partial report
* **sdk:** add typed Responses client

### Bug Fixes

* **1841:** stamp sealed metric units on the Modal collector, and make the CP filters cover what the suite reads
* **adapters:** serialize output_hidden_states forwards — close the #2144 recorder-race class
* **adapters:** serialize output_hidden_states forwards to close the #2144 recorder-race class
* **api:** address multimodal contract review
* **api:** close multimodal parity gaps
* **bench:** add faithful PyLate official recipe
* **bench:** address launch evidence review
* **bench:** align Any2Any dispatch semantics
* **bench:** align ColBERT reference validation
* **bench:** align pinned dataset metadata
* **bench:** bind generation evidence to runtime protocol
* **bench:** bind Modal measurement provenance
* **bench:** bind Qwen3.6 launch evidence to runtime identity
* **bench:** canonicalize campaign route identities
* **bench:** clean up failed server startups
* **bench:** close operation evidence review gaps
* **bench:** discard partial server results
* **bench:** emit strict quality JSON
* **bench:** fail closed on benchmark operation identity
* **bench:** forward only encode/score to EvalRunner in quality path
* **bench:** harden generation quality sharding
* **bench:** harden multimodal batch validation
* **bench:** harden queue readiness for generation defaults
* **bench:** harden shard failure capture
* **bench:** keep image types out of runtime imports
* **bench:** pin ragged PyLate reranking fix
* **bench:** pin trusted dataset loaders
* **bench:** preserve canonical FewRel dataset identity
* **bench:** preserve generation evidence imports
* **bench:** preserve invalid campaign archives
* **bench:** preserve profile benchmark identity
* **bench:** preserve safe Modal shard failures
* **bench:** reject extraction item errors
* **bench:** reject incompatible reference operations early
* **bench:** reject incomplete shard assembly
* **bench:** reject partial performance measurements
* **bench:** remove stale generation import
* **bench:** report empty perf evidence explicitly, not as ambiguous
* **bench:** restore governed retrieval floors
* **bench:** route profiled performance requests
* **bench:** satisfy review static analysis
* **bench:** start named profile routes
* **bench:** stop plural options poisoning choice extraction
* **bench:** type bounded image dispatch
* **bench:** validate multimodal encode batches
* **bench:** version routed rate candidates
* **billing:** bind Modal tags to pretraffic contract
* **billing:** bind zero-page job evidence
* **billing:** preserve zero-page job evidence
* **billing:** project observed units through reservation plans
* **billing:** restrict zero terminals to pages
* **billing:** verify multimodal fixture units
* **catalog:** bind named output evidence
* **catalog:** disable unsafe qwen speculative default
* **catalog:** harden native acceptance validation
* **catalog:** raise GTE Muvera repetitions
* **catalog:** route mxbai edge through native attention
* **catalog:** use pinned GTE query length
* **catalog:** validate named profile evidence
* **ci:** authenticate quality dependency sync
* **ci:** diff quality evals from current PR base
* **ci:** fail closed on mixed Modal output
* **ci:** isolate privileged req14 completion
* **ci:** isolate quality eval evidence
* **ci:** isolate quality runner credentials
* **ci:** let scalr-bootstrap track the opentofu pin
* **ci:** repair Qwen quality shard workflow
* **ci:** verify trusted req14 callback checkout
* **cloud:** accept canonical Modal volume mount
* **cloud:** accept Modal GCP provider identity
* **cloud:** address release workflow review follow-ups
* **cloud:** align launch evidence contracts
* **cloud:** align release lock namespace
* **cloud:** align release probes with scoped listing
* **cloud:** allow release image manifest readback
* **cloud:** allow runtime secret namespace
* **cloud:** allow scoped release key discovery
* **cloud:** allowlist release child environments
* **cloud:** approve permissive launch licenses
* **cloud:** attest LightOnOCR image compatibility
* **cloud:** attest Modal snapshot memory limits
* **cloud:** bind billing to executing Modal app
* **cloud:** bind billing to Modal environment
* **cloud:** bind commercial evidence to routed candidates
* **cloud:** bind ECR login to OIDC identity
* **cloud:** bind rate probes to archived model pins
* **cloud:** bind score probes to campaign batches
* **cloud:** bind topology plan semantics
* **cloud:** bootstrap purpose-bound Slack sinks
* **cloud:** bound and isolate release preflight reads
* **cloud:** classify ECS start failures first
* **cloud:** complete managed release operability follow-ups
* **cloud:** complete release operability audit
* **cloud:** dispatch exact catalog profiles
* **cloud:** dispatch quality shards in Modal batch mode
* **cloud:** fail closed in candidate image lookup
* **cloud:** fail closed on ECR lookup errors
* **cloud:** finish release child isolation
* **cloud:** finish release preflight hardening
* **cloud:** guard acceptance projection freshness
* **cloud:** guard catalog projection freshness
* **cloud:** harden commercial evidence joins
* **cloud:** harden launch identity and approval invariants
* **cloud:** harden launch rate campaign probes
* **cloud:** harden managed release promotion
* **cloud:** harden notification launch paths
* **cloud:** harden screenshot acceptance
* **cloud:** harden vision acceptance bounds
* **cloud:** include smart chat task evidence
* **cloud:** isolate launch credit approvals
* **cloud:** keep LightOn preflight CPU-safe
* **cloud:** make topology imports typecheck
* **cloud:** map static lanes to catalog pools
* **cloud:** measure both SigLIP towers
* **cloud:** pin batch worker execution identity
* **cloud:** pin dev worker identity
* **cloud:** pin launch measurement resources
* **cloud:** pin native worker execution identity
* **cloud:** pin staging batch memory limit
* **cloud:** preflight live release dependencies
* **cloud:** preflight Modal SGLang runtime
* **cloud:** preflight release task arguments
* **cloud:** preserve concurrent topology output
* **cloud:** preserve governed release tool paths
* **cloud:** preserve sidecar image contract
* **cloud:** prime catalog before gateway rollout
* **cloud:** probe exact campaign request batches
* **cloud:** reconcile commercial rates with launch profiles
* **cloud:** reconcile measurement topology with launch catalog
* **cloud:** refresh staging evidence pins
* **cloud:** reject duplicate generation shards
* **cloud:** reject malformed release metadata
* **cloud:** reject mismatched catalog evidence
* **cloud:** reject undecodable evidence
* **cloud:** reject unversioned S3 evidence
* **cloud:** require complete pii redaction proof
* **cloud:** reseal licensed launch release locks
* **cloud:** reseal production rate release locks
* **cloud:** restore CUDA toolkit to SGLang sidecars
* **cloud:** retain published measurement evidence
* **cloud:** retire legacy environment ECR state
* **cloud:** scope Better Stack credentials by product
* **cloud:** separate guardrail policy fixtures
* **cloud:** use lightweight generation evidence contract
* **cloud:** validate launch evidence integrity
* **cloud:** verify control-plane runtime dependencies
* **cloud:** verify Modal markers from files
* **cloud:** verify redact pii acceptance
* **cloud:** verify redact PII postprocessing
* **deploy:** pin internal scalr workspaces to opentofu
* **gateway:** validate schema keywords by context
* **generation:** close structured grammar review gaps
* **grammar:** bound pointer index parsing
* **grammar:** harden schema normalization
* **loadtest:** manage github-token in terraform + sync via ESO, pin sie-perf-lab ref
* **loadtest:** persist github-token in terraform, sync via ESO, pin perf-lab ref
* **loadtest:** remove illegal variable interpolation from output description
* **model-ops:** close review gaps in the AUTO verdict contract
* **model-ops:** require an explicit matrix_task key in auto_authoring
* **profiles:** honor routed outputs and retrieval recipes
* **quality-eval:** address Copilot review on the wall-clock stop
* **quality-eval:** remove shadow tree on exception paths, not just returns
* **quality-eval:** select governed official metric
* **quality-eval:** serve non-default smoke bundles from a pip --target shadow so the overlay reaches the bundle floor
* **quality-eval:** validate wall_budget_minutes inside run_pipeline
* **release:** authenticate gateway rollback probe
* **release:** keep sie-audio-prep internal to builds
* **sdk:** accept nullable grammar metadata
* **sdk:** keep grammar types on supported module path
* **security:** include aws edge in standing review
* **security:** pin hf_revision on the 10 unpinned trust_remote_code models
* **server:** apply trained Rotary ColBERT projection
* **server:** defer LightOn OCR compatibility patch
* **server:** hide unsupported visual muvera profiles
* **server:** mask ColBERT expansion keys in flash attention
* **server:** meter SigLIP tokens from attention masks
* **server:** pin launch guardrail policies
* **server:** preserve stream helper compatibility
* **server:** preserve trusted code revision on fallback
* **server:** restore ColBERTv2 retrieval recipe
* **server:** restore Jina ColBERT retrieval recipe
* **server:** serialize remaining tokenizer access
* **sie_server:** register remote-code snapshot dir on sys.path for ST-Router models
* **sie-bench:** isolate PyLate reference evaluation
* **sie-bench:** preserve expansion token attention
* **smoke:** fail loud when a bundle overlay undershoots its declared pins
* **terraform:** manage github-token container only, drop placeholder version
* **tooling:** derive overlay cache version from lock
* **tooling:** gate Modal quality shard fanout
* **tooling:** harden launch failure handling
* **tooling:** isolate Modal shell worktree environment
* **tooling:** make empty Modal status succeed
* **tooling:** mount tracked playbook fixtures on Modal
* **tooling:** namespace Modal JIT caches by ABI
* **tooling:** preserve Modal status errors
* **tooling:** preserve multiline Modal commands
* **tooling:** read Modal lock from mounted workspace
* **tooling:** resolve modal from task path
* **tooling:** scope Modal inventory conflicts
* **tooling:** validate Modal description types
* **tooling:** version Modal datasets cache
* **vision:** harden launch evidence
* **vision:** pad fused detection batches

### Performance Improvements

* **catalog:** bind bounded extraction evidence
* **catalog:** bind current nonvision launch evidence
* **generation:** recover granite guardian queue evidence
* **vision:** reuse siglip preprocessing

## v0.6.22 (2026-07-22)

### Highlights

- **New capabilities:** bind exact rate measurement windows; shard OCR quality by page; support package revision evidence; attest worker execution identity; fan out quality shards on Modal; pin immutable Docling artifacts
- **Reliability and operations:** harden OCR shard provenance; harden package revision evidence; fail fast on managed gateway lock drift; attest OCR shard execution identity; clarify complete rate evidence
- **Performance:** concurrent OCR page dispatch so server batching actually engages

### Features

* **bench:** bind exact rate measurement windows
* **bench:** shard OCR quality by page
* **bench:** support package revision evidence
* **cloud:** attest worker execution identity
* **cloud:** fan out quality shards on Modal
* **server:** pin immutable Docling artifacts
* **server:** support immutable Docling model artifacts
* **tooling:** attach named Modal shell secrets

### Bug Fixes

* **bench:** attest OCR shard execution identity
* **bench:** clarify complete rate evidence
* **bench:** clarify rate-window mismatch errors
* **bench:** harden OCR shard provenance
* **bench:** harden package revision evidence
* **bench:** mark serving basis non-commercial
* **bench:** require exact rate resource evidence
* **ci:** install sdk deps for catalog tests
* **ci:** queue release rerun when main moves
* **ci:** regenerate launch catalog on release PR
* **ci:** regenerate release PR when main moves
* **cloud:** bind benchmark evidence to deployed workers
* **cloud:** bind benchmark model identity
* **cloud:** bind quiesced ECS candidates before scaling
* **cloud:** bind reviewed catalog recipe executors
* **cloud:** bind worker resource attestation
* **cloud:** fail fast on managed gateway lock drift
* **cloud:** include compiler version dependency
* **cloud:** isolate catalog compiler dependencies
* **cloud:** meter detection catalog smoke images
* **cloud:** preserve deployment validation ordering
* **cloud:** preserve identity across billing handoff
* **cloud:** reject noninteger GPU identity counts
* **cloud:** remove member key spend ceiling
* **cloud:** require clean governed release sources
* **cloud:** validate integral worker identity fields
* **cloud:** validate managed gateway lock before release
* **eval:** constrain temporary artifact root
* **eval:** serve canonical profile routes
* **tooling:** bind release locks from clean head
* **tooling:** close Modal fanout lifecycle races
* **tooling:** fail closed on Modal volume sync
* **tooling:** isolate Modal uv cache
* **tooling:** isolate quality shards on Modal v2
* **tooling:** keep active modal shells alive
* **tooling:** preserve Modal fanout failures
* **tooling:** reject symlink fanout outputs
* **tooling:** resolve modal in detached shells

### Performance Improvements

* **sie-bench:** concurrent OCR page dispatch so server batching actually engages

## v0.6.21 (2026-07-21)

### Highlights

- **New capabilities:** add rate-grade evidence contract; add reusable local queue stack; add upstream GLiNER NER reference; govern vision launch evidence; seed load-test request selection; SIESparseImageWrapper for served visual sparse retrieval
- **Reliability and operations:** capture generation timeout evidence; harden local queue stack lifecycle; drop stale metrics registry dependency; finalize billing money hardening; harden account lifecycle rehearsal
- **Performance:** record SPLADE 512-token run; record extraction L4 measurements; compact sparse output in native precision; fuse packed SPLADE pooling

### Features

* **bench:** add rate-grade evidence contract
* **bench:** add reusable local queue stack
* **bench:** add upstream GLiNER NER reference
* **bench:** govern vision launch evidence
* **bench:** seed load-test request selection
* **bench:** SIESparseImageWrapper for served visual sparse retrieval
* **billing:** meter multimodal catalog units
* **candle:** add and optimize native SPLADE sparse encoding
* **candle:** add native SPLADE sparse encoding
* **cloud:** add exact qwen35 perf probe
* **cloud:** add fail-closed catalog acceptance gate #1943
* **cloud:** add fail-closed catalog acceptance runner
* **cloud:** add fast chat catalog acceptance
* **cloud:** add managed-edge abuse and tenant gates
* **cloud:** add OCR and speech acceptance probes
* **cloud:** add operator key administration
* **cloud:** add operator key administration and keyed usage
* **cloud:** add OWLv2 acceptance executor
* **cloud:** add OWLv2 catalog acceptance
* **cloud:** add staging account lifecycle rehearsal
* **cloud:** add staging rehearsal rate authority
* **cloud:** audit measured launch rate-book coverage
* **cloud:** audit measured rate book coverage
* **cloud:** bind verified DINO and Whisper evidence
* **cloud:** declare exact lane placement resources
* **cloud:** declare exact lane resources
* **cloud:** enforce conserved usage rating
* **cloud:** enforce exact managed catalog authority
* **cloud:** execute text acceptance recipes
* **cloud:** exercise DINO catalog acceptance
* **cloud:** expose operator billing health
* **cloud:** gate catalog tasks on evidence
* **cloud:** govern launch catalog topology
* **cloud:** govern vision launch catalog
* **cloud:** integrate control-plane launch readiness
* **cloud:** integrate exact request billing authority
* **cloud:** notify operators of new signups
* **cloud:** pack Whisper into shared launch pool
* **cloud:** record generation launch evidence
* **cloud:** record Qwen3.5 launch evidence
* **cloud:** rotate staging rehearsal rate book
* **cloud:** serve release-locked model catalog
* **cloud:** serve release-locked model catalog #1944
* **cloud:** settle exact pair and generation usage
* **cloud:** ship proprietary vision launch catalog
* **cloud:** validate extraction catalog recipes
* **cloud:** validate structured output acceptance
* **cloud:** verify qwen35 A100 model evidence
* **cloud:** wire privacy-safe console RUM
* **cluster:** opt-in bake cache write via SIE_BAKE_CACHE_WRITE
* **core:** expose authoritative request usage
* **core:** report authoritative score pair usage
* **evidence:** propagate worker execution identity
* **extract:** add pinned PII quality gate
* **gateway:** add strict Cohere rerank compatibility
* **model-ops:** feasibility accepts transformers-5-line bundle choice
* **model-ops:** scaffold re-renders cloud release locks with the model YAML
* **models:** onboard Snowflake/snowflake-arctic-embed-s (scaffolded)
* **models:** raise qwen36 launch context to 8k
* **playbooks:** upstream SparseEncoder reference arm for sparse-visual models
* **quality-eval:** wire the signature-DB loop — adjudicated-skip + verdict write-back
* **reranking:** aggregate launch runtime and release locks
* **reranking:** harden Qwen score runtime
* **sdk:** expose terminal billing metadata
* **sdk:** preserve terminal request metadata on errors
* serve visual-SPLADE via transformers5 (sparse-vision adapter + measurement path)
* **server:** bind package artifact identity
* **server:** bind package-backed model artifacts
* **server:** transformers514 bundle + SparseEncoder vision sparse adapter
* **sie-bench:** add generation quality sharding
* **sie-bench:** add generic generation quality sharding
* **telemetry:** centralize observability on OTLP
* **vision:** harden OSS runtime contracts
* **vision:** ship OSS launch runtime and evidence

### Bug Fixes

* **audio:** sandbox image builds the audio wheel opportunistically
* **bench:** align local generation lane routing
* **bench:** attest generation execution revision
* **bench:** bind broker artifact ARN
* **bench:** capture generation timeout evidence
* **bench:** clarify direct revision semantics
* **bench:** close execution identity provenance gaps
* **bench:** exercise generation queue sidecar
* **bench:** harden local queue stack lifecycle
* **bench:** isolate generation smoke stack lifecycle
* **bench:** isolate queue smoke ports
* **bench:** make fake stack no-op explicit
* **bench:** preserve ChemProt entity offsets
* **bench:** preserve exact reranker task identity
* **bench:** preserve NER task parameters
* **bench:** reject coerced execution identity
* **bench:** remove failure log content
* **bench:** render generation smoke help
* **bench:** resolve canonical AG News dataset
* **bench:** route canonical model profiles
* **bench:** select encode output families
* **bench:** separate model and machine profiles
* **bench:** serialize image rerank inputs
* **bench:** use SPIA field labels for PII eval
* **bench:** verify broker artifact bucket before reads
* **bench:** warm generation at measured concurrency
* **billing:** preserve legacy outbox drain
* **billing:** reject unexpected metronome rates
* **billing:** verify vendor-bound raw usage
* **candle:** bound CUDA compiler parallelism
* **candle:** include CUDA math constants for SPLADE
* **candle:** preserve tiny SPLADE log1p weights
* **ci:** bind cloud gateway cache to OSS source
* **ci:** install catalog compiler dependency
* **ci:** load sidecar image for release smoke tests
* **ci:** make cloud gateway checks deterministic
* **ci:** supersede obsolete pull request runs
* **ci:** validate release guard wait settings
* **ci:** wait for release check reruns
* **cloud:** accept numeric whisper fixture transcript
* **cloud:** address billing review findings
* **cloud:** address rate authority review
* **cloud:** align acceptance with native labels
* **cloud:** align rehearsal rates with measured audit
* **cloud:** benchmark qwen35 through native generate
* **cloud:** bind acceptance hash to source artifact
* **cloud:** bind exact lane cpu limits
* **cloud:** bind generation grammar pool
* **cloud:** bind structured schema semantics
* **cloud:** bound outbox poison isolation
* **cloud:** clarify idempotent deletion tests
* **cloud:** clean up metered rehearsal credits
* **cloud:** close console release preflight gaps
* **cloud:** close launch integration review gaps
* **cloud:** close launch review gaps
* **cloud:** close operator key launch gaps
* **cloud:** converge managed catalog exactly
* **cloud:** cover multimodal rehearsal units
* **cloud:** declare redact PII labels
* **cloud:** derive migration log stream
* **cloud:** drop stale metrics registry dependency
* **cloud:** enforce billing activation invariants
* **cloud:** finalize billing money hardening
* **cloud:** govern legacy billing cutover
* **cloud:** harden account lifecycle rehearsal
* **cloud:** harden billing money flows
* **cloud:** harden generation acceptance evidence
* **cloud:** harden launch rate evidence
* **cloud:** harden lifecycle rehearsal cleanup
* **cloud:** harden signup alert routing
* **cloud:** harden Stripe clawbacks and usage delivery
* **cloud:** include registry compiler dependency
* **cloud:** keep cleanup accounts suspended
* **cloud:** keep live secret values out of state
* **cloud:** keep proxy dry-run provenance current
* **cloud:** keep qwen35 measurements pending
* **cloud:** keep qwen35 performance pending
* **cloud:** match Better Stack canonical alert fields
* **cloud:** pin exact console release commits
* **cloud:** preserve charged structured failures
* **cloud:** preserve committed key recovery
* **cloud:** preserve custom catalog evidence
* **cloud:** preserve proxy token shape in dry-run
* **cloud:** rebase catalog acceptance gate
* **cloud:** reconcile exact metronome rating
* **cloud:** reconcile native vision label bindings
* **cloud:** refresh release lock artifact pins
* **cloud:** reject duplicate billing bindings
* **cloud:** reject empty key scopes
* **cloud:** reject malformed catalog exports
* **cloud:** reject unknown extraction recipes
* **cloud:** reject unsupported lane providers
* **cloud:** report gross Stripe funding
* **cloud:** require console release domain
* **cloud:** require ordered whisper transcript
* **cloud:** restore Metronome billing authority
* **cloud:** restore telemetry branch CI
* **cloud:** retain charged acceptance failures
* **cloud:** tighten billing terminal invariants
* **cloud:** tolerate inaccessible gateway candidates
* **cloud:** validate acceptance readiness reasons
* **cloud:** validate edge origin before secret sync
* **cloud:** validate lazy snapshot lanes
* **cloud:** validate modal cloud placement
* **cloud:** validate native label binding kinds
* **cloud:** validate rate tariff shape
* **colpali:** serialize model forwards — transformers output-recorder race
* **colpali:** stop GPU memory accumulation on Vidore evals
* **control-plane:** make wallet authority explicit
* **core:** preserve unit count wire order
* **deploy:** bound generation readiness process
* **deploy:** harden generation readiness gate
* **deploy:** honor local weight source precedence
* **deploy:** prove exact generation revision readiness
* **deploy:** resolve model staging weight sources
* **eval:** add faithful Qwen3-VL reference
* **eval:** bound multimodal reference scoring
* **eval:** bound reranker candidates by corpus
* **eval:** isolate VL reference dependencies
* **eval:** keep gate plan aligned with dispatch
* **eval:** preserve VL text coverage
* **eval:** reject empty reranker qrels
* **evidence:** bind immutable model identity
* **evidence:** close worker identity contract gaps
* **evidence:** harden worker identity boundaries
* **extract:** close review validation gaps
* **extract:** filter unconstrained relation endpoints
* **keda:** keep gateway coordination under autoscaling
* **keda:** simplify forward OTLP upgrade
* **keda:** support forward-only OTLP upgrade
* **keda:** verify lean forward upgrade
* keep audio prep release pin in sync
* **loadtest:** skip doomed bench-only scenarios when no release deployed
* **metering:** reject invalid score counts
* **model-ops:** align exact vision gate evidence
* **model-ops:** gate parses the trigger's 'Triggered by #N' issue reference
* **model-ops:** void gate attempts that provably began no spend
* **model:** restore SPLADE quality sequence cap
* **models:** pin hf_revision for gte-Qwen2-7B remote-code load
* **models:** restore SPLADE 512-token capacity
* **models:** serve gte-Qwen2-7B via its own bidirectional forward
* **playbooks:** v2-native task extraction + timeout contract for the sparse arm
* **quality-eval:** identity-based signature dedup + branch hash (#2096 review)
* **quality-eval:** make runner wake parallel, fail-loud, and demand-accurate
* **quality-eval:** reset stale WakeRequestedAt against launch time
* **quality-eval:** tolerate YAML date-typed fields in signature identity (#2096 review)
* **quality:** tolerate missing gpu identity tooling
* refresh managed release locks
* **release:** refresh deployment locks during release regeneration
* **release:** run deployment lock refresh in project env
* **reranking:** address exact-head review
* **reranking:** emit authoritative fake usage
* **reranking:** emit fake score usage
* **reranking:** reject blank compatibility inputs
* **server:** empty bundle model_filter must advertise zero models, not all
* **server:** remove unused fake adapter import
* **sie-bench:** harden generation sharding evidence
* **sie-tools:** keep session failures auth-scoped
* **sie-tools:** preserve account lifecycle errors
* **smoke:** align OTel collector contract checks
* **smoke:** align OTLP metric assertions
* **smoke:** restore kind smoke on main
* **splade:** align telemetry and model docs
* **splade:** harden profile and sparse defaults
* **tooling:** fail fast before Modal dataplane runs
* **tooling:** stop modal shells by app id
* **tools:** preserve baked mise in modal shells
* **vision:** address review edge cases
* **worker:** propagate effective sequence caps
* **worker:** release OOM tracebacks in the no-recovery branch too (CodeRabbit)

### Performance Improvements

* **benchmarks:** record SPLADE 512-token run
* **bench:** record extraction L4 measurements
* **candle:** compact sparse output in native precision
* **candle:** fuse packed SPLADE pooling
* **candle:** pack SPLADE BERT inference
* **vision:** bulk-convert detection outputs

## v0.6.20 (2026-07-18)

### Highlights

- **New capabilities:** add GLiGuard and fix extraction catalog behavior; add GLiGuard and fix LightOnOCR and GLiREL extraction; add cloud benchmark campaigns; add native R3-Skill evaluation; add staging cloud smoke path; add two-client staging rehearsal
- **Reliability and operations:** fence broker rollback recovery; harden cloud campaign evidence; harden cloud campaign readiness; harden live campaign preflight; harden managed campaign control plane
- **Performance:** cache exact bf16 gelu values; fuse ModernBERT exact GeGLU; optimize ModernBERT CUDA hot path; vectorize bf16 gelu lookup gate

### Features

* add GLiGuard and fix extraction catalog behavior
* add GLiGuard and fix LightOnOCR and GLiREL extraction
* **bench:** add cloud benchmark campaigns
* **bench:** add native R3-Skill evaluation
* **bench:** add staging cloud smoke path
* **bench:** add two-client staging rehearsal
* **bench:** complete managed cloud campaigns
* **bench:** cover cloud task modalities
* **bench:** inspect diagnostic campaign archives
* **bench:** operationalize cloud measurement campaigns
* **bench:** operationalize managed SIE Cloud measurement campaigns
* **ci:** fake-stack-test job — Fake Engine regression suite + Cloud dispatcher loopback
* **ci:** fake-stack-test job + five Fake Engine regression tests + Cloud loopback
* **cloud-gateway:** serve job chunk results via signed capability URLs
* **cloud-gateway:** serve offload job results via signed capability URLs
* **cloud:** apex-domain cutover + console deploy model
* **cloud:** control-plane audit-read + org-settings + residency-read endpoints (console Tier-2)
* **cloud:** export worker-lane OTLP traces to the collector (lane-only, #1740)
* **cloud:** gateway OTLP trace+metric export — clean end-to-end telemetry spine
* **cloud:** govern Modal deployments with typed specs
* **cloud:** per-org alert-threshold toggles (0006 alert_prefs) gating _fire_alerts
* **cloud:** per-org console service keys via /internal/console-keys
* **cloud:** per-org console service keys via POST /internal/console-keys
* **cloud:** relocate the schedule runner in-VPC (EventBridge -&gt; /internal/schedules/run-due)
* **cloud:** resolve fresh keys on-demand to kill the ~5-10s gateway 401 window
* **cloud:** surface the saved card's brand/last4/exp (Stripe payment-method retrieve)
* **cloud:** WP4 billing-depth control-plane endpoints (invoices, usage, alert toggles, payment method)
* **cloud:** WP4 billing-depth endpoints — invoices, usage series, alert-config toggles
* **control-plane:** make auto-recharge fire end-to-end
* **mac-smoke:** additive fake phase 0 — weightless surface contract
* **model-ops:** add stage-5 quality+perf onboarding gate
* **model-ops:** add-model trigger — labeled issue starts the onboarding agent
* **model-ops:** stage-5 quality+perf onboarding gate — measure vs target, bounded fix loop, evidence packet
* **models:** add Tencent R3-Skill support
* **models:** onboard intfloat/multilingual-e5-small (scaffolded)
* **quality-eval:** #1810 experiment runner — parallel Modal sessions, GPU selection, budget backstop
* **quality-eval:** autonomous experiment runner — parallel Modal sessions, GPU selection, budget backstop (Req14 v2 3/4)
* **quality-eval:** end-to-end pipeline wiring + dedup/budget semantics (Req14 v2 4/4)
* **quality-eval:** wire triage→planner→runner→autofix into a bounded pipeline
* **server:** add deterministic weightless fake adapter family
* **server:** failure injection for the fake adapter family
* **server:** fake adapter family — deterministic weightless sie-fake models
* **server:** synthetic memory-tracker mode — deterministic pressure/eviction
* **server:** synthetic memory-tracker mode for deterministic eviction
* support Iso ModernColBERT on Candle
* **tests:** lost-result 120s characterization on the full queue topology
* **tests:** lost-result characterization on the full queue topology

### Bug Fixes

* **bench:** accept canonical UTF-8 broker artifacts
* **bench:** accept omitted optional request refs
* **bench:** align broker contract boundary nesting
* **bench:** align broker log retention
* **bench:** align campaign runtime dependencies
* **bench:** avoid imagebuilder template collision
* **bench:** bind ami runtime kernel evidence
* **bench:** bound R3 quality regression gate
* **bench:** bound staging smoke below service knee
* **bench:** close campaign review and ci gaps
* **bench:** close campaign review edge cases
* **bench:** expose safe broker failure diagnostics
* **bench:** externalize broker deployment settings
* **bench:** fence broker rollback recovery
* **bench:** harden cloud campaign evidence
* **bench:** harden cloud campaign readiness
* **bench:** harden live campaign preflight
* **bench:** harden managed campaign control plane
* **bench:** make SSM dispatch at most once
* **bench:** pass authorized request to orchestrator
* **bench:** pin runnable worker manifests
* **bench:** prepare campaign clients before start
* **bench:** preserve broker terminal state
* **bench:** preserve cloud campaign timing evidence
* **bench:** preserve Fargate scratch permissions
* **bench:** probe exact start control key
* **bench:** publish worker image to owned ecr path
* **bench:** reject malformed canonical artifacts
* **bench:** relocate runtime caches to scratch
* **bench:** remove stale quality import
* **bench:** retain ami assertion diagnostics
* **bench:** retry transient ECS role propagation
* **bench:** retry transient R3 score timeouts
* **bench:** roll task revisions without downtime
* **bench:** secure ephemeral identity creation
* **bench:** secure staging cloud validation
* **bench:** skip absent profile deletion
* **bench:** support repeated maintenance pauses
* **bench:** trust scheduler group for worker expiry
* **bench:** validate R3 evaluation inputs
* **candle:** close ModernBERT audit gaps
* **ci:** disable implicit cloud tool installs
* **ci:** pin nested mise runtime
* **ci:** verify cloud IaC toolchain
* **cloud-gateway:** complete loopback class-fix via one canonical helper; align Python mirror + docs (#1906 final review)
* **cloud-gateway:** IP-literal loopback check + recover stuck payload-volume reload (CodeRabbit #1906 r2)
* **cloud-gateway:** reject non-ASCII signature hex without panicking
* **cloud-gateway:** validate capability-URL origins, bound TTL, and fix object-store minting (CodeRabbit #1906)
* **cloud:** address CodeRabbit — same-item batch id pairing + metrics protocol env + /metrics deny log
* **cloud:** address CodeRabbit — serialize env-mutating metrics-protocol test with a module-local ENV_LOCK
* **cloud:** address follow-up review
* **cloud:** address Modal deployment review
* **cloud:** alert on zero-metering invoice drift; reject ambiguous usage_events apportionment
* **cloud:** align dev backend benchmark
* **cloud:** align vendor billing with ledger debits
* **cloud:** batch/offload jobs retry retryable per-item MODEL_LOADING instead of failing the chunk
* **cloud:** bound on-miss key resolver's single-flight map + harden lookup
* **cloud:** CodeRabbit CP review — alert-configs PUT 404s, invoice-payload 502 guard, ty-ignore format
* **cloud:** console SIE_GATEWAY_URL is the public edge, not the raw Modal URL
* **cloud:** customer-scope the saved-card lookup — close cross-tenant card-PII leak
* **cloud:** deploy otel-collector + merge OTLP endpoint before lane snapshot so lane tracing activates
* **cloud:** derive outbox credits from the ledger debit and alert on reconciliation drift
* **cloud:** diagnose stale-checkout CP destroy + bridge VERCEL_API_TOKEN
* **cloud:** diagnose stale-checkout CP destroy + bridge VERCEL_API_TOKEN in cloud-provision
* **cloud:** don't persist raw billing_email in audit_log detail (PII)
* **cloud:** final WP4 review — strict invoice currency, atomic alert toggle, UTC-naive parse, card guard
* **cloud:** gateway OTLP export — blocking reqwest client for span batch processor + HTTP/TLS metric transport
* **cloud:** gateway OTLP export — Tokio runtime for span batch processor + HTTP/TLS metric exporter
* **cloud:** harden console-key lifecycle
* **cloud:** harden whole-estate finalization
* **cloud:** hide console keys from /me/usage + /me revoke (security review)
* **cloud:** preserve billing alert and meter attribution
* **cloud:** refresh recovery handle on batch chunk re-spawn (avoid orphaned container + 409 spam)
* **cloud:** release the scheduler lock_conn on any pre-hand-off failure
* **cloud:** stage the gen-lane model in model-stage and env-qualify the unstaged-weights hint
* **cloud:** stop deploy from dirtying satellite uv.locks
* **cloud:** stop deploy from dirtying satellite uv.locks (clean-tree guard)
* **cloud:** WP4 review follow-ups — usage upper bound, invoice-fallback diagnosability, card exp contract
* **control-plane:** hand auto-recharge sweep off the EventBridge request path
* **control-plane:** harden auto-recharge collection
* **control-plane:** make recharge retries lease-safe
* **fake-engine:** address two-axis review findings across the stack
* **fake-engine:** CI loopback interop + CodeRabbit review findings
* **gateway:** attest deployed execution revision
* **gateway:** include unanimous error_code in all_items_failed 500 logs
* **gateway:** reject malformed queue-request bodies with a typed 400
* harden classification quality verification
* **model-ops:** apply CodeRabbit review findings on the stage-5 gate
* **model-ops:** apply review findings on the friction PR
* **model-ops:** deterministic escalation when the agent itself crashes
* **model-ops:** filter vanilla floor material to fleet metric families
* **model-ops:** harden add-model trigger per review
* **model-ops:** raise directly in _read_json per code-quality finding
* **model-ops:** read vanilla numbers from the verdict's evidence field
* **model-ops:** rehearsal follow-ups — agent guardrail, record valves, auto-review, log-tail noise
* **model-ops:** rehearsal follow-ups — guardrails, record valves, log-tail noise
* **model-ops:** scaffolder crash on present-but-null sibling fields
* **model-ops:** seed-floor emits fleet metric families, not the raw mteb vector
* **model-ops:** stage-5 gate read vanilla numbers from verdict evidence, not probe measurements
* **model-ops:** tighten the agent-crash fallback per review
* **ocr:** address review lifecycle findings
* preserve parameterized quality gate targets
* **quality-eval:** accurate backstop messages for CPU playbooks (review S1)
* **quality-eval:** guard non-dict cluster signals in the pipeline (review)
* **quality-eval:** harden experiment runner per PR #1883 review (traversal, stale-verdict, rollback, build-clock)
* **quality-eval:** harden pipeline per PR #1891 review (report/verdict/workflow)
* **quality-eval:** reap process groups in rollback to avoid zombies (re-review)
* **quality-eval:** reject non-object verdict JSON in escalate_verdict (re-review)
* **quality:** address R3 review feedback
* **quality:** complete R3 faithful gate workflow
* resolve file-backed adapter impacts
* resolve onboarding bundle metadata
* resolve query-qualified quality targets
* reuse resolved uv in vanilla probe
* **server:** allow dependency-free model selections
* **server:** recommend bundle for full selection
* **server:** support cloud catalogs in bundle checks
* **server:** validate LightOnOCR bundle runtime

### Performance Improvements

* **candle:** cache exact bf16 gelu values
* **candle:** fuse ModernBERT exact GeGLU
* **candle:** optimize ModernBERT CUDA hot path
* **candle:** vectorize bf16 gelu lookup gate
* **cloud:** index usage_events_outbox (account_id, occurred_at) for the usage read (0007)
* **ocr:** serve generative models with SGLang
* **ocr:** serve generative OCR with SGLang continuous batching

## v0.6.19 (2026-07-14)

### Highlights

- **New capabilities:** add catalog-smoke step — per-model serve_i6pn load verification; add env-specific Modal deploy provisioning; align prod-us manifest to latest i6pn fast-path topology; commit generated price book so the provisioner locates it; deploy-both IaC hardening — env-derived Modal env + env-targeted B6 smoke; event-driven fast-lane promotion + credit ramp
- **Reliability and operations:** drop the speculative mise install retry chain; harden deploy tooling — modal --json ANSI, async snapshot-commit race, Vercel sensitive-var upsert; harden lane_deploy driver — modal --json ANSI + async snapshot-commit race; retry server launch on startup death; scope the port sweep to our serve chain and harden the start guard

### Features

* **cloud:** add catalog-smoke step — per-model serve_i6pn load verification
* **cloud:** add env-specific Modal deploy provisioning
* **cloud:** align prod-us manifest to latest i6pn fast-path topology
* **cloud:** commit generated price book so the provisioner locates it
* **cloud:** deploy-both IaC hardening — env-derived Modal env + env-targeted B6 smoke
* **cloud:** event-driven fast-lane promotion + credit ramp
* **cloud:** normalize the OSS terminal `failed` readiness state to MODEL_LOAD_FAILED on the Modal lane
* **cloud:** productionize the i6pn realtime dispatch fast path end-to-end
* **cloud:** readiness-gated i6pn dispatch — thread loaded_models to the wire, gate cold workers
* **model-ops:** deterministic add-a-model scaffolder — verdict → YAML + matrix + target
* **model-ops:** deterministic add-a-model scaffolder + onboarding runbook
* **model-ops:** structured feasibility verdict — eval-model stage-2 gate
* **model-ops:** structured feasibility verdict for add-a-model stage 2
* **quality-eval:** add fix-proposal planner (tiered Fable-5) + experiment-plan schema
* **quality-eval:** bounded auto-fix — agent PRs for floor re-baselines, human merge gate (Req 14 stage 4)
* **quality-eval:** bounded auto-fix CLI for floor re-baseline
* **quality-eval:** fix-proposal planner (tiered Fable-5) + experiment-plan schema (Req14 v2 2/4)
* **quality-eval:** make the planner reason from signals + consume the LLM's predicted evidence
* **quality-eval:** onboarding smoke playbook — served-route bring-up on Modal
* **quality-eval:** provenance-gated preflight + vanilla-lane floor template
* **quality-eval:** provenance-gated preflight + vanilla-lane floor template (Req14 v2 1/4)

### Bug Fixes

* **ci:** drop the speculative mise install retry chain
* **ci:** give the scheduled quality run its own concurrency group
* **ci:** stop rolling the un-retried rustup ensure in mac-smoke-nightly
* **cloud:** address catalog-smoke review — gate internal token, harden revoke + warm-up bound
* **cloud:** address CodeRabbit round-2 review on #1782
* **cloud:** address env deploy review comments
* **cloud:** address PR review comments
* **cloud:** bound wedged model loads + typed dead-letter on the lazy lane
* **cloud:** bound/adapt slab credit holds so a funded org stops self-wedging at 402
* **cloud:** catalog-smoke presents SIE_API_KEY when the gateway enforces managed keys
* **cloud:** don't send `type` on Vercel env update (sensitive-var 400)
* **cloud:** drop candle-only bge-large-en from the default L4 lane's models_served
* **cloud:** gateway must not 401 valid keys during cold-replica key-snapshot warmup
* **cloud:** gen-lane warm floor + gateway-config secret merge in deploy-edge
* **cloud:** harden deploy tooling — modal --json ANSI, async snapshot-commit race, Vercel sensitive-var upsert
* **cloud:** harden lane_deploy driver — modal --json ANSI + async snapshot-commit race
* **cloud:** heartbeat TTL for i6pn discovery keys (corpse expiry on hard stop)
* **cloud:** narrow B6 local gateway bundle to the deployed models_served (#1771 hash parity)
* **cloud:** pg-connector smoke builds ModalLaneExecutor from SIE_LANE_APP_MAP
* **cloud:** pg-connector smoke builds ModalLaneExecutor from SIE_LANE_APP_MAP (#1782 missed caller)
* **cloud:** scope _modal_cli TERM=dumb to captured calls only (CodeRabbit #1807)
* **cloud:** skip B6 offload-batch in external-CP mode (unshareable payload store)
* **cloud:** stage refs/main alias so offline lanes resolve pinned-SHA models
* **cloud:** stop provisioning unbacked Stripe metered prices
* **cloud:** stop provisioning unbacked Stripe metered prices (prepaid-credits + Metronome rating is the model)
* **cloud:** tighten edge deploy conflict resolution
* **cloud:** Vercel env update sends value only (sensitive-var target 400)
* **cloud:** warm managed key before catalog-smoke sweep to beat the key-propagation race
* **cloud:** wire the default gen lane into the gateway census + app map
* **gateway:** bump yanked spin 0.10.0 to 0.10.1 in Cargo.lock
* mac-smoke nightly score-path contract + unknown-model 500s
* **mac-smoke:** call /v1/score with the slash-form model id
* **mac-smoke:** kill stale server by port between phases
* **mac-smoke:** retry server launch on startup death
* **mac-smoke:** scope the port sweep to our serve chain and harden the start guard
* **model-ops:** harden feasibility validator per review
* **model-ops:** harden scaffolder per review
* **poc:** pass explicit github_token to claude-code-action
* **quality-eval:** charset-gate vanilla_vector metric keys (review C1)
* **quality-eval:** classify malformed encode responses as REFUTED, not a crash
* **quality-eval:** controlled exit on planner output-write failure (CodeRabbit)
* **quality-eval:** harden autofix per CodeRabbit review round 2
* **quality-eval:** harden autofix preflight against output injection (review round 1)
* **quality-eval:** require string cell fields before charset validation
* **server:** colbert-small multivector encode returns one vector-set per input, worker-item order
* **server:** Florence-2 processor compat with transformers 4.57
* **server:** guard cross_encoder predict tokenization against the concurrent-metering race (class-fix for #1800)
* **server:** load tokenizer by pinned revision so offline chat-template rendering doesn't need refs/main
* **server:** return 404 instead of 500 for unknown models on inference endpoints
* **server:** terminal Failed readiness + Florence-2 processor compat + pin launch-set revisions
* **server:** terminal Failed readiness state so the sidecar dead-letters permanently-failed model loads

## v0.6.18 (2026-07-12)

### Highlights

- **New capabilities:** add flat→per-account Files-store migration; co-located AWS us-east-1 reverse-proxy edge for api.superlinked.com; consolidate data-plane deploy + manifest-drive proxy-auth/i6pn; namespace Files-store objects by account; opt-in Modal proxy-auth layer for config/lane/OTLP ingress; parameterize Modal environment + prod-us tf root by SIE env (§B2)
- **Reliability and operations:** fail fast on terraform plan errors before the destroy-check; bake SIE_MODAL_PROXY_AUTH into image env so the proxy-auth knob doesn't diverge local↔remote; make proxy_auth_env authoritative in both directions; pin npm for TypeScript release publishing; aws-edge S3 backend uses native state locking (use_lockfile)

### Features

* **cloud:** add flat→per-account Files-store migration
* **cloud:** co-located AWS us-east-1 reverse-proxy edge for api.superlinked.com
* **cloud:** consolidate data-plane deploy + manifest-drive proxy-auth/i6pn
* **cloud:** namespace Files-store objects by account
* **cloud:** opt-in Modal proxy-auth layer for config/lane/OTLP ingress
* **cloud:** parameterize Modal environment + prod-us tf root by SIE env (§B2)
* **cloud:** per-account Files-store namespacing + migration
* **cloud:** prod bring-up terraform prerequisites
* **cloud:** prod bring-up terraform prerequisites — prod-us auth-token outputs, aws-edge prod state
* **cloud:** prod enablement — aws-edge provisioner step, prod-us root, Modal-env parameterization
* **cloud:** prod-us terraform root — env-isolated control plane, separate state bucket
* **cloud:** reconcile OTLP transport + proxy-auth sender headers
* **cloud:** stand up the AWS edge via cloud-provision
* **dispatcher:** full IP-pin for object-store + postgres verify-full egress
* **dispatcher:** org-ownership assertion on upload:// ends — defense in depth behind the gateway check
* **dispatcher:** resolve-then-pin connector egress — SSRF TOCTOU backstop

### Bug Fixes

* **ci:** fail fast on terraform plan errors before the destroy-check
* **ci:** pin npm for TypeScript release publishing
* **cloud:** address CodeRabbit review on the prod-enablement PR
* **cloud:** aws-edge S3 backend uses native state locking (use_lockfile)
* **cloud:** bake SIE_MODAL_PROXY_AUTH into image env so the proxy-auth knob doesn't diverge local↔remote
* **cloud:** collision-aware Files migration move
* **cloud:** correct console callback route + guard the CP terraform apply against silent destroys
* **cloud:** correct scheduler/slo-tuner DSN-secret comments + enforce non-ingress secret refs in deploy-posture audit
* **cloud:** default Metronome credit type to fiat USD-cents so low-balance alerts and grants provision
* **cloud:** gateway payload store must be a local path, not the s3 payload_store_url
* **cloud:** gateway payload store must be a local path, not the s3 payload_store_url (#1743 follow-through)
* **cloud:** gateway ships the cloud-storage backend; provisioner wires SIE_CONFIG_SERVICE_URL
* **cloud:** guard config-URL resolver against a malformed manifest
* **cloud:** harden Phase-3 CI + sibling-sweep per adversarial review (Refs #1757 #1758)
* **cloud:** make proxy_auth_env authoritative in both directions
* **cloud:** parameterize the deploy-env slug in managed Modal app names
* **cloud:** parameterize the OTLP collector app name with the deploy-env slug
* **cloud:** provision §7.6 prepaid credit-pack price + STRIPE_PRICE_ID
* **cloud:** redact all secret-bearing Config fields in Debug
* **cloud:** redact ModalProxyToken Debug + cover /gen handshake forwarding
* **cloud:** resolve gateway app name from gateway_app.APP_NAME, not ctx.env
* **dispatcher:** address CodeRabbit review on the #1742 egress pins
* **dispatcher:** harden connector egress pin against CodeRabbit findings
* stabilize tilt e2e regressions
* use pinned cargo-zigbuild in tilt builds

## v0.6.17 (2026-07-09)

### Highlights

- **New capabilities:** add GTE RoPE embedding kernel; CLI WorkOS device-flow auth — retire the shared-token shortcut for customer commands; SIE Cloud managed service — AWS control plane, Modal data plane, provisioning + billing kernel; OSS amendments from the SIE Cloud build — engine, SDKs, gateway/sidecar seams; add deterministic regression-ledger detector for scheduled quality runs; add floor-archaeology playbook (CPU-only)
- **Reliability and operations:** retry model-dependent calls on MODEL_LOADING; harden regression-ledger live mode and audit input validation; harden multi-gpu worker routing; compile fused gated activation op; import CUDA pointer trait

### Features

* **candle:** add GTE RoPE embedding kernel
* **cloud:** CLI WorkOS device-flow auth — retire the shared-token shortcut for customer commands
* **cloud:** SIE Cloud managed service — AWS control plane, Modal data plane, provisioning + billing kernel
* **oss:** OSS amendments from the SIE Cloud build — engine, SDKs, gateway/sidecar seams
* **quality-eval:** add deterministic regression-ledger detector for scheduled quality runs
* **quality-eval:** add floor-archaeology playbook (CPU-only)
* **quality-eval:** add Modal parity-probe playbook (diagnostic-only)
* **quality-eval:** add official-recipe, fp32-ab, config-sweep, determinism playbooks
* **quality-eval:** add report-only triage classifier + workflow for scheduled quality runs
* **quality-eval:** add validation-playbook harness + hypothesis/verdict contract
* **quality-eval:** report-only triage classifier + gated workflow for scheduled quality runs
* **quality-eval:** seed triage signature DB and short-circuit known clusters
* **quality-eval:** spread runner fleet across capacity pools with always-on floors
* **quality-eval:** validation playbooks — adversarial repro harness (Req 14 Proj 1, stage 3)
* **server:** support multi-gpu worker placement
* **worker:** route by child queue pressure
* **worker:** route multi-gpu sidecar children
* **worker:** support multi-gpu sidecar children

### Bug Fixes

* **candle:** compile fused gated activation op
* **candle:** import CUDA pointer trait
* **mac-smoke:** retry model-dependent calls on MODEL_LOADING
* **quality-eval:** address CodeRabbit + code-quality review comments
* **quality-eval:** address CodeRabbit review findings
* **quality-eval:** address triage review findings
* **quality-eval:** detect EBS cache volume via lsblk model string
* **quality-eval:** enforce budget cap on all eval playbooks + budget-stop artifacts
* **quality-eval:** fp32 ST arm via throwaway -st model, not a profile
* **quality-eval:** harden playbook verdict logic (review round 1)
* **quality-eval:** harden regression-ledger live mode and audit input validation
* **quality-eval:** harden triage input handling per CodeRabbit review
* **quality-eval:** keep register-runners.sh bash-3.2 compatible
* **quality-eval:** tolerate stop/start race in wake instance-stopped wait
* **quality-triage:** share data-bearing baseline discovery with triage live mode
* **quality-triage:** skip data-less scheduled runs in ledger baseline auto-discovery
* **req6:** align candle chart after rebase
* **req6:** harden multi-gpu worker review fixes
* **req6:** harden multi-gpu worker routing
* **req6:** resolve main rebase fallout
* **req6:** wire pressure-aware multi-gpu routing
* **server:** keep ipc ping alive on gpu health errors
* **sidecar:** bound queue admission through scheduler completion
* **sidecar:** refresh ready slots on cancel fanout failure
* **sidecar:** release cancelled scheduler pressure
* **sidecar:** wire local ingest through adapter pool

## v0.6.16 (2026-07-07)

### Highlights

- **New capabilities:** opt-in template_as_prompt for ST prompt-length-aware templates; use_model_encode — delegate to checkpoint-native encode(); add Candle OOM pressure recovery; add Candle residency eviction; per-pool gpu_driver option + confidential H100 tracing example
- **Reliability and operations:** harden splade flash gate + add flash-vs-native parity test; splade-v3 floors were cosine-epoch artifacts — re-baseline to dot self-measures + harden flash gate; restore executable bit on mac-smoke mise task; faithful ModernBERT forward + hardened 1_Dense loading; restore e5 query/passage templates on the multilingual-e5-large sentence_transformer profile

### Features

* **adapters:** opt-in template_as_prompt for ST prompt-length-aware templates
* **adapters:** use_model_encode — delegate to checkpoint-native encode()
* add Candle OOM pressure recovery
* add Candle residency eviction
* **terraform/azure:** per-pool gpu_driver option + confidential H100 tracing example

### Bug Fixes

* **adapters:** harden splade flash gate + add flash-vs-native parity test
* **adapters:** honor caller instruction in use_model_encode without a template
* address candle catalog review comments
* address Candle residency review comments
* align Candle idle eviction lifecycle
* align candle workers with catalog lane config
* **bench:** cap docling pubtabnet-html at 2000 pages to fit the job budget
* **bench:** splade-v3 floors were cosine-epoch artifacts — re-baseline to dot self-measures + harden flash gate
* **bundles:** pin flash-attn-4 prerelease so sglang 0.5.10.post1 resolves
* **ci:** restore executable bit on mac-smoke mise task
* **colbert_modernbert_flash:** apply trained ColBERT Dense projection (+ guard, tokenizer, re-baseline)
* **colbert_modernbert_flash:** faithful ModernBERT forward + hardened 1_Dense loading
* **colbert:** GTE muvera also at simhash 6 - CMedQA/MMarco evals OOM at dim 81920
* **colbert:** mxbai muvera at simhash 6 (FDE dim 20480) - runner dies at 81920
* **colbert:** retune GTE/mxbai muvera for the Dense-projected representation
* **colbert:** serve pylate Dense chains and align fallback with flash
* **deploy:** cap tester l4 lanes
* **deploy:** zero and cap tester L4 lanes
* **deploy:** zero AWS worker warm floor
* **deploy:** zero tester worker warm lanes
* **gateway:** fan out gpu-agnostic demand across pool profiles for scale-from-zero
* **gateway:** normalize explicit GPU demand labels
* **helm:** wake a deterministic owner lane on gpu-agnostic demand for multi-profile bundles
* **models:** NV-Embed-v2 — serve via checkpoint-native encode() (latent-attention recipe), NFCorpus 0.3715→0.4479
* **models:** restore e5 query/passage templates on the multilingual-e5-large sentence_transformer profile
* **models:** restore e5 templates on multilingual-e5-large ST profile
* **models:** serve NV-Embed-v2 via its native latent-attention ST path
* preserve Candle idle evictor failure logs
* **quality-eval:** fail fast when sie-server dies during readiness wait
* **quality-eval:** unbreak 7B sglang lanes (flash-attn-4 resolution) + fail-fast readiness + docling job budget
* restore kind smoke candle routing

## v0.6.15 (2026-07-03)

### Highlights

- **New capabilities:** add Candle profile routing and catalog models; add candle runtime diagnostics; add candle xlm-roberta flash attention; add dev AMI bake script; add dev ec2 cleanup workflow; add dev EC2 launch script
- **Reliability and operations:** harden native worker routing; harden native worker runtime; restore sidecar model readiness; delete Python NATS pull loop from worker hot path; harden SIE dev AMI workflow
- **Performance:** add GEMM diagnostics switch; fuse xlm-r qkv projection; optimize bge-m3 xlm-r runtime path; enable candle reduced precision gemms

### Features

* add Candle profile routing and catalog models
* add candle runtime diagnostics
* add candle xlm-roberta flash attention
* add dev AMI bake script
* add dev ec2 cleanup workflow
* add dev EC2 launch script
* add SIE API key rotation tooling
* **candle:** add native adapter runtime
* **dev:** add personal SIE CUDA/Rust AMI workflow
* document ssm ssh dev launcher
* **gateway:** count dropped stale NAKs from abandoned attempts
* **helm:** add opt-in OTLP tracing wiring + bundled collector to sie-cluster
* **helm:** add optional bundled Tempo tracing backend
* **helm:** add SIE tracing dashboard and tempo datasource uid
* **helm:** opt-in OTLP tracing wiring + bundled collector in sie-cluster
* **helm:** optional bundled Tempo tracing backend + Grafana datasource for sie-cluster
* improve candle replica concurrency
* improve dev AMI cache scalability
* **observability:** bound tracing shutdown + trim env inputs across all four runtimes
* **observability:** bound tracing shutdown and trim env inputs across all four runtimes
* **observability:** gate Rust tracing on SIE_TRACING_ENABLED
* **observability:** unify tracing enable switch across gateway, sidecar, and worker
* **quality-eval:** gate score-on-retrieval candidate cells once floor committed
* reattach dev cache volume
* **server:** land Apple Silicon as a primary device onto main
* **server:** land Apple Silicon as a primary device onto main (re-land of #1455)
* **sie_gateway:** emit gateway.publish span on non-streaming queue-publish path
* **sie_server_rust:** add OTLP distributed tracing
* **sie_server_rust:** add OTLP distributed tracing to the Rust native worker
* **terraform:** dedicated single-AZ node group for stateful observability
* **tilt:** enable tracing with bundled Tempo
* **tracing:** emit sidecar.dispatch span from sie_server_sidecar
* **tracing:** inject queue trace context for non-generation endpoints
* **tracing:** instrument sie_server_sidecar with OTLP (sidecar.dispatch span)
* **tracing:** propagate trace context through the non-streaming worker loop
* tune candle forward concurrency
* tune dev EC2 launch defaults
* wire dev AMI cache disk helpers

### Bug Fixes

* **#1340:** stella per-task instruction + re-baseline stale dense/colbert floors (blast radius was #1432)
* **adapters:** guard empty-document MaxSim batches; strengthen single-doc test
* address API key review feedback
* address API key rollback review
* address rust worker smoke review comments
* address sidecar PR review feedback
* align rust worker kind smoke wiring
* align rust worker warm smoke behavior
* align worker pool config hashes
* back up API key secret before apply
* **candle:** align cublaslt with candle cuda stream
* **candle:** align readiness with loaded model state
* **candle:** copy vendored cuda dependency in image
* **candle:** harden native worker routing
* **candle:** harden native worker runtime
* **candle:** import cublaslt cuda extension
* **candle:** keep catalog additions candle-only
* **candle:** publish loaded models in health
* **candle:** repair layer norm cuda lifetimes
* **candle:** restore sidecar model readiness
* **candle:** run bge-m3 profile in bf16
* **candle:** vendor cublaslt for cuda build
* **candle:** vendor layer norm cuda dependency
* clean stale Python queue ownership drift
* complete empty preformed worker requests
* compress dev ami user data
* **config:** defer cancellation after config commit starts
* **config:** handle idempotency in-flight races
* **config:** shield idempotency failure cleanup
* copy rust worker vendor deps before docker cook
* **dashboard:** harden run_status backfill per review feedback
* **dashboard:** keep cancelled nightlies off the Quality on main widget
* delete Python NATS pull loop from worker hot path
* **dev:** address dev AMI review hardening
* **dev:** harden SIE dev AMI workflow
* **dev:** persist Codex state on dev cache disk
* **docs:** drop otel.name from span-attribute table; tighten tracing dashboard test
* export profile bundle targets
* **gateway:** append forwarded embeddings headers to preserve multi-value tracestate
* **gateway:** cancel abandoned direct batch fallback work
* **gateway:** don't report a succeeded SSE generation as inter_chunk_timeout on consumer lag
* **gateway:** drop stale NAKs from abandoned attempts in handle_nak
* **gateway:** forward traceparent/tracestate through /v1/embeddings rewrite
* **gateway:** harden batch direct fallback cancellation
* **gateway:** ignore pre-latch stale NAKs after republish
* **gateway:** ignore quick-xml &lt;0.41 RustSec advisories in cargo-deny
* **gateway:** keep signalling SSE incompleteness on consumer lag; skip redundant cancel
* **gateway:** keep signalling SSE incompleteness on lag; only skip cancel when done
* **gateway:** record demand on backpressure so KEDA scales the lane
* **gateway:** route engine-pin errors through endpoint_error_response
* **gateway:** separate queue timeouts from model loading
* **gateway:** use real RUSTSEC id for second quick-xml advisory
* harden candle bge-m3 runtime bring-up
* harden dev EC2 launcher dry-run
* harden dev profile home fallback
* harden SIE API key tooling
* **helm:** make bundled collector downstream OTLP TLS configurable
* **helm:** set OTEL_EXPORTER_OTLP_PROTOCOL=grpc on traced pods
* **helm:** suppress bundled collector when explicit tracing endpoint set
* improve candle embedding runtime throughput
* install gateway pool config hashes
* install gpg agent for dev AMI bake
* keep dev AMI user data under limit
* log candle batch failure context
* **models:** load stock tokenizer for gte-Qwen2-7B-instruct
* **models:** restore per-task MTEB instruction for stella_en_v5
* **models:** serve embeddinggemma-300m via SentenceTransformerDenseAdapter
* **models:** serve gte-Qwen2-1.5B via faithful sentence_transformer adapter
* **models:** serve gte-Qwen2-1.5B via sentence_transformer adapter
* **muvera:** center tokens before SimHash to fix answerai [@muvera](https://github.com/muvera) Blackwell collapse
* **muvera:** drop same-dim count-sketch + retune colbert muvera configs
* **muvera:** drop same-dim count-sketch, retune colbert configs + floors
* persist dev cache volume
* persist dev kubeconfig
* **playground:** keyboard-accessible tabs + associated form labels
* polish rollback backup examples
* preserve candle compute cap diagnostics
* preserve control-plane bundle hash in sidecar
* quiet dev AMI tool installs
* remove candle replica fallback path
* remove Python worker pull-loop path
* repair candle ci routing checks
* repair candle rebase fallout
* repeat dev AMI bake markers
* require explicit API key rollback path
* resolve latest dev AMI base
* respect hf cache env in rust worker
* restore default candle cuda stream
* route kind rust worker build through docker task
* **sdk:** raise actionable error on encode result-count desync
* select rustls provider at binary startup
* select rustls provider for rust services
* **server:** /v1/embeddings maps OOM to 503 RESOURCE_EXHAUSTED, not 500
* **server:** address CI + review comments on the Apple Silicon re-land
* **server:** apply profile runtime options + append EOS on cluster encode path
* **server:** apply profile runtime options and append EOS on cluster encode path
* **server:** bound continuous-batch drain so a busy LoRA can't starve siblings
* **server:** bound the continuous-batching drain so a busy LoRA can't starve siblings
* **server:** cap LoRA drain to snapshot backlog so post-snapshot arrivals don't overshoot
* **server:** make SGLang load eviction headroom-aware
* **server:** move load headroom behind adapter contract
* **server:** serialize hot reload registry mutations
* **server:** skip EOS on empty input; cover non-default profile on queue path
* **server:** stop mislabelling uint8 embeddings as "binary" on the queue path
* **server:** stop mislabelling uint8 embeddings as binary on the queue path
* **server:** unregister from MemoryManager even if adapter.unload() raises
* **server:** unregister model from MemoryManager even if adapter.unload() raises
* set home during dev AMI bake
* **sidecar:** bound scheduler bulk enqueue waves
* **sidecar:** keep only config hash wiring
* **sidecar:** start batch cancel before direct pulls
* **sie_server_rust:** only log OTLP init success when exporter built
* stabilize rust worker kind smoke
* **terraform:** address CodeRabbit review on obs node group
* **terraform:** auto-derive observability node group AMI from instance arch
* **terraform:** make observability node group ami_type/disk_type configurable
* tighten rust worker smoke diagnostics
* **tilt:** use single string for candle-helm-smoke cmd (Starlark has no implicit concat)
* **tilt:** wait for complete trace before asserting span chain
* tolerate missing zstd devel package
* **tracing:** assert items/metadata alignment before zip in scheduler batch
* **tracing:** attach inbound trace context for non-generation publishes
* **tracing:** parent sidecar.dispatch on first valid inbound context
* verify dev AMI bake via ssm

### Performance Improvements

* **candle:** add GEMM diagnostics switch
* **candle:** fuse xlm-r qkv projection
* **candle:** optimize bge-m3 xlm-r runtime path
* enable candle reduced precision gemms
* fuse candle xlm-roberta qkv projection
* **gateway:** build the HRW ring from the by_bundle index, not a whole-cluster scan
* improve candle worker concurrency and build cache
* restore candle xlm-roberta qkv layout
* **server:** vectorize bge_m3_flash CLS gather to drop N per-encode GPU syncs
* **sidecar:** bulk enqueue scheduler batches

### Reverts

* restore gateway sidecar rustls defaults

## v0.6.14 (2026-06-26)

### Highlights

- **Reliability and operations:** unblock ColBERT score-on-retrieval — score_pairs 500 + muvera dense-encode + gate co-serve; include sie-bench in image context

### Bug Fixes

* **#1430:** unblock ColBERT score-on-retrieval — score_pairs 500 + muvera dense-encode + gate co-serve
* **docker:** include sie-bench in image context

## v0.6.13 (2026-06-25)

### Highlights

- **New capabilities:** publish sie-bench image to GHCR on release; per-tool GPU routing and env-overridable GLiNER models
- **Reliability and operations:** extend tester structured generation timeouts; retry failed payload cleanup deletes; align sie-test qwen27 lane routing; build sie-bench image CPU-only with headless opencv; tolerate cold-start provisioning errors in scale-from-zero loadtests

### Features

* **ci:** publish sie-bench image to GHCR on release
* **sie-mcp:** per-tool GPU routing and env-overridable GLiNER models

### Bug Fixes

* align sie-test qwen27 lane routing
* **bench:** address CodeRabbit review (non-root image, clearer comment, checkout creds)
* **bench:** build sie-bench image CPU-only with headless opencv
* **bench:** tolerate cold-start provisioning errors in scale-from-zero loadtests
* extend tester structured generation timeouts
* **gateway:** dedupe offloaded payload cleanup keys
* **gateway:** retry failed payload cleanup deletes
* **gateway:** track exact offloaded payload keys
* **loadtest:** tolerate cold-start provisioning errors + publish sie-bench to GHCR
* persist sie-test qwen warm lanes
* scale sie-test spot lanes to zero
* **server:** drain removed profile variants during config apply
* **server:** preflight profile variant config updates
* **server:** preserve profile variants during config reconcile
* **server:** serialize async config registry updates
* **server:** tighten profile variant reconciliation
* **sie-mcp:** chunk GLiNER extract/redact for large docs; flag the other 4B

## v0.6.12 (2026-06-24)

### Highlights

- **New capabilities:** enable gateway-API model pinning on sie-test; add per-profile quality floors for bge-m3 sparse/multivector
- **Reliability and operations:** exclude per-profile target.&lt;profile&gt;.json floors from result loaders; run docker task in project env; skip payload-store cleanup DELETEs for inline requests; guard label-window overflow + defer LEDGAR for v1.0 uni-encoders; gate subset-parameterized tasks; keep skipping candidate-gen dirs
- **Performance:** release MaxSim per-batch intermediates to cap peak memory

### Features

* **deploy:** enable gateway-API model pinning on sie-test
* **quality-eval:** add per-profile quality floors for bge-m3 sparse/multivector

### Bug Fixes

* **bench:** exclude per-profile target.&lt;profile&gt;.json floors from result loaders
* **ci:** run docker task in project env
* **deploy:** address CodeRabbit review on sie-test pinning
* **gateway:** record offload before the store write (PR review)
* **gateway:** skip payload-store cleanup DELETEs for inline requests
* **gliclass:** guard label-window overflow + defer LEDGAR for v1.0 uni-encoders
* **quality-eval:** gate subset-parameterized tasks; keep skipping candidate-gen dirs
* **sie_bench:** batch-local query padding in MaxSim to fix #1461 OOM

### Performance Improvements

* **sie_bench:** release MaxSim per-batch intermediates to cap peak memory

## v0.6.11 (2026-06-23)

### Highlights

- **New capabilities:** floor Florence-2-base on COCO detection (AP 0.2785); floor Florence-2-base on olmOCR-bench (accuracy 0.0824); fail fast when a bundle is placed on the wrong platform; add Gemma 4 generation models (E2B, E4B, 26B-A4B); first-class cuda13 platform for the gemma bundle; first-class cuda13 platform for the gemma bundle (build/CI/release/deploy)
- **Reliability and operations:** raise uv HTTP timeout for the gemma cu130 wheel downloads; harden logical pool queue routing; harden worker-direct batch fallback; return retryable model loading for generation; gate mteb task instruction to instruction-following models

### Features

* **bench:** floor Florence-2-base on COCO detection (AP 0.2785)
* **bench:** floor Florence-2-base on olmOCR-bench (accuracy 0.0824)
* **deploy:** fail fast when a bundle is placed on the wrong platform
* **server:** add Gemma 4 generation models (E2B, E4B, 26B-A4B)
* **server:** first-class cuda13 platform for the gemma bundle
* **server:** first-class cuda13 platform for the gemma bundle (build/CI/release/deploy)
* **server:** keep pinned models loaded and exclude them from LRU eviction
* support logical pools backed by queue pools
* **terraform:** add marci-dev Azure A10 single-node smoke-test example
* **terraform:** Azure marci-dev A10 smoke-test example + node-subnet NSG LB fix

### Bug Fixes

* align retryable generation review feedback
* **bench:** gate mteb task instruction to instruction-following models
* **bench:** gate MTEB task instructions to instruction-following models
* **bench:** resolve profile output_types from adapter_options.runtime for bge-m3 sparse/multivector evals
* **ci:** raise uv HTTP timeout for the gemma cu130 wheel downloads
* **ci:** run the gemma cuda13 build smoke + cover it in release verify/stamp/warm
* disable payload store in kind smoke
* handle sidecar progress ack failures
* harden logical pool queue routing
* harden worker-direct batch fallback
* **loadtest:** disable payload store in nightly deploy
* merge runtime pool capacity changes
* preserve queue delivery budget
* return retryable model loading for generation
* **server:** correct stale cuda13/gemma platform references
* **server:** make docker bake/verify platform-aware (cuda13/gemma in mixed builds)
* **server:** make transformers version bundle-controlled
* **terraform:** allow public LoadBalancer/ingress inbound on AKS node subnet NSG
* **terraform:** clear internal mise references from public Azure module
* **terraform:** clear internal mise/tooling references from public AWS and GCP modules
* **terraform:** drop internal references from public Azure module README and variables

## v0.6.10 (2026-06-22)

### Highlights

- **New capabilities:** floor measure-first dark models, retire donut-rvlcdip; provision and enable the payload store by default; fix &lt;OD&gt; bbox conversion, re-task base-ft/large to COCO detection
- **Reliability and operations:** align gliner label inputs; align gliner v2.5 conll protocol; improve generation failure observability; coalesce whitespace-only instruction to default; pool post-RMSNorm last_hidden_state

### Features

* **bench:** floor measure-first dark models, retire donut-rvlcdip
* **deploy:** provision and enable the payload store by default
* **florence2:** fix &lt;OD&gt; bbox conversion, re-task base-ft/large to COCO detection

### Bug Fixes

* **bench:** align gliner label inputs
* **bench:** align gliner v2.5 conll protocol
* **deploy:** address payload-store review feedback
* improve generation failure observability
* **qwen3_vl_embedding:** coalesce whitespace-only instruction to default
* **qwen3_vl_embedding:** pool post-RMSNorm last_hidden_state
* **sie_bench:** coerce ClassLabel int names to str in classification loader
* **sie_bench:** repoint extract-kie FUNSD to live nielsr/funsd dataset
* **sie-cluster:** make the self-host deploy skill robust to its own runtime
* **sie-server:** make transformers5 bundle reachable via pip
* sync generation visibility api contract

## v0.6.9 (2026-06-19)

### Highlights

- **New capabilities:** add per-pool pinned-model set to the pool config API; per-pool pinned-model set in the pool config API; support profile-qualified ids in per-pool pinned-model set; enable S3 payload store for &gt;1MB work items
- **Reliability and operations:** bound pinned-model metric labels and add handler validation tests; scope worker preload models by lane; alias PEFT LoRA adapter names; expose model_cache_bucket_url as a root output

### Features

* **gateway:** add per-pool pinned-model set to the pool config API
* **gateway:** per-pool pinned-model set in the pool config API
* **gateway:** support profile-qualified ids in per-pool pinned-model set
* **tester-cluster:** enable S3 payload store for &gt;1MB work items

### Bug Fixes

* **gateway:** bound pinned-model metric labels and add handler validation tests
* scope worker preload models by lane
* **server:** alias PEFT LoRA adapter names
* **tester-cluster:** expose model_cache_bucket_url as a root output

## v0.6.8 (2026-06-16)

### Highlights

- **Reliability and operations:** harden release artifact workflows; keep score traffic from idle-evicting rerankers; serialize worker use with unload state; share score media batch cost

### Bug Fixes

* **ci:** harden release artifact workflows
* **server:** keep score traffic from idle-evicting rerankers
* **server:** serialize worker use with unload state
* **server:** share score media batch cost

## v0.6.7 (2026-06-16)

### Highlights

- **New capabilities:** answer_questions transient-QA tool (Req 12 #1309)
- **Reliability and operations:** guard model ready timeout against liveness budget; wire model ready timeout through helm; add Qwen3.6 27B 32K RTX serving profile; route grammar requests to a non-speculative profile (NEXTN bypasses Outlines FSM); make worker config reconciliation no-op safe
- **Performance:** extract shared vision patch-embed rebind helper; batch generate() across pages to close #601 L4 throughput gap

### Features

* **mcp:** answer_questions transient-QA tool (Req 12 #1309)

### Bug Fixes

* add Qwen3.6 27B 32K RTX serving profile
* address qwen 27b review feedback
* **generate:** route grammar requests to a non-speculative profile (NEXTN bypasses Outlines FSM)
* guard model ready timeout against liveness budget
* make worker config reconciliation no-op safe
* route qwen 27b fp8 cold starts
* wire model ready timeout through helm

### Performance Improvements

* **adapters:** extract shared vision patch-embed rebind helper
* **lighton_ocr:** batch generate() across pages to close #601 L4 throughput gap

## v0.6.6 (2026-06-14)

### Highlights

- **Reliability and operations:** align pool-scoped bundle hashes; avoid sticky missing bundle hashes; clarify missing profile inheritance; fail closed on missing bundle metadata; stabilize keda all-marker e2e

### Bug Fixes

* **config:** align pool-scoped bundle hashes
* **config:** avoid sticky missing bundle hashes
* **config:** clarify missing profile inheritance
* **config:** fail closed on missing bundle metadata
* **tilt:** stabilize keda all-marker e2e

## v0.6.5 (2026-06-13)

### Highlights

- **New capabilities:** demand-side token-reduction benchmark (Req 12, #1311); add describe_image tool (caption + zero-shot tags); add describe_image tool (caption + zero-shot tags) — Req 12 #1310; cap describe_image payload size before cluster calls; claude.ai connector surface — OAuth bridge + skill ZIP (Req 12 #1312); sie_mcp edge with docs_to_markdown tool (Req 12 #1306)
- **Reliability and operations:** harden replace snapshot IPC; return retryable OpenAI provisioning errors; bind OAuth authorization codes to client_id; doctor classifies probe read-timeouts as cold, not unreachable; align GPU memory pressure defaults
- **Performance:** skip tag embedding when top_k &lt;= 0

### Features

* **bench:** demand-side token-reduction benchmark (Req 12, #1311)
* **mcp:** add describe_image tool (caption + zero-shot tags)
* **mcp:** add describe_image tool (caption + zero-shot tags) — Req 12 #1310
* **mcp:** cap describe_image payload size before cluster calls
* **mcp:** claude.ai connector surface — OAuth bridge + skill ZIP (Req 12 #1312)
* **mcp:** sie_mcp edge with docs_to_markdown tool (Req 12 #1306)
* **mcp:** structured extraction + structured generation tools (Req 12 #1308)
* **mcp:** wire measured token-reduction figures into savings metadata
* **tools:** add sie doctor — per-capability cluster diagnostics
* **tools:** Florence-2 fallback for image OCR
* **tools:** sie_tools — Claude Code context-offload client for managed clusters

### Bug Fixes

* align GPU memory pressure defaults
* **config:** detect bundle config hash drift
* **config:** fingerprint model pool ownership
* **config:** harden replace snapshot IPC
* **config:** replace drifted export snapshots
* **deps:** bump sidecar prometheus for protobuf advisory
* **gateway:** address provisioning review feedback
* **gateway:** align provisioning contract docs
* **gateway:** decode native media JSON bytes
* **gateway:** dereference structured output schema refs
* **gateway:** make provisioning non-2xx universally
* **gateway:** preserve ref sibling schema semantics
* **gateway:** return retryable OpenAI provisioning errors
* **mcp:** address review feedback on structured tools
* **mcp:** bind OAuth authorization codes to client_id
* **mcp:** blank-env fallback for model ids; honor SIE_MCP_IMAGE_TOP_K=0
* **mcp:** deep-copy committed token-reduction figures in build_metadata
* **mcp:** validate embedding shapes in _top_k_tags
* **sdk:** normalize score image payloads for wire transport
* **sdk:** normalize score images for wire transport
* **server:** guard readiness for removed configs
* **server:** honor pool-aware model configs
* **server:** render qwen3 vl reranker document images in user prompt
* **server:** render Qwen3-VL reranker document images in user prompt
* **sie-cluster:** add spot toleration to AKS worker pool
* **tools:** address doctor review feedback
* **tools:** doctor classifies probe read-timeouts as cold, not unreachable
* **worker:** keep SGLang loads off event loop

### Performance Improvements

* **mcp:** skip tag embedding when top_k &lt;= 0

## v0.6.4 (2026-06-11)

### Highlights

- **New capabilities:** add `grant_admin_to_creator` (opt-in AAD-RBAC for caller); lock model-cache storage account to cluster VNet by default; install kubelogin and convert kubeconfig after AKS get-credentials; wire Azure provider tooling; add azure (AKS) terraform module; ship values-aks.yaml AKS overlay with the Azure module
- **Reliability and operations:** harden storage_allowed_ip_ranges CIDR validation; harden release guarded merge checks; resubscribe stale NATS health stream; emit `az aks get-credentials --overwrite-existing`; drop unreachable final_registry guard so ACR path can fire

### Features

* **azure-terraform:** add `grant_admin_to_creator` (opt-in AAD-RBAC for caller)
* **azure-terraform:** lock model-cache storage account to cluster VNet by default
* **cluster:** install kubelogin and convert kubeconfig after AKS get-credentials
* **cluster:** wire Azure provider tooling
* **deploy:** add azure (AKS) terraform module
* **helm:** ship values-aks.yaml AKS overlay with the Azure module

### Bug Fixes

* **azure-terraform:** emit `az aks get-credentials --overwrite-existing`
* **azure-terraform:** harden storage_allowed_ip_ranges CIDR validation
* **ci:** harden release guarded merge checks
* **cluster:** address review feedback on Azure provider wiring
* **cluster:** drop unreachable final_registry guard so ACR path can fire
* **cluster:** set TF_VAR_* on Azure destroy path (same as create)
* **deploy:** revert system pool default to Standard_D4s_v3 (zoned everywhere)
* **gateway:** resubscribe stale NATS health stream
* **sidecar:** preserve msgpack work item payloads

## v0.6.3 (2026-06-10)

### Highlights

- **New capabilities:** add azure blob payload store support; add server-side copy fast path for cloud weight sync; informational generation eval CI gate over committed floors; add vision (image) input to generate(); preserve text/image content-part ordering; vision (image) input for generate()
- **Reliability and operations:** harden cloud cache sync paths; clear HIGH Dependabot alerts (docling, rustls-webpki); ensure cloud weight sync creates local parents; evict stale gateway workers on shutdown; fall back to relay on S3/GCS server-side copy failure
- **Performance:** engage conformant image preprocessing for v1; engage conformant image preprocessing for v1 (1.8x)

### Features

* add azure blob payload store support
* add server-side copy fast path for cloud weight sync
* **bench:** informational generation eval CI gate over committed floors
* **generate:** add vision (image) input to generate()
* **generate:** preserve text/image content-part ordering
* **generate:** vision (image) input for generate()
* support azure blob cluster cache
* **tester-cluster:** rtx6000 g7e.4xlarge + sglang preload + hf-token wiring

### Bug Fixes

* address azure cache review feedback
* address final cloud storage review issues
* **bench:** harden generation eval gate per review
* **deps:** clear HIGH Dependabot alerts (docling, rustls-webpki)
* ensure cloud weight sync creates local parents
* evict stale gateway workers on shutdown
* fall back to relay on S3/GCS server-side copy failure
* **generate:** address CodeRabbit review on vision input
* **generate:** address huronat review on vision input (F2-F8)
* **generate:** image-free content_parts field must not shadow layout
* **generate:** reject both-present image-bearing content layouts
* harden cloud cache sync paths
* **loadtest-ci:** self-heal orphaned cluster + stale lock in preflight
* **nemo_colembed:** trim left-padding rows from v1 conformant doc embeddings
* normalize local weight sync destination
* **quality-adapter:** gate v1 Vidore3 on English; finalize ?lang= plumbing
* **skill:** add bash language tag to hfCache --set fenced block (MD040)
* **skill:** move inline comments off shell continuation lines so the helm snippet pastes cleanly
* support cloud source weight sync
* **tester-cluster:** update rtx6000-spot machineType doc to g7e.4xlarge to match terraform

### Performance Improvements

* **nemo_colembed:** engage conformant image preprocessing for v1
* **nemo_colembed:** engage conformant image preprocessing for v1 (1.8x)

## v0.6.2 (2026-06-08)

### Highlights

- **New capabilities:** defer sie-config NATS startup and honor log levels; refresh KEDA Tilt local dev branch; M4 dense encoders — mxbai-embed-large-v1, arctic-embed-l-v2.0, modernbert-embed-base; add daily guarded stable releases
- **Reliability and operations:** accept dense dim in qwen3 vl embedding adapter; preserve model query templates in mteb eval; scale single-profile bundles on gpu-agnostic demand; consolidate runtime ninja install; install ninja in cuda runtime

### Features

* defer sie-config NATS startup and honor log levels
* **dev:** refresh KEDA Tilt local dev branch
* **models:** M4 dense encoders — mxbai-embed-large-v1, arctic-embed-l-v2.0, modernbert-embed-base
* **release:** add daily guarded stable releases

### Bug Fixes

* accept dense dim in qwen3 vl embedding adapter
* **bench:** preserve model query templates in mteb eval
* **dev:** address KEDA Tilt PR review
* **helm:** scale single-profile bundles on gpu-agnostic demand
* **server:** consolidate runtime ninja install
* **server:** install ninja in cuda runtime
* **server:** install ninja in CUDA SGLang runtime
* **terraform:** deny non-HTTPS access on state and quality-eval S3 buckets

## v0.6.1 (2026-06-07)

### Highlights

- **New capabilities:** configure GPU disk sizing and generate smoke; support static queue pools
- **Reliability and operations:** fail fast on invalid static pool config; pin kind smoke workers to default queue pool; canonicalize static queue pool names; stabilize GPU disk Terraform test

### Features

* configure GPU disk sizing and generate smoke
* **gateway:** support static queue pools

### Bug Fixes

* address GPU disk review comments
* **ci:** pin kind smoke workers to default queue pool
* **gateway:** canonicalize static queue pool names
* **gateway:** fail fast on invalid static pool config
* stabilize GPU disk Terraform test

## v0.6.0 (2026-06-07)

### Highlights

- **Breaking change:** Queue work subjects and pool streams use the new sie.work.\{pool\}.\{machine_profile\}.\{bundle\}.\{model\} shape only; legacy subject filters are intentionally not preserved.; workers will subscribe to `sie.work.*.<poolName>` instead of `sie.work.*.default`. Deployed alone (without the matching gateway/sidecar update that publishes/filters on the new subject) this will break routing on every cluster. To preserve the old shared-queue behavior, set `workers.common.queuePool: "default"` explicitly.
- **New capabilities:** route work by queue pool lanes; default SIE_POOL to pool name (not "default")
- **Reliability and operations:** harden queue lane admission; bump vitest 2.1.9 -&gt; 4.1.0 (CVE-2026-47429); align lane defaults and tilt e2e; preserve worker-group queue defaults

### ⚠ BREAKING CHANGES

* **gateway:** Queue work subjects and pool streams use the new sie.work.\{pool\}.\{machine_profile\}.\{bundle\}.\{model\} shape only; legacy subject filters are intentionally not preserved.
* **helm:** workers will subscribe to `sie.work.*.<poolName>` instead of `sie.work.*.default`. Deployed alone (without the matching gateway/sidecar update that publishes/filters on the new subject) this will break routing on every cluster. To preserve the old shared-queue behavior, set `workers.common.queuePool: "default"` explicitly.

### Features

* **gateway:** route work by queue pool lanes
* **helm:** default SIE_POOL to pool name (not "default")

### Bug Fixes

* **deps:** bump vitest 2.1.9 -&gt; 4.1.0 (CVE-2026-47429)
* **gateway:** harden queue lane admission
* **helm:** align lane defaults and tilt e2e
* **helm:** preserve worker-group queue defaults

## v0.5.0 (2026-06-04)

### Highlights

- **Breaking change:** `workers.pools.<name>.bundle` (string), `workers.pools.<name>.minReplicas`, `workers.pools.<name>.maxReplicas`, `workers.pools.<name>.extraEnv`, and `workers.pools.<name>.imageBundle` are replaced by `workers.pools.<name>.bundles.<bundle>.{minReplicas, maxReplicas, extraEnv, imageBundle, enabled}`. `workers.common.bundle` is removed (no longer consumed). StatefulSet, ScaledObject, PDB, and image-prepull DaemonSet names change from `worker-{pool}` to `worker-{pool}-{bundle}`, so in-place upgrades require deleting the old resources first.
- **New capabilities:** agent-jobs text-gen readiness — code/SQL/tools/guard evals + Qwen3.6-27B + precision routing; transfer sie-cluster claude skill; P(unsafe) logprob threshold for CHECK POLICY precision; split worker pools into pool × bundles schema; surface code/sql/guard capabilities; resolve job aliases in configs/resolve; add sglang worker pool for generative models
- **Reliability and operations:** expose unauthenticated metrics scrape port; expose unauthenticated metrics scrape port safely for prom; preserve gateway metrics scrape labels; drop unsupported ebnf advertisement + restore guardian a100 guard threshold; fail-fast on missing Spider DBs + order-sensitive SQL exec accuracy

### ⚠ BREAKING CHANGES

* **helm:** `workers.pools.<name>.bundle` (string), `workers.pools.<name>.minReplicas`, `workers.pools.<name>.maxReplicas`, `workers.pools.<name>.extraEnv`, and `workers.pools.<name>.imageBundle` are replaced by `workers.pools.<name>.bundles.<bundle>.{minReplicas, maxReplicas, extraEnv, imageBundle, enabled}`. `workers.common.bundle` is removed (no longer consumed). StatefulSet, ScaledObject, PDB, and image-prepull DaemonSet names change from `worker-{pool}` to `worker-{pool}-{bundle}`, so in-place upgrades require deleting the old resources first.

### Features

* agent-jobs text-gen readiness — code/SQL/tools/guard evals + Qwen3.6-27B + precision routing
* **agents:** transfer sie-cluster claude skill
* **guard:** P(unsafe) logprob threshold for CHECK POLICY precision
* **helm:** split worker pools into pool × bundles schema
* **models:** surface code/sql/guard capabilities; resolve job aliases in configs/resolve
* **tester-cluster:** add sglang worker pool for generative models

### Bug Fixes

* **agents:** address sie cluster review comments
* **bench:** fail-fast on missing Spider DBs + order-sensitive SQL exec accuracy
* **gateway:** expose unauthenticated metrics scrape port
* **gateway:** expose unauthenticated metrics scrape port safely for prom
* **guard:** reject multi-candidate sampling + keep logprobs consistent on rewrite
* **guard:** robust verdict thresholding, logprob hygiene, decoded-token logprobs
* **helm:** fail-fast on missing/invalid bundle replica bounds
* **helm:** preserve gateway metrics scrape labels
* **helm:** use sidecar binary for image pre-pull
* **models:** drop unsupported ebnf advertisement + restore guardian a100 guard threshold
* **sie_server:** honor params.instruction in Florence-2 extract
* **tester-cluster:** cap rtx6000 default bundle to avoid over-subscription
* **tools:** via-SIE EBNF response_format shape + request/preload model split

## v0.4.2 (2026-06-03)

### Highlights

- **New capabilities:** 5-domain generation bench + via-sie quality matrix + gateway schema gaps; add e0-02 all-minilm time-share experiment; land coalesce_ms=5 + max_batch_requests=12 as Rust defaults; add --via-sie smoke path (route through sie_server); add min_tokens + system_prompt + temperature for G4 retry; close Qwen3.6-27B gap — min_tokens=10 + max=768 + ctx=4096
- **Reliability and operations:** set verbose=True on SIEServer so launch errors surface; document worker-sidecar metrics wiring; gate sidecar nats reconnect refresh; harden sidecar config recovery; budget loadtest barrier timeouts
- **Performance:** anchor min_batch_cost floor at max_batch_tokens // 4; tighten adaptive wait ceiling + revert gte-multilingual 32k; rebind vision Conv3d patch-embed to F.linear; raise max_batch_tokens 16k → 32k to stop IPC-batch shred

### Features

* 5-domain generation bench + via-sie quality matrix + gateway schema gaps
* add e0-02 all-minilm time-share experiment
* **batch_config:** land coalesce_ms=5 + max_batch_requests=12 as Rust defaults
* **bench-27b:** add --via-sie smoke path (route through sie_server)
* **bench-27b:** add min_tokens + system_prompt + temperature for G4 retry
* **bench-27b:** close Qwen3.6-27B gap — min_tokens=10 + max=768 + ctx=4096
* **bench-27b:** launch full SIE stack (NATS+worker+gateway) for --via-sie
* **bench+model:** via-sie 4-task n=300 sweep + NEXTN smaller-draft on 27B
* **bench:** 0.6B via-sie validated; harness + 27B config gains
* **bench:** 5-shot CoT for CaseHOLD (item 5 — close 27B target gap)
* **bench:** fix Qwen3-0.6B GPQA (parrot bug) + 27B diagnostics; final matrix
* **bench:** improve perf eval output handling
* **docling:** accept image input + run on OCR-bench quality path
* **gateway+worker:** chat surface accepts min_tokens + chat_template_kwargs
* **gateway:** strengthen generation isolation guardrails
* **latency:** tighten FetchExpiryController defaults to 2/15/50
* **model+bench:** RTX-PRO-6000 FP8 profile for Qwen3.6-27B + 6000 validation
* **model:** bump Qwen3-0.6B serving context 1024→4096 for prod simple-task use
* **models:** add Marqo/marqo-fashionSigLIP (SigLIP open_clip, fashion image-text)
* **ocr:** docling accepts images + quality eval prefers documents
* reconcile live worker config in sidecar
* RTX PRO 6000 FP8 profile for Qwen3.6-27B + SIE-on-6000 generative benchmark matrix
* **scheduler:** load-aware pipeline_depth autotune (S14 follow-up)
* **scheduler:** production-parity defaults + serial pipeline (carveout p99 fix)
* **scheduler:** restore SIE_RUST_PIPELINE_DEPTH=2 default (deep-saturation fix)
* **scheduler:** SIE_PULL_QUANTUM_INCLUDE_QUEUE_MS for Py-main parity
* **scheduler:** SIE_RUST_WAVE_CADENCE env toggle (default on)
* **scheduler:** step adaptive controller once per wave (Python parity)
* **sidecar:** add worker config and pool admission reconciliation
* **sidecar:** wire generation direct dispatch
* **sie_server:** add MinerU2.5-Pro-2604-1.2B doc OCR adapter
* **sie_server:** carve out QueueExecutor + IPC types for Rust worker POC
* **sie_server:** integrate MinerU2.5-Pro-2604-1.2B doc OCR adapter
* **sie_server:** UDS msgpack IPC server for Rust worker sidecar
* **sie_worker_rust:** close parity gaps with Python pull loop + smoke test
* **sie_worker_rust:** scaffold Rust worker sidecar crate (Phase 1c)
* **sie_worker_rust:** wire end-to-end NATS -&gt; IPC -&gt; publish loop (Phase 1d)
* **sie-bench:** synchronize loadtest measurement start
* **worker/rust:** IPC connection pool — lift the sidecar's last serialization bottleneck
* **worker/rust:** narrate the hot path — structured INFO, slow-RPC + heartbeat-streak WARNs, full error chains
* **worker:** introduce InferenceBackend trait + BackendRouter
* **worker:** native Candle BERT backend behind `candle` feature

### Bug Fixes

* accept dense_dim in dense adapters
* **adapters:** replace Qwen3-VL vision Conv3d patch-embed with matmul
* **adapters:** route Qwen3-VL VLMs through flash attention (Vidore3 throughput)
* address pr review quality issues
* **bench-27b:** drop bundle from SIEServer (sie-server rejects bundle+models combo)
* **bench-27b:** set verbose=True on SIEServer so launch errors surface
* **bench-27b:** skip chat_template_kwargs on via-sie (gateway rejects unsupported field)
* **bench-27b:** wait for sie-server /healthz (not /health)
* **bench:** bump casehold/gpqa max_tokens to 2048 (CoT truncation)
* **bench:** let via-SIE smoke serve a profile-variant model end-to-end
* **bench:** resolve CPU deps for quality server
* **catalog:** include eval-matrix tasks so dispatch filter accepts them
* **ci:** address analyzer findings and stale queue test
* **ci:** avoid nested mise in integration fixture
* **ci:** keep sidecar out of warm cache
* **ci:** refresh gateway openapi contract
* correct e0 vm runbook paths
* **deploy:** add sidecar registry resources
* **deploy:** address server sidecar review feedback
* **deploy:** align server sidecar naming
* **deploy:** align server sidecar naming and kind preload smoke
* **deploy:** align tilt sidecar image naming
* **deploy:** document worker-sidecar metrics wiring
* **deploy:** keep sidecar on GHCR by default
* **deploy:** normalize server sidecar naming
* **deploy:** publish server sidecar image
* **deploy:** rename sidecar container to worker-sidecar
* **deploy:** wire SIE server sidecar for kind smoke
* **deploy:** wire worker sidecar image across kind and cloud
* gate sidecar nats reconnect refresh
* **gateway+server:** queue is the only mode — kill direct-mode cruft
* **gateway:** suppress H9 first-chunk-fallback on single-worker pools
* harden sidecar config recovery
* **impact-map:** keep profiles distinct when adapter_options differ
* keep generation machinery off default queue path
* **loader:** wire profile runtime.default_sampling into the adapter
* **modal:** report actual GPU on remote, not stale env-default
* **model:** bump Qwen3.6-27B default/h100 mem_fraction_static 0.85 → 0.92
* **orchestrator:** thread CLI -p profile through to client.extract
* preserve worker batch identity and publish image
* **product:** update design audit for topical docs
* **quality_eval:** take results-bearing JSON envelope in load_eval_json
* **quality:** batch3 of CodeQL findings + bench KIE bug
* **quality:** batch3 of CodeQL findings + bench KIE root-cause
* **quality:** batches 1+2 of CodeQL quality findings
* **quality:** close CodeQL quality-tab findings
* **quality:** drop redundant inline imports in donut + registry
* **quality:** repair adapter eval harness regressions
* remove e0 preflight httpx dependency
* require rust sidecar for queue workers
* **review:** 0.6B ctx test 1024-&gt;4096, loader except logs, README gaps resolved
* **review:** recompute 27B target delta_vs_baseline for the 2048 scores
* run directory creation
* run e0 vm scripts via uv
* **scheduler:** autotune signal — observed_p50/target_p50 ratio
* scope bundle config hash cache per registry
* **security:** bump astro to ^6.4.2 for website
* **security:** bump gateway deps to patched versions
* **security:** bump product/gtm Python lockfiles
* **security:** bump product/gtm/content/slides npm transitives
* **security:** bump root pnpm deps + add overrides for transitives
* **security:** bump root Python deps to patched versions
* **security:** bump sie_dashboard npm deps to patched versions
* **security:** bump sie_ts_sdk standalone pnpm transitives
* **security:** bump sst to ^4 to drop vulnerable aws-sdk v2
* **security:** cap vite at ^6 + add Node engines to website
* **security:** close ~190 Dependabot alerts across 9 manifests
* **security:** sanitize one-pager template with DOMPurify
* **security:** use Reflect.construct for WebSocket headers shim
* **sie_bench:** send SIE profile via X-SIE-MACHINE-PROFILE header
* **sie_bench:** use rapidfuzz for OmniDocBench edit distance
* **sie_server:** clear CUDA cache on uncovered VLM paths + drop private sem _value access
* **sie_server:** VLM cache clears on uncovered paths + drop private sem _value access
* **sie-bench:** budget loadtest barrier timeouts
* slow sidecar nats consumer reconcile
* **smoke:** launch sie_server worker with -b sglang, not -m &lt;model&gt;
* **smoke:** preload the target model in via-sie worker
* **test:** restore donut helper call contract
* **worker-sidecar:** harden queue carveout contracts
* **worker/rust:** one long-lived pull stream — kill 30s ack_wait stall
* **worker/rust:** re-copy src after cargo chef cook so real build isn't a stub
* **worker/rust:** set CUDA_COMPUTE_CAP at build time (default 89, L4)
* **worker/rust:** stop shipping the cargo-chef stub binary as the real build
* **worker:** harden Candle backend + align dispatcher error contract
* **worker:** harden payload store + error paths; surface silent success bugs
* **worker:** SGLang adapter accepts min_new_tokens kwarg + 27B via-sie validated

### Performance Improvements

* **adaptive:** anchor min_batch_cost floor at max_batch_tokens // 4
* **batching:** tighten adaptive wait ceiling + revert gte-multilingual 32k
* **glm_ocr:** rebind vision Conv3d patch-embed to F.linear
* **gte-multilingual-base:** raise max_batch_tokens 16k → 32k to stop IPC-batch shred
* **mineru_vl:** O(L) incremental no-repeat-ngram for greedy decode
* **ocr:** swap pure-Python Levenshtein DP for rapidfuzz
* **rope_flash:** vectorize CLS/mean pooling, eliminate per-item .item() sync
* **server:** FP16 on GPU, coalesce sized for IPC bursts, starvation self-heal

### Reverts

* restore adaptive batching defaults to 15/50ms
* **scheduler:** drop depth autotune (signal didn't pan out in S17)

## v0.4.1 (2026-05-28)

### Highlights

- **New capabilities:** add Qwen3.6-27B model + migrate to CUDA 12.9
- **Reliability and operations:** isolate generation direct dispatch from shared queues; resolve 18 open CodeQL alerts; use SHA256 (not SHA1) for actor_id log tag; colocate tests under infra/, update sync contract

### Features

* **server:** add Qwen3.6-27B model + migrate to CUDA 12.9

### Bug Fixes

* isolate generation direct dispatch from shared queues
* **security:** resolve 18 open CodeQL alerts
* **security:** use SHA256 (not SHA1) for actor_id log tag
* **terraform-sync:** colocate tests under infra/, update sync contract

### Reverts

* **security:** drop advanced CodeQL setup

## v0.4.0 (2026-05-27)

### Highlights

- **Breaking change:** fail-closed authentication (default-deny)
- **New capabilities:** generation quality-gate scoring core (roadmap §5, trust-critical); generation-quality regression gate over the existing scorers; add cohere measurements for us-east-1; add openai measurements for us-east-1; add voyage measurements for us-east-1; regex/EBNF response_format + developer role (roadmap 1.7)
- **Reliability and operations:** forward provision_timeout_s in SIEImageTextWrapper.encode; raise image-task eval timeouts to fix Flickr30k nightly; exclude favicon + OG image from auth middleware; keep public surfaces vague about what's behind auth; revert NextAuth function-form, use try/catch on Resource
- **Performance:** warm one Lambda, bump timeout, narrow S3 verdict fetch

### ⚠ BREAKING CHANGES

* **gateway:** fail-closed authentication (default-deny)

### Features

* **bench:** generation quality-gate scoring core (roadmap §5, trust-critical)
* **bench:** generation-quality regression gate over the existing scorers
* **benchmarks:** add cohere measurements for us-east-1
* **benchmarks:** add openai measurements for us-east-1
* **benchmarks:** add voyage measurements for us-east-1
* **chat:** regex/EBNF response_format + developer role (roadmap 1.7)
* **dashboard:** add executive quality summary widget on landing
* **dashboard:** add hover-tooltip on 'to verify' explaining WARN
* **dashboard:** brand alignment foundation - palette, fonts, header
* **dashboard:** brand favicon, opengraph image, light-mode hover fix
* **dashboard:** brand foundation - palette, fonts, header logo
* **dashboard:** brand surfaces - retokenise landing page + widget
* **dashboard:** public preview shell on / so Slack unfurls work
* **dashboard:** retokenise badges - brand-elevated neutrals, DM Mono labels
* **dashboard:** retokenise landing surfaces onto brand palette
* **dashboard:** retokenise loadtest pages onto brand surfaces
* **dashboard:** retokenise quality pages onto brand surfaces
* **dashboard:** retokenise shared components and sign-in onto brand palette
* **dashboard:** treat WARN as passing-with-verify, shade gauge amber
* **gateway,sdk,server:** add native generate endpoint with improved admission control and validation
* **gateway,sdk,server:** add OpenAI-compatible chat completions with streaming and sampling extensions
* **gateway,server:** add multi-turn tool-use support with OpenAI-compatible message format
* **gateway:** /v1/completions (legacy OpenAI Completions, raw-prompt)
* **gateway:** /v1/completions streaming (text_completion SSE)
* **gateway:** /v1/generate accepts seed/logprobs/logit_bias/n/best_of/lora_adapter (M8)
* **gateway:** /v1/responses (OpenAI Responses API, MVP)
* **gateway:** /v1/responses structured array input (conversation history)
* **gateway+worker:** per-choice OpenAI streaming for n&gt;1 (H4, H5, M4)
* **gateway:** accept OpenAI multimodal content-parts; reject images (no VL model)
* **gateway:** add routing salt + byte-preserving key mode (M11)
* **gateway:** advertise lora_adapters on /v1/models + pre-validate unknown names
* **gateway:** fail-closed authentication (default-deny)
* **gateway:** meaningful system_fingerprint on chat responses (roadmap 1.3/§5)
* **gateway:** refactor streaming and routing with improved error handling and metrics
* **gateway:** register /v1/moderations as explicit 501 (roadmap 1.8, phase 3)
* **gateway:** serve a rendered API reference at /docs (Redoc)
* **gateway:** unify /v1/embeddings on the OpenAI error envelope (roadmap 1.4)
* **generation:** best_of — over-generate + rank by logprob, return top n
* **generation:** complete M4 req2 generation primitive with streaming, structured outputs, and routing
* **generation:** multi-candidate n&gt;1 (non-streaming) end-to-end (roadmap 1.5)
* **generation:** multi-LoRA serving (one base, N adapters, per-request) (roadmap 6.2)
* **generation:** ship generate() primitive — Qwen3.5-4B + NEXTN/MTP + xgrammar, adapter perf at parity with raw SGLang
* **generation:** streaming n&gt;1 — per-candidate SSE interleave
* **helm/sie-cluster:** bundle cert-manager + trust-manager (opt-in) with self-signed TLS mode
* **openapi:** add tool_calls support to chat completion schema
* **python-sdk:** expose typed params for chat n/logprobs/lora_adapter/etc (M7)
* **quality-eval:** add heartbeat logging and improve long-running process observability
* **routing:** cache-aware (prefix-hash) routing (roadmap §6.3)
* **sie_bench:** add Cohere as a first-class eval source
* **sie_bench:** add Cohere multimodal embeddings
* **sie_bench:** add Cohere rerank backend for native MTEB rerank tasks
* **sie_bench:** add OpenAI Embeddings as a first-class eval source
* **sie_bench:** add Voyage provider source plumbing
* **sie_bench:** implement Voyage text embedding runner
* **sie_dashboard:** add /quality/compare to diff two quality runs
* **terraform:** add default_tags Project=sie/Cluster on all AWS providers
* **terraform:** add on-demand RTX 6000 baseline pool to tester-cluster
* **terraform:** add uniform project=sie label across all GCP clusters
* **terraform:** idle-stop and on-demand wake for quality-eval runner fleet
* **terraform:** scale quality-eval fleet to 5+5, smart-wake, 4h timeout
* **terraform:** wake-runners retries + watchdog queued-jobs backstop
* **tester-cluster:** add on-demand L4 worker pool + capacityType node pins
* **ts-sdk:** handle 202 provisioning in chatCompletions + expose missing fields (H1+M6)
* **worker:** SGLang owns grammar; worker preflight opt-in only (H8, ADR-0002)
* **worker:** wire mixed-pool fairness scheduler into the pull-loop (opt-in)
* **worker:** WorkClassScheduler core for mixed-pool fairness (roadmap §6.1)

### Bug Fixes

* **bench:** declare olmocr[bench] dep for OCR-bench quality eval
* **bench:** drop --with-deps from playwright install (no sudo on g7e)
* **bench:** forward provision_timeout_s in SIEImageTextWrapper.encode
* **bench:** install playwright chromium for OCR-bench KaTeX rendering
* **bench:** playwright install --with-deps for OCR-bench chromium
* **bench:** raise image-task eval timeouts to fix Flickr30k nightly
* **bench:** respect similarity() inputs in ColBERT/ColPali wrappers
* **bench:** run olmocr.bench.tests off orchestrator's asyncio loop
* centralize worker_id subject normalization (M5)
* **chart:** pin image-prepull DaemonSets to GPU nodes
* **ci:** copy assets/ into the gateway Docker build (redoc bundle)
* **dashboard:** close remaining 'SIE Dashboard' leaks in &lt;title&gt; and og:image:alt
* **dashboard:** drop edge runtime on opengraph-image for OpenNext
* **dashboard:** exclude favicon + OG image from auth middleware
* **dashboard:** keep public surfaces vague about what's behind auth
* **dashboard:** pick healthy daily via coverage + health gates
* **dashboard:** require &gt;=50 pairs on main-run fallback in nightly picker
* **dashboard:** revert NextAuth function-form, use try/catch on Resource
* **dashboard:** short tab title for authed users, neutral for unauth
* **dev:** set explicit auth opt-in for local gateway launchers (post fail-closed)
* **gateway,sdk,server:** add wire-level validation, improve resource cleanup, and enhance observability across request lifecycle
* **gateway,sdk,server:** prevent metric cardinality DoS and fix non-idempotent retry logic
* **gateway,sdk,server:** strengthen validation and eliminate silent failures across request lifecycle
* **gateway,sdk,server:** validate numeric fields and improve error handling
* **gateway:** add NATS config trusted producers helm override
* **gateway:** document ModelCapabilities in OpenAPI + refresh on profile delta-update
* **gateway:** generation timeouts bypass legacy request-timeout ceiling (H7)
* **gateway:** scope LoRA adapter capabilities per profile (M10)
* **gateway:** strict allow-list + 400 contract on /v1/completions (H3)
* **gateway:** strict allow-list + 400 contract on /v1/responses (H2)
* **gateway:** tighten chat sampler/token-cap + tool-history validation (M1, M13)
* **gateway:** trust chart-rendered sie-config pod name for NATS deltas
* **generation:** cancel tombstone prevents first-chunk fallback double-execution (H9)
* **generation:** LoRA lora_path is a top-level /generate field, not a sampling param
* **generation:** tighten lossy tool-control flags (M14)
* **grammar:** resolve tokenizer adapter for Outlines processor factories and remove anchors from regex patterns
* **helm/sie-cluster:** guard validateTls probe against deployments with nil labels
* **helm/sie-cluster:** guard validateTls probe against nil deployment labels
* **helm/sie-cluster:** include cert-manager mode in presence-check gate
* **helm/sie-cluster:** label-based cert-manager detection + bidirectional runtime check
* **helm/sie-cluster:** make self-signed root-CA namespace configurable
* **helm/sie-cluster:** one-step bundled cert-manager install with self-signed TLS
* **helm/sie-cluster:** probe cert-manager controllers cluster-wide
* **helm/sie-cluster:** regenerate Chart.lock with synced digest
* **helm:** trim and drop empty entries in ingress.hosts
* **modal:** exclude cargo target/ from sandbox image mount
* **pytorch_embedding:** accept and forward revision kwarg
* **quality_eval:** tolerate stdout noise around eval JSON envelope
* **quality-eval:** handle paginated jobs API and align last two filters
* **release-docker,warm-cache:** address review findings
* **server,gateway:** add GPU-aware health probes to detect and recover from wedged CUDA contexts
* **server,gateway:** GPU-aware health probes to detect & recover from wedged CUDA contexts
* **sie_bench:** register OPENAI_SOURCE so --save-targets openai actually saves
* **sie_dashboard:** address compare-page review nits
* **sie_dashboard:** label compare log links with run IDs
* **sie_server:** base64-decode JSON image inputs
* **sie_server:** enforce media bytes contract at every consumer
* **sie_server:** install cv2 system libs for docling extract
* **streaming:** no-silent-drop on chunk-queue backpressure (H6)
* **terraform:** detect in-flight workflow runs via explicit status query
* **terraform:** drop redundant Project overrides in quality-eval-l4
* **terraform:** drop watchdog_idle_minutes default to 5, ignore PIP drift
* **terraform:** grant ec2:DescribeInstanceStatus to wake role
* **terraform:** per-runner idle-stop via GitHub Actions runners API
* **terraform:** require positive activity observation before idle-stop
* **terraform:** scope quality-eval IAM on Role tag instead of Project
* **terraform:** seed GPU node-group desired_size from min_size
* **terraform:** watchdog backstop covers queued-status runs
* **terraform:** watchdog grants HANG_MINUTES grace from LaunchTime
* **terraform:** watchdog ignores LastBusyAt older than LaunchTime
* **tester-cluster:** pin worker pool nodeSelectors to gpu-type as well

### Performance Improvements

* **dashboard:** warm one Lambda, bump timeout, narrow S3 verdict fetch

### Reverts

* **release-docker:** drop matrix consolidation, keep deps push retry

## v0.3.4 (2026-05-14)

### Highlights

- **New capabilities:** default payload store to model-cache bucket /payloads; typed InputTooLongError for extract 400 INPUT_TOO_LONG
- **Reliability and operations:** bump dev-g6-spot to g6.2xlarge; bump dev-g6-spot to g6.2xlarge so default worker pool fits; default workers to shared queue pool; pin opencv-python-headless to drop X11 runtime deps; tolerate config conflicts in bootstrap, gate on sie-config

### Features

* **infra:** default payload store to model-cache bucket /payloads
* **sdks:** typed InputTooLongError for extract 400 INPUT_TOO_LONG

### Bug Fixes

* **aws-example:** bump dev-g6-spot to g6.2xlarge
* **aws-example:** bump dev-g6-spot to g6.2xlarge so default worker pool fits
* **chart:** default workers to shared queue pool
* **cluster.py,aws.py:** address review suggestions 1 & 2
* **deps:** pin opencv-python-headless to drop X11 runtime deps
* **gateway:** tolerate config conflicts in bootstrap, gate on sie-config
* **gateway:** tolerate config conflicts in bootstrap, gate on sie-config ready
* **sdk:** widen sie-sdk requires-python to &gt;=3.12
* **terraform-aws:** default ECR creation off, prefix repo names with project_name
* **terraform-aws:** trim slashes from ecr_repository_prefix
* **terraform-google-sie:** wait for identity pool before binding WI
* **terraform:** relax required_version from ~&gt; 1.14.3 to &gt;= 1.14

## v0.3.3 (2026-05-13)

### Highlights

- **New capabilities:** add ColQwen3 + Nemotron ColEmbed v2 visual doc retrieval; add text classification task support; add post-download load timeout with stall-based download bounds; add scope-able workflow_dispatch with model/profile/task filters; add new INPUT_TOO_LONG ErrorCode; enforce overflow_policy in gliclass adapter
- **Reliability and operations:** surface empty matrix and add measurement-mode for unbaselined adapters; emit task_class in quality-adapter JSON output; score detection eval predictions from result["objects"]; annotate empty-diff path that bypasses impact_map; capture real exit code from impact_map in resolve-impact.sh

### Features

* **adapters:** add ColQwen3 + Nemotron ColEmbed v2 visual doc retrieval
* **extraction:** add text classification task support
* **model-loader:** add post-download load timeout with stall-based download bounds
* **quality-adapter:** add scope-able workflow_dispatch with model/profile/task filters
* **server:** add new INPUT_TOO_LONG ErrorCode
* **server:** enforce overflow_policy in gliclass adapter
* **server:** route INPUT_TOO_LONG to HTTP 400 in extract API
* **server:** validate overflow_policy in resolve_runtime_options

### Bug Fixes

* **bench:** emit task_class in quality-adapter JSON output
* **bench:** score detection eval predictions from result["objects"]
* **ci:** annotate empty-diff path that bypasses impact_map
* **ci:** capture real exit code from impact_map in resolve-impact.sh
* **ci:** collapse adapter-equivalent profiles in quality-adapter matrix
* **ci:** pin mise to 2026.5.5 in loadtest workflows
* **ci:** surface empty matrix and add measurement-mode for unbaselined adapters
* **docker:** install libspatialindex-c6 in worker images
* **gliclass:** catch IndexError empty-tensor crash as InputTooLongError
* **gliclass:** raise InputTooLongError from argmax-empty backstop
* **probe-chart:** make 3-OLD/2-NEW sample asymmetry explicit in title
* **quality-adapter:** namespace quality_eval tests to avoid conftest collision
* **quality-adapter:** split adapter_paths on commas before --changed-dirs
* **quality-adapter:** split Pair column so pair_key | stops breaking the table
* **quality-adapter:** split Pair column to stop pair_key | breaking the table

## v0.3.2 (2026-05-08)

### Highlights

- **New capabilities:** default score_pairs() in BaseAdapter + baseline reranking targets; bump cold-start schema to v6 with deserialize/warmup split; per-model perf concurrency defaults for OCR adapters; adapter-triggered quality eval on persistent L4 runner; make destroy conditional on workflow_dispatch input; nightly loadtest pipeline + baseline recorder
- **Reliability and operations:** raise loadtest job timeout to GH Actions ceiling (360 min / 6h); clarify experimental NATS health mode; lang tag on fenced block; fail fast on invalid scenarios; surface error/no_results rows in MD; harden parse_label against unexpected filenames; adapt OCR perf Item shape to model.inputs and fail loudly on errors

### Features

* **adapters:** default score_pairs() in BaseAdapter + baseline reranking targets
* **bench,charts:** bump cold-start schema to v6 with deserialize/warmup split
* **bench:** per-model perf concurrency defaults for OCR adapters
* **ci:** adapter-triggered quality eval on persistent L4 runner
* **ci:** make destroy conditional on workflow_dispatch input
* **ci:** nightly loadtest pipeline + baseline recorder
* **colbert:** add score_pairs support and expand model coverage
* **dashboard:** add status and kind filters to runs list
* **dashboard:** introduce run-group concept (run = 3 scenarios)
* **dashboard:** loadtest results dashboard (Next.js + SST + DynamoDB)
* **dashboard:** render every metric in the perf-lab archive
* **dashboard:** scaffold loadtest dashboard (Next.js + SST)
* **dashboard:** status and kind filters on runs list
* **dashboard:** track run_status; gh-API one-time backfill
* **docling:** add ocr profile defaulting do_ocr=true
* **gateway:** expose OpenAPI contract
* **gateway:** unify API errors and align probe contracts
* **helm:** expose probes value trees for worker/gateway/config
* **helm:** tighten startup/readiness probes for faster pod-ready
* **helm:** TLS termination via cert-manager + BYO matrix docs
* **helm:** wire probe templates to values trees
* **infra:** opt-in S3 cluster model cache
* **ltfr:** cache-vs-no-cache compare chart, with-cache run data, and 8 single-mode chart refresh
* **matrix:** add task_class stamping to eval measurements
* **server,bench:** split deserialize/warmup in cold-start instrumentation (v6)
* **server:** cap torch CPU threads at worker startup
* **sie_server:** per-stage timing markers in lifespan for engine_boot attribution
* **sie_server:** split adapter.warmup() out of load() with cold-start log markers
* **tools:** bump cold-start bench to v5 with scenario flag
* **tools:** LTFR per-scenario bench tooling + results (issue #652)
* **tools:** ltfr-bench orchestrator (issue #652)

### Bug Fixes

* **bench:** adapt OCR perf Item shape to model.inputs and fail loudly on errors
* **bench:** address CodeRabbit review on PR #779
* **bench:** correctly detect v6 split presence in flattened runs[]
* **bench:** derive emitted gpu_load_s from v6 deserialize+warmup split when available
* **chart:** pass --cluster-cache to sie-server and correct populate command in docs
* **charts:** vertical legend so 'image pull + container init' and 'node prov' aren't clipped
* **ci+terraform:** three deterministic root causes for loadtest pipeline
* **ci:** address CodeRabbit findings on quality-adapter PR
* **ci:** address CodeRabbit's second-pass review on quality-adapter
* **ci:** address CodeRabbit's third-pass review on quality-adapter
* **ci:** auto-clear stale terraform state lock from prior runner crashes
* **ci:** drop double cuda12 suffix + force codebuild for missing images
* **ci:** forensic dump on argo failure + LB-ENI release before destroy
* **ci:** gate stale-lock clear behind force_unlock input + pass --aws-region to destroy
* **ci:** override registry/gpu-selector/tolerations + Python heredoc
* **ci:** parse markdown bench output → result.json synthesis
* **ci:** pass WORKFLOW env to run_scenarios.sh in loadtest.yml
* **ci:** preflight env-var check in run_scenarios + finalize scripts
* **ci:** provision GH PAT secret + in-cluster github-token before bootstrap
* **ci:** raise loadtest job timeout to GH Actions ceiling (360 min / 6h)
* **ci:** read bench-config from local clone, not raw.githubusercontent.com
* **ci:** right-size bench pod + worker pod resources for cluster shape
* **cluster:** move orphan-LB sweep into cmd_destroy, drop parallel script
* **cluster:** use project_name (not example name) for orphan-LB VPC tag
* **cluster:** use project_name for orphan-LB VPC tag lookup
* **dashboard,ci:** keep run_status consistent between S3 and DynamoDB
* **dashboard,ci:** wire real Prometheus matrix shape + extend headlines
* **dashboard:** drop time-based legacy run grouping (was unsafe)
* **dashboard:** GPU util shown as 0-100 (was being multiplied by 100 again)
* **dashboard:** include duration_seconds in run-meta.json (was DynamoDB-only)
* **dashboard:** normalize array-shaped searchParams before .trim()
* **deps:** bump plotly to &gt;=6.1.1 for kaleido compat
* **docker:** stub bundles/ and models/ in deps stage
* **docling:** cache DocumentConverter per (device, ocr_enabled)
* **docling:** mark adapter unloaded in unload()
* **docling:** thread device through PdfPipelineOptions accelerator_options
* **gateway:** address latest coderabbit contract notes
* **gateway:** address PR review for NATS health mode
* **gateway:** address probe and SDK review findings
* **gateway:** align CreatePoolRequest OpenAPI with runtime validation
* **gateway:** clarify experimental NATS health mode
* **gateway:** close remaining review contract gaps
* **gateway:** preserve embeddings timing headers
* **gateway:** preserve scale-from-zero request path
* **gateway:** reject unsupported embeddings token arrays
* **helm:** validate ACME server and privateKeySecretRef in validateTls
* **infra:** grant kms:Decrypt to workers when model cache uses SSE-KMS
* **infra:** normalize whitespace-only model cache string inputs
* **infra:** treat empty model_cache_kms_key_id as unset
* **infra:** use flat lifecycle key for s3-bucket module v5
* **loadtest-ci:** force-delete orphan elbv2 LBs before terraform destroy
* **loadtest-ci:** force-delete orphan LBs and stop swallowing destroy failures
* **loadtest-ci:** poll workflow phase instead of argo submit --wait
* **loadtest-ci:** poll workflow phase instead of relying on argo --wait
* **ltfr-bench,notes:** lang tag on fenced block; fail fast on invalid scenarios; surface error/no_results rows in MD
* **ltfr-bench:** hoist imports to top; guard payload.results shape
* **ltfr-bench:** mark request_failed rows in scenario-a/b MD tables
* **ltfr-bench:** preserve failure context in aggregated rows; add request_failed status
* **ltfr-bench:** treat no_results cells as failures in exit code
* **ltfr-charts:** strip legend clip-path so labels render full width
* **ltfr:** tighten UID/timestamp guards in capture_image_pull_events
* **multi_pod_cold_start:** raise on ASG terminate fail; UID-filter pull events; isolate scenario-c pod
* **paddleocr_vl:** pass use_cache=True to generate
* **paddleocr_vl:** pass use_cache=True to generate to enable KV-cache
* **review:** tighten score_pairs options handling and query text validation
* **sie-server:** include model and bundle directories in wheel distribution
* **terraform:** detect HF-cache EBS by NVMe model + size, not by Linux name
* **terraform:** set resolve_conflicts_on_* = OVERWRITE on EKS addons
* **tools:** drop module-level docstrings (AGENTS.md rule)
* **tools:** guard fig_per_cell_table aggregation against empty results
* **tools:** guard mean() against empty engine_boot_s in aggregate()
* **tools:** harden parse_label against unexpected filenames
* **tools:** mark cold-start-bench.py executable (EXE001)
* **tools:** remove module docstring from cold_start_charts.py (repo rule)

## v0.3.1 (2026-04-29)

### Highlights

- **New capabilities:** add OmniDocBench OCR quality loader; support /v1/score with dense/sparse/colbert/hybrid modes; add Marqo/marqo-ecommerce-embeddings-B via open_clip backend
- **Reliability and operations:** add terminal failed state to model registry (sie-test#85)

### Features

* **bench:** add OmniDocBench OCR quality loader
* **bge-m3:** support /v1/score with dense/sparse/colbert/hybrid modes
* **siglip:** add Marqo/marqo-ecommerce-embeddings-B via open_clip backend

### Bug Fixes

* **server:** add terminal failed state to model registry (sie-test#85)

## v0.3.0 (2026-04-29)

### Highlights

- **Breaking change:** openapi.json is now a committed artifact that must be regenerated and committed when API changes are made
- **New capabilities:** add GLM-OCR adapter; add Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B multimodal adapters; add GLiNER2 and GLiNER-bi adapters; add Qwen3-Reranker-0.6B and 4B causal LM reranker support; add SigLIP 2 base-patch16-224 vision-language encoder; add minimal cache weights snapshot command for offline deployments
- **Reliability and operations:** retry only transient connection errors under wait_for_capacity; surface unrouteable models loudly and helm-repo-add on pristine hosts; emit identical NATS payload to bundle and _all subjects; surface mixed-profile unrouteable models and keep snapshot consistent on writes; add retry logic for deadsnakes PPA to handle Launchpad outages
- **Performance:** cache JPEG-encoded corpus images across queries; lazily JPEG-encode corpus images on first use; cache SDK version parse, integer audit latency, UUIDv7; cut hot-path allocations, fuse numpy decode, tighten backpressure

### ⚠ BREAKING CHANGES

* **openapi:** openapi.json is now a committed artifact that must be regenerated and committed when API changes are made

### Features

* **adapters:** add GLM-OCR adapter
* **adapters:** add Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B multimodal adapters
* add GLiNER2 and GLiNER-bi adapters
* add Qwen3-Reranker-0.6B and 4B causal LM reranker support
* add SigLIP 2 base-patch16-224 vision-language encoder
* **admin:** add minimal cache weights snapshot command for offline deployments
* **bench:** honor SIE_BENCH_SERVER_READY_TIMEOUT in eval orchestrator
* **ci:** nightly loadtest gate against dedicated EKS cluster
* **ci:** nightly loadtest gate, ephemeral cluster per run
* **extract:** add Docling adapter for PDF/DOCX/HTML extraction
* **extract:** add Docling adapter for PDF/DOCX/HTML parsing
* **extract:** plumb document items and structured `data` results
* **observability:** add Prometheus metrics to sie-config and expand sie-gateway coverage
* **observability:** Prometheus metrics for sie-config and sie-gateway
* **oom:** implement defensive exception fan-out and improve recovery metrics
* **oom:** improve error semantics and budget exhaustion detection
* **openapi:** add static spec export and validation
* **router:** import Rust gateway source tree
* **server:** add reactive OOM recovery and proactive idle eviction
* **social:** daily social content pipeline with 5-source drafts + engagement
* **types:** add `document` input modality across SDKs, server, and metadata

### Bug Fixes

* **adapters:** add input validation guards for empty/failed visual inputs
* **adapters:** address review findings for Qwen3-VL adapters
* **adapters:** clarify video placeholder, validate token IDs, fix torch_dtype key
* add client-side hour filter to search_x_posts (was date-level only)
* address CodeRabbit review feedback
* address follow-up PR review nits
* address remaining CodeRabbit feedback (round 2)
* address review findings — negative truncation guard, score() options, constant dedup
* **bench:** show correct unit labels for MP/s throughput in --print-gap
* **bundles:** declare Qwen3-VL adapters in default bundle
* **ci:** use Blacksmith runner in CI
* **client:** retry only transient connection errors under wait_for_capacity
* **cluster:** address PR #701 review comments
* **cluster:** correct kubectl flag combo and reorder LB sweep before helm uninstall
* **cluster:** helm uninstall before terraform destroy to clean up AWS LB leftovers
* **cluster:** unblock end-to-end `mise run cluster create --build`
* **config,cluster:** surface unrouteable models loudly and helm-repo-add on pristine hosts
* **config:** emit identical NATS payload to bundle and _all subjects
* **config:** surface mixed-profile unrouteable models and keep snapshot consistent on writes
* **docker:** add retry logic for deadsnakes PPA to handle Launchpad outages
* **docker:** propagate failure when all add-apt-repository retries exhausted
* **docling:** per-task converter, hf_revision guard, callable typing (CodeRabbit)
* **docs:** Update packages/sie_server/Dockerfile.cuda11
* fail closed on missing/unparsable timestamps in lookback filter
* **gateway,config,sdk:** resiliency, concurrency, and cross-service hash parity
* **gateway,config:** address PR review -- 404 for unknown models, 202 on default routing, full YAML propagation
* **gateway,config:** harden auth, trusted NATS producers, and recovery path; drop gateway HA default
* **gateway,sdk:** map upstream timeouts to 503+MODEL_LOADING for SDK retry
* **gateway:** add GET /v1/models/\{model\} detail route
* **gateway:** address PR #716 review feedback
* **gateway:** align /v1/models error and list shapes
* **gateway:** drop double-counted REQUEST_COUNT / REQUEST_LATENCY emit
* **gateway:** emit X-SIE-Error-Code header on model-loading 503
* **gateway:** keep record_request async to match main's call shape
* **gateway:** make sie-config single source of truth for bundles with live resync
* **gateway:** normalize model ids in NATS work subjects + docs/tooling/ha cleanup
* **gateway:** pre-instantiate request/demand metric families on startup
* **gateway:** prioritize epoch-rewind branch; harden no-thrash test; correct arch-guide on ephemeral restart
* guard score() and score_pairs() against empty input lists
* **helm:** default clusterRouting to "queue" on import-sie-router-rust
* **helm:** enable NATS + JetStream by default to match queue clusterRouting
* **helm:** fail fast when gateway has no bundle source
* **kind-smoke:** add --no-pool-isolation for static clusters + contract-drift fixes
* **kind-smoke:** address bot review feedback
* **kind-smoke:** enable configStore and harden config/gateway tests
* **kind-smoke:** enable JetStream on test NATS and drop duplicate subchart
* **kind-smoke:** wire sie-config image and helm overrides into kind cluster fixture
* **kind-smoke:** wire sie-config image into kind cluster fixture
* **observability:** address PR review blockers on metrics PR
* **sdk:** cluster cache prefix probe uses list, not head
* **sdk:** cluster cache prefix probe uses list, not head (Refs #732, #654)
* **sdk:** has_children filters folder-marker objects (Refs #732, #654)
* **sdk:** preserve caller-supplied document format over inferred (CodeRabbit)
* **sdk:** retry mid-flight transport disconnects, not just timeouts
* **sdk:** retry on connection errors and generic 503s
* **sie_config:** address PR review feedback
* **terraform/aws:** set 100GB root volume on cpu node group to avoid DiskPressure
* **terraform/gcp:** undo router→gateway rename on GCP Cloud Router + NAT
* **tests:** include sie-config in expected missing-image list
* **tests:** restore docker gateway smoke test after router rename
* **tmux-scripts:** improve robustness of session parsing and argument handling
* **types:** adapt to ty 0.0.32 stricter ignore handling
* use searchTerms for X tweet-scraper actor (was searchQueries)

### Performance Improvements

* **bench:** cache JPEG-encoded corpus images across queries
* **bench:** lazily JPEG-encode corpus images on first use
* **docker:** add --link + move ARG BUNDLE to eliminate cross-bundle layer noise
* **docker:** normalize mtimes so shared venv layer is dedupable
* **docker:** reorder stages for maximum BuildKit cache reuse
* **docker:** split worker venv into shared + bundle-specific layers
* **gateway:** cache SDK version parse, integer audit latency, UUIDv7
* **gateway:** cut hot-path allocations, fuse numpy decode, tighten backpressure
* **gateway:** fuse msgpack_numpy decode into the response path
* **gateway:** move score-endpoint unwrap instead of cloning
* **gateway:** pass msgpack items through as rmpv::Value
* **gateway:** publish work items concurrently + borrow shared fields
* **gateway:** tighten cold-pool backpressure + cheaper QPS counter
* **gateway:** trim per-request work on the inference hot path

## v0.2.0 (2026-04-17)

### Highlights

- **Breaking change:** Removed `--model` CLI args from worker startup; use `SIE_PRELOAD_MODELS` env var or `--preload` flag instead
- **New capabilities:** add ModernBERT flash dense embedding support with fallback mechanism; add OCR quality benchmarks (olmOCR-bench); add OCR quality benchmarks with olmOCR-bench; add pages/sec throughput metric for OCR perf eval; add perf metrics to OCR eval pipeline; also report query throughput in mpix/s for image queries
- **Reliability and operations:** add missing NATS Helm repo to release workflow; don't set NODE_AUTH_TOKEN for OIDC npm publishes; harden affinity spill with bounds check, clamp, and debug log; make rejected requests visible to KEDA scaling metrics; remove redundant tokenizer validation and unused template parameter

### ⚠ BREAKING CHANGES

* **workers:** Removed `--model` CLI args from worker startup; use `SIE_PRELOAD_MODELS` env var or `--preload` flag instead

### Features

* **adapters:** add ModernBERT flash dense embedding support with fallback mechanism
* **bench:** add OCR quality benchmarks (olmOCR-bench)
* **bench:** add OCR quality benchmarks with olmOCR-bench
* **bench:** add pages/sec throughput metric for OCR perf eval
* **bench:** add perf metrics to OCR eval pipeline
* **bench:** also report query throughput in mpix/s for image queries
* **benchmarks:** add MTEB NFCorpus evaluation results for ModernBERT-based embedders
* **bench:** report vision corpus throughput in mpix/s
* **bench:** report vision corpus throughput in mpix/s instead of items/s
* **deps:** migrate from pynvml to nvidia-ml-py package
* **haystack:** add haystack_integrations namespace aliases
* **haystack:** add namespace-convention aliases
* **observability:** add anonymous usage telemetry
* **sdk:** add max_concurrency param to SIEAsyncClient to prevent connection pool exhaustion
* **server:** add lightonai/LightOnOCR-2-1B OCR adapter with next bundle
* **workers:** implement model preloading at startup to reduce first-request latency

### Bug Fixes

* **adapters:** remove redundant tokenizer validation and unused template parameter
* address PR review — panel title, namespace variable
* **bench:** handle unloaded images in pixel count computation
* **bench:** use concurrent async requests for OCR perf eval
* **bench:** validate image entries before computing pixel counts
* **bench:** validate pixel counts before using them for image corpus throughput
* **build:** downgrade dockerfile syntax version to 1 for broader compatibility
* **ci:** add missing NATS Helm repo to release workflow
* **ci:** don't set NODE_AUTH_TOKEN for OIDC npm publishes
* **dashboard:** queue routing dashboard accuracy and usability
* **docs:** add update date to portfolio header
* **docs:** correct PR reference in reranker reclassification note
* **docs:** populate reranker data and simplify table header
* **docs:** update stale model counts after reranker reclassification
* **haystack:** rename namespace alias to sie
* install uv via curl instead of COPY --from ghcr.io
* preload smoke test checks model.loaded instead of nonexistent workers field
* **readme:** heading format
* **release:** add LanceDB integrations to release-please config
* **router:** add overflow spill to break model affinity deadlock
* **router:** harden affinity spill with bounds check, clamp, and debug log
* **router:** make rejected requests visible to KEDA scaling metrics
* **sie-bench:** account for in-flight drain in throughput calculation
* **sie-bench:** use union wall-clock for multiprocess throughput merge
* **tester-cluster:** patient KEDA scale-down for worker pools

## v0.1.10 (2026-04-09)

### Highlights

- **New capabilities:** add async, chunking, and streaming to Weaviate document enricher; improve DLQ routing and score response handling; implement Config Management API with NATS-based distribution and review fixes; add LanceDB integration (Python + TypeScript); queue routing dashboard + NATS exporter + router image tag; queue routing dashboard, NATS prom exporter, router image tag
- **Reliability and operations:** correct cluster routing condition, stream max_age units, and reconnect state ordering; add recreate strategy for router deployment when nats config restore is enabled; restore Chart.yaml deps from main, keep appVersion v-prefix; queue routing dashboard PromQL for NATS wait; configurable NATS fetch budget, Helm-wired queue params
- **Performance:** decouple scanner and SIE batch sizes in enrich_table; stream enrich_table batch-by-batch instead of full materialization; use Lance scanner for column projection in enrich_table; bypass FastAPI for hot proxy paths via raw ASGI middleware

### Features

* add async, chunking, and streaming to Weaviate document enricher
* **dlq,pull-loop:** improve DLQ routing and score response handling
* implement Config Management API with NATS-based distribution and review fixes
* **integrations:** add LanceDB integration (Python + TypeScript)
* **observability:** queue routing dashboard + NATS exporter + router image tag
* **observability:** queue routing dashboard, NATS prom exporter, router image tag
* **sdk:** add get_model() and configure LanceDB release workflows
* **terraform:** add AWS eval-eu EKS cluster with multi-GPU support
* **terraform:** add evaluation cluster setup for AWS with multi-GPU support and updated configurations
* **terraform:** add node labels, adjust pool sizes for tester cluster

### Bug Fixes

* **config,queue,nats:** correct cluster routing condition, stream max_age units, and reconnect state ordering
* handle BytesIO images in LlamaIndex and validate Weaviate classify config
* **helm:** add recreate strategy for router deployment when nats config restore is enabled
* **helm:** restore Chart.yaml deps from main, keep appVersion v-prefix
* **helm:** use generic release-please updater for appVersion
* **helm:** use generic updater for both Chart.yaml version fields
* **helm:** use l4-spot/rtx6000-spot naming convention for spot profiles
* **integrations:** address CodeRabbit review findings for LanceDB PR
* **observability:** queue routing dashboard PromQL for NATS wait
* **queue-routing:** configurable NATS fetch budget, Helm-wired queue params
* **queue-routing:** resolve bugs, add configurable NATS params, fix score wire format
* **queue-routing:** score response format and DLQ fallback routing key
* **release:** use NPM_TOKEN for initial sie-lancedb publish
* **router:** use "scores" key in queue-mode score responses
* **terraform:** add GPU subnet coverage validation
* **terraform:** relax AZ validation and clarify defaults
* **terraform:** review fixes for tester cluster infra
* **terraform:** switch tester-cluster to us-east-2
* **terraform:** Switch tester-cluster to us-east-2 and update deployment docs
* **terraform:** validate gpu_node_groups for duplicate and reserved names
* **test:** add buildx builder pause recovery and improve build error diagnostics
* update adapter tests and address code review feedback
* use OCI registry URI for helm chart in README

### Performance Improvements

* **lancedb:** decouple scanner and SIE batch sizes in enrich_table
* **lancedb:** stream enrich_table batch-by-batch instead of full materialization
* **lancedb:** use Lance scanner for column projection in enrich_table
* **router:** bypass FastAPI for hot proxy paths via raw ASGI middleware
* **router:** reduce thread pool pressure by inlining small deserialization
* **router:** remove msgpack_numpy global patch and BaseHTTPMiddleware
* **router:** replace stdlib json with orjson for 3-10x faster serialization
* **sdk+router:** lazy msgpack_numpy.patch and pure ASGI middleware

## v0.1.9 (2026-04-02)

### Highlights

- **Reliability and operations:** increase docker smoke test timeouts and add retry; include $platform in worker image tag format; revert pool names to machine profile names; remove --provenance flag (requires public repo)

### Bug Fixes

* **helm:** include $platform in worker image tag format
* **helm:** revert pool names to machine profile names
* increase docker smoke test timeouts and add retry
* remove --provenance flag (requires public repo)

## v0.1.8 (2026-04-01)

### Highlights

- **Reliability and operations:** add sie-qdrant and sie-weaviate to release-please config; point sync-terraform default repos to production; remove --provenance from npm publish for private repo; correct image.tag comment to reflect actual format; remove duplicate platform suffix from worker image tag

### Bug Fixes

* add sie-qdrant and sie-weaviate to release-please config
* **ci:** point sync-terraform default repos to production
* **ci:** remove --provenance from npm publish for private repo
* **helm:** correct image.tag comment to reflect actual format
* **helm:** remove duplicate platform suffix from worker image tag
* remove internal-only references from COMPATIBILITY.md

## v0.1.7 (2026-04-01)

### Highlights

- **New capabilities:** add profiling script for sparse encoding hot path; add GitHub Actions workflow to sync Terraform modules to registry repos; apply QoL improvements from PR #484 review comments; switch default GPU from g5 (A10G) to g6 (L4); add rerank/score support to TEI runner; implement configurable document length limits and custom prefix token registration
- **Reliability and operations:** restore triggering ref for source checkout; restore quality by enabling causal attention and QK-normalization; restore dev-l4-spot zones to us-central1 for GPU availability; check /metrics endpoint in test_prometheus_metrics_exist; add per-attempt timeout to lease renewal fetch
- **Performance:** optimize MoE expert dispatch with sorted-expert routing; batch MaxSim scoring across documents on GPU; batch sparse aggregation with segment_reduce and fuse relu; batch split_embeddings + validate ColBERT performance

### Features

* **adapters:** add profiling script for sparse encoding hot path
* add GitHub Actions workflow to sync Terraform modules to registry repos
* apply QoL improvements from PR #484 review comments
* **aws:** switch default GPU from g5 (A10G) to g6 (L4)
* **bench:** add rerank/score support to TEI runner
* **colbert:** implement configurable document length limits and custom prefix token registration
* **deploy:** move namespace, SA, and HF token secret management to Helm chart
* **deploy:** prepare Terraform modules for public registry publishing
* **deploy:** rewrite example module sources to registry references
* **deploy:** rewrite Helm and internal references for public release
* **deploy:** two-artifact model — GCP Terraform infra-only, batteries-included Helm chart
* **docker:** add --docker-platform flag to docker build task
* extend create_pool API/SDK with minimum_worker_count and bundle
* **helm:** add batteries-included sub-chart dependencies to sie-cluster
* **helm:** add image pre-pull DaemonSet for GPU worker pools
* **helm:** add step to build Helm chart dependencies in Kind smoke tests
* **helm:** default router to image-embedded model configs
* **helm:** enable image pre-pull DaemonSet by default
* **helm:** port health gates from Terraform to Helm post-install hooks
* **helm:** remove prometheus alias, bump to v0.2.0, standardize chart
* **infra:** add Modal GPU sandbox for remote benchmark execution
* **infra:** add rollout warning and explicit image_type for GCFS
* **infra:** enable GCFS image streaming on GPU node pools
* **infra:** set min_node_count=1 on L4 spot GPU node pools
* **integrations:** add Qdrant integration
* **integrations:** add Qdrant integration with native sparse vector support
* **integrations:** add Weaviate v4 integration with Go module spec
* multiprocess loadtest + SDK aiohttp migration
* **sdk:** add version negotiation headers between SDK and server
* **sdk:** default wait_for_capacity=True and timeout=900s
* **sdk:** version negotiation header (SDK ↔ server)
* **sie-bench:** add dataset/input_type fields for mTEB corpus inputs
* **sie-bench:** built-in multiprocess loadtest mode
* **skills:** add eval-model skill for HF model assessment
* **skills:** add eval-model skill for HF model integration assessment
* sync Terraform modules to registry repos
* **tei-runner:** add /embed_sparse support for sparse models
* **tei:** add /embed_sparse support and auto-detect pooling mode
* **terraform/aws:** restore cluster autoscaler helm release to infra module
* **terraform/aws:** strip k8s resources, restructure as infra module with examples
* **terraform:** add cluster name and artifact registry variables; update node pool configuration
* **terraform:** add EBS CSI driver, NVIDIA device plugin, default StorageClass
* **terraform:** strip gcp k8s/ layer; examples use infra-only module
* **tools:** add ColBERT query vs document profiling script
* **tools:** add dense P50 latency profiling script

### Bug Fixes

* **adapters:** sort IDF unique_ids to satisfy SparseVector contract
* add missing production example to tf validate; fix tempfile leak; remove module docstring
* address PR #478 review feedback
* address PR review — GPU alert formula, kubectl parsing, CI path filter
* address review feedback for npm publish
* address review findings - race prevention, cleanup, lighter checkout
* **alloy:** add stage.cri\{\} before stage.json to unwrap CRI log envelopes
* **alloy:** explicitly set configMap name and key for sub-chart wiring
* **alloy:** scope pod discovery to current node via field selector
* **bench:** complete g5 to g6 migration in AWS eval configs and GPU mapping
* **benchmarks:** use TEI /embed_all for ColBERT multi-vector models
* **bench:** skip loading candidates_model for single-model servers
* **chart:** update home URL and Helm install command in README
* CI compatibility and consistent env var usage
* CI compatibility for sync-terraform workflow
* **ci:** add contents: read permission to publish-pypi-oidc job
* **ci:** add helm repo add + dep build to kind-smoke workflow
* **ci:** g5 refactored to g6 already
* **cluster:** build concrete helm command in status from infra_outputs
* **cluster:** guard helm/kubectl post-create log when outputs are empty
* **deploy:** clean terraform init artifacts before push
* **deploy:** correct smoke test TypedDict access and helm dry-run args
* **deploy:** correct StatefulSet rollout semantics, PDB scope, and KEDA pause
* **deploy:** remove dangling kubernetes_namespace_v1.sie references from health_gates.tf
* **deploy:** restore triggering ref for source checkout
* **deploy:** update default destination repos for GCP and AWS modules
* **deploy:** use triggering ref for source checkout in sync-terraform
* disable LoRA adapter layers after loading to prevent quality corruption
* **docs:** clarify optional image push in AWS and GCP README files
* **docs:** update Helm chart path in AWS and GCP README files
* fix integration test
* **helm,hook:** deploy/helm/sie-cluster/templates/hooks/prometheus-ready-test.yaml
* **helm:** add before-hook-creation to Job delete policies; document count==0 expectation
* **helm:** address coderabbit findings on health gate hooks
* **helm:** address non-blocking review findings from PR #336
* **helm:** address review findings in batteries-included sub-chart PR
* **helm:** address reviewer suggestions for health gate hooks
* **helm:** address second-pass review findings
* **helm:** aggregate buckets by le in p95 latency alert
* **helm:** bump chart version to 0.1.1 (patch, not minor)
* **helm:** clarify prometheusAddress comment — ignored when sub-chart is installed
* **helm:** correct kube-prometheus-stack semver constraint and remove hardcoded grafana password
* **helm:** correct misleading validation comment in router-deployment.yaml
* **helm:** don't emit ScaledObject CRDs unless KEDA is confirmed present
* **helm:** downscope KEDA RBAC to Role/RoleBinding; remove runtime apk installs
* **helm:** fix loki service URL and extract alloy config to file
* **helm:** fix three blocking review issues in sub-chart dependencies
* **helm:** improve temporary values file handling in helm_template function
* **helm:** improve, simplify, and modularize sie-cluster chart
* **helm:** move 'app.kubernetes.io/part-of' label to selector labels for consistency
* **helm:** remove autoscaling.enabled from values-aws.yaml
* **helm:** render KEDA ScaledObjects via post-install hook to avoid CRD chicken-and-egg
* **helm:** replace hardcoded namespace in provisioning alert rules
* **helm:** require non-empty hfToken.value when hfToken.create is true
* **helm:** sub-chart naming, Loki compactor, event exporter ECR, Grafana folders
* **helm:** use autoscaling.prometheusAddress in prometheus hook; remove stub health_gates.tf
* **helm:** use full FQDN for Prometheus service in KEDA and health gates
* **helm:** use router.service.port in NOTES.txt instead of hardcoded 8080
* **infra:** update min_node_count default in top-level GCP module
* normalize SDK version warned-set key to major.minor
* pool error types, add pool/progress test coverage
* **profiling:** add flash variant registry, device validation, top-level import
* **profiling:** sync GPU before tensor timing, move script to tools/
* **profiling:** use in-place relu_ to match production code path
* **qwen3:** restore quality by enabling causal attention and QK-normalization
* **readme:** correct helm chart path
* **readme:** correct helm install command
* **release:** track all package versions via release-please extra-files
* **release:** track TS SDK version.ts via release-please
* replace corrupted bge-m3 NanoFiQA2018 target + set bfloat16 precision
* review items
* **router:** increase pool lease TTL to survive rolling upgrades
* **router:** resolve default pool GPU for scale-up when gpu/pool omitted
* **router:** use effective_pool instead of pool_name for default pool GPU extraction
* **sdk:** defer aiohttp session creation to fix "no running event loop" in SIEAsyncClient
* **sie_bench:** improve --print-gap report accuracy and readability
* **tei-runner:** validate /embed_all returns per-token embeddings
* **tei-runner:** validate output_type in TEIRunner init
* **terraform/aws:** add full -backend-config flags to production init command
* **terraform/aws:** add precondition asserting &gt;=2 GPU-capable AZs exist
* **terraform/aws:** address review findings post-restructure
* **terraform/aws:** correct helm chart path in dev-g5-spot example comment
* **terraform/aws:** filter VPC AZs to only zones offering the GPU instance type
* **terraform/aws:** fix invalid splat on instance type offerings locations
* **terraform/aws:** remove provider aws block from child module
* **terraform/aws:** use var.project_name in VPC subnet cluster tags
* **terraform:** add validation for GPU node pool zones to ensure they match the configured region
* **terraform:** restore dev-l4-spot zones to us-central1 for GPU availability
* **terraform:** update GPU instance type description for clarity and add dev-g6-spot example
* **terraform:** update stale k8s module references in comments
* **terraform:** upgrade AWS modules and fix deprecations
* **test:** check /metrics endpoint in test_prometheus_metrics_exist
* **test:** update EKS tests from g5 to g6 after GPU instance type change
* **ts-sdk:** add per-attempt timeout to lease renewal fetch
* use prepack instead of prepublishOnly
* validate minimum_worker_count input and soften docstrings

### Performance Improvements

* **adapter:** optimize MoE expert dispatch with sorted-expert routing
* **adapters:** batch MaxSim scoring across documents on GPU
* **adapters:** batch sparse aggregation with segment_reduce and fuse relu
* **adapters:** batch split_embeddings + validate ColBERT performance
* **adapters:** batch split_embeddings in ColBERT adapters
* **adapters:** eliminate GPU overhead from IDF query encode path
* **florence2:** greedy decoding for OCR (-23% P50)
* **florence2:** switch OCR configs from beam search to greedy decoding
* **server,bench:** add batch coalescing, query warmup, and benchmark stability improvements
* **server:** dispatch immediately when worker is idle
* **server:** optimize BertFlashAdapter inference path (+35% corpus throughput)
* **server:** reduce batch wait timeout 10ms \u2192 2ms for lower Doc P50

## v0.1.6 (2026-03-12)

### Highlights

- **Breaking change:** remove `florence2` and `gliner` standalone bundles — extraction adapters (gliner, glirel, gliclass) are now included in the `default` bundle
- **New capabilities:** add native MTEB reranking task support with MRR metric; encode-dense matrix eval — 3 models × 8 tasks; add date-prefixed versioning for chronological filename ordering; reorganize and expand model size lookup table with alphabetical ordering; add detailed perf metrics, metric filter, and threshold selector; add marimo benchmark dashboard notebook
- **Reliability and operations:** restore /var/cache/apt mounts, keep /var/lib/apt removed; add HF_TOKEN auth and config kwargs and fix stella models; add dense projection support to Qwen2FlashAdapter; apply query_template from runtime options in SentenceTransformerDenseAdapter; replace pip install with uv add in docker error messages
- **Performance:** vectorize GTE sparse encode path; vectorize tokenization and packing for gte-multilingual-base; vectorize tokenization and packing to reduce throughput gap; switch Qwen/GTE models to flash attention adapter

### Breaking Changes

* **bundles:** remove `florence2` and `gliner` standalone bundles — extraction adapters (gliner, glirel, gliclass) are now included in the `default` bundle

### Features

* **bench:** add native MTEB reranking task support with MRR metric
* **bench:** encode-dense matrix eval — 3 models × 8 tasks
* **benchmarks:** add date-prefixed versioning for chronological filename ordering
* **benchmarks:** reorganize and expand model size lookup table with alphabetical ordering
* **benchview:** add detailed perf metrics, metric filter, and threshold selector
* **benchview:** add marimo benchmark dashboard notebook
* **benchview:** add perf metric selector to Model Size tab
* **benchview:** detailed perf metrics, metric filter, threshold selector
* **router,bench,sdk:** improve throughput with inflight tracking, batching, and connection pooling
* **server:** typed request parsing with msgspec
* **sie_server:** add gliner, glirel, and gliclass extraction dependencies to the default bundle

### Bug Fixes

* **adapter:** add dense projection support to Qwen2FlashAdapter
* **adapter:** apply query_template from runtime options in SentenceTransformerDenseAdapter
* apply CodeRabbit auto-fixes
* **bench:** replace pip install with uv add in docker error messages
* **benchview:** add missing statistics import and use _median helper
* **bundles:** include sglang bundle in default cluster and eval-matrix configs
* **client:** update websocket header parameter name from extra_headers to additional_headers
* **colbert:** enable native mode fallback for non-CUDA devices and add Matryoshka truncation
* **deps:** cap timm upper bound and fix lazy handler init
* **docker:** clear stale apt lists before update to prevent 404s
* **docker:** remove all apt cache mounts from Dockerfiles
* **docker:** remove no-op /var/lib/apt cache mount from apt RUN blocks
* **docker:** restore /var/cache/apt mounts, keep /var/lib/apt removed
* **helm:** increase CPU worker pool memory limits for expanded default bundle
* **model:** add missing query_template to stella_en_400M_v5
* **models:** switch all-MiniLM-L6-v2 to SentenceTransformerDenseAdapter
* **multilingual-e5-large-instruct:** use instruct query template, NFCorpus 0.3521 → 0.3567
* replace invalid HTML entities in SVG with XML numeric entities
* **rope_flash:** clear cached _rope_dummy on unload and use torch.cat for packing
* **router:** also resolve pool-derived GPU names to spot variants
* **router:** resolve bare GPU types to spot variants for KEDA scaling
* **sdk:** resolve sync/async client inconsistencies in score() and encode()
* **server:** centralize request validation to prevent 500s from malformed items
* **server:** use BertFlashAdapter for e5-small-v2, resolve e5 perf anomalies
* **server:** use BertFlashAdapter for intfloat/e5-small-v2 and remove stale benchmarks
* set compute_precision to bfloat16 for stella_en_1.5B_v5
* **sie_server:** add HF_TOKEN auth and config kwargs and fix stella models
* **sie_server:** resolve BGE-M3 linear weights loading for HF model IDs and fix test fixtures
* **sie_server:** support NV-Embed-v2 with PyTorch embedding adapter
* **splade:** align special token filtering and guard empty batches
* **test:** always rebuild Docker images to pick up code changes
* **typecheck:** move ty type checker from mise tool to uv dependency

### Performance Improvements

* **adapters:** vectorize GTE sparse encode path
* **rope_flash:** vectorize tokenization and packing for gte-multilingual-base
* **rope_flash:** vectorize tokenization and packing to reduce throughput gap
* **server:** switch Qwen/GTE models to flash attention adapter
* **splade:** vectorize tokenization and sparse aggregation (1.5x throughput)
* **splade:** vectorize tokenization and sparse aggregation in SPLADEFlashAdapter

## v0.1.5 (2026-02-27)

### Highlights

- **New capabilities:** add GLiNER v2.5 model configs; stream request bodies through proxy instead of buffering; add classification model configs for GLiClass-large and cross-encoder NLI
- **Reliability and operations:** release pipeline cache collision and smoke test timeout; strip content-length header from streamed proxy responses
- **Performance:** stream response body to eliminate bytes.join bottleneck

### Features

* **models:** add GLiNER v2.5 model configs
* **router:** stream request bodies through proxy instead of buffering
* **sie_server:** add classification model configs for GLiClass-large and cross-encoder NLI

### Bug Fixes

* release pipeline cache collision and smoke test timeout
* **router:** strip content-length header from streamed proxy responses

### Performance Improvements

* **router:** stream response body to eliminate bytes.join bottleneck

## v0.1.4 (2026-02-27)

### Highlights

- **Reliability and operations:** revert sharing=locked, add cache-read-only for build step

### Bug Fixes

* revert sharing=locked, add cache-read-only for build step

## v0.1.3 (2026-02-26)

### Highlights

- **Reliability and operations:** revert token lifetime extension, re-auth before push instead; revert token lifetime, re-auth before push

### Bug Fixes

* revert token lifetime extension, re-auth before push instead
* revert token lifetime, re-auth before push

## v0.1.2 (2026-02-26)

### Highlights

- **Reliability and operations:** release image builds failing from GCP token expiry

### Bug Fixes

* release image builds failing from GCP token expiry

## v0.1.1 (2026-02-26)

### Highlights

- **Reliability and operations:** update bundle definitions to replace legacy and gte-qwen2 with gliner

### Bug Fixes

* **bundles:** update bundle definitions to replace legacy and gte-qwen2 with gliner

## v0.1.0 (2026-02-26)

### Highlights

- **Breaking change:** HTTP 409 dependency conflict responses are removed from all API endpoints; the DEPENDENCY_CONFLICT error code no longer exists; .beads/ issue tracking data removed from repository
- **New capabilities:** add X-SIE-Worker response header for per-worker metrics tracking; add encode-image-text measurements to benchmarks dir; add encode-multivector perf measurements; add encode-multivector performance measurements; add encode-visual-document perf measurements; add encode-visual-document performance measurements
- **Reliability and operations:** increase helm install timeout from 10m to 15m; add trailing empty line to gitignore; align release-images workflow with docker task flags; register GLiClass and DeBERTa models in bundles; build and deploy gliner bundle in Kind smoke tests
- **Performance:** add connection pooling load test results (Feb 24); pool httpx client and add X-SIE-Worker header in router proxy; pool httpx client in router proxy to eliminate per-request TCP overhead; move transformers imports to module level

### ⚠ BREAKING CHANGES

* **deps:** HTTP 409 dependency conflict responses are removed from all API endpoints; the DEPENDENCY_CONFLICT error code no longer exists
* .beads/ issue tracking data removed from repository
* **deps:** model config files no longer support the `dependencies` field

### Features

* add X-SIE-Worker response header for per-worker metrics tracking
* **benchmarks:** add encode-image-text measurements to benchmarks dir
* **benchmarks:** add encode-multivector perf measurements
* **benchmarks:** add encode-multivector performance measurements
* **benchmarks:** add encode-visual-document perf measurements
* **benchmarks:** add encode-visual-document performance measurements
* **benchmarks:** add extract-detection L4-SPOT performance measurements
* **benchmarks:** add extract-kie-docvqa measurements to benchmarks dir
* **benchmarks:** add extract-relation L4-SPOT performance measurement
* **benchmarks:** add score-colbert perf measurements
* **benchmarks:** add score-colbert performance measurements
* **models:** add encode-image-text measurements
* **models:** add extract-detection measurements
* **models:** add extract-kie-docvqa measurements
* **models:** add extract-relation measurements
* **router:** add structured audit logging for API requests

### Bug Fixes

* **.claude:** add trailing empty line to gitignore
* align release-images workflow with docker task flags
* **bundles:** register GLiClass and DeBERTa models in bundles
* **ci:** build and deploy gliner bundle in Kind smoke tests
* **colbert:** remove CUDA requirement and improve device compatibility
* **eval:** read 'sie_id' instead of 'name' from model configs in runner
* **extract:** use dict access for Entity TypedDict in sort
* **gliner:** relax stale transformers&lt;4.52 pin
* increase helm install timeout from 10m to 15m
* reduce cpu-gliner resource requests for Kind CI
* **router:** read 'sie_id' instead of 'name' from model configs
* **server:** migrate NLI adapter to classifications and improve API consistency
* **server:** migrate nli_classification adapter and improve type annotations
* **server:** populate classifications instead of entities in GLiClass adapter
* use manifest mode for release-please and reset to v0.0.0
* use nested .gitignore for .claude/ directory

### Performance Improvements

* add connection pooling load test results (Feb 24)
* pool httpx client and add X-SIE-Worker header in router proxy
* pool httpx client in router proxy to eliminate per-request TCP overhead
* **pytorch-embedding:** move transformers imports to module level
* **server:** use uvloop as default event loop for uvicorn

### Reverts

* keep CONTRIBUTING.md clone URLs pointing to sie.git

### Miscellaneous Chores

* remove beads, agent prompts, mypy refs; consolidate ty config

### Code Refactoring

* **deps:** move adapter dependencies from per-adapter pyproject.toml to bundle YAML
* **deps:** remove model-level dependencies feature
