Release Notes
Latest version: v0.7.1 (2026-08-09).
v0.7.1 (2026-08-09)
Section titled “v0.7.1 (2026-08-09)”Highlights
Section titled “Highlights”- New capabilities: serve docling OCR with the English recogniser
- Reliability and operations: meter executed GLiNER token window
Features
Section titled “Features”- server: serve docling OCR with the English recogniser
Bug Fixes
Section titled “Bug Fixes”- server: meter executed GLiNER token window
v0.7.0 (2026-08-08)
Section titled “v0.7.0 (2026-08-08)”Highlights
Section titled “Highlights”- Breaking change: aws-control-plane no longer accepts component_contract_sha256 or injects its environment marker.
- New capabilities: add fail-closed estate-bootstrap saved-plan inspection; add frozen assembly apply seam; add minimal promotion eligibility guard; admit estates by declared capability and add the sandbox desired state; allow benchmark campaign self-review; assert apply-free and same-scope second convergence runs
- Reliability and operations: align catalog smoke timeout budget; apply the checkpoint whitelist to record, harden the git guard; bind the published price table to the active billing authority; harden document field fixture validation; harden generation response handling
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- terraform: aws-control-plane no longer accepts component_contract_sha256 or injects its environment marker.
Features
Section titled “Features”- cloud: add fail-closed estate-bootstrap saved-plan inspection
- cloud: add frozen assembly apply seam
- cloud: add minimal promotion eligibility guard
- cloud: admit estates by declared capability and add the sandbox desired state
- cloud: allow benchmark campaign self-review
- cloud: assert apply-free and same-scope second convergence runs
- cloud: batch the priced rate-grade cells behind one approval
- cloud: bill audio_ms at a 1,150 ms minimum duration
- cloud: bind the inspection verdict to the exact plan digest
- cloud: declare the terraform backend kind per estate cell
- cloud: deepen the read-only estate bootstrap preflight
- cloud: default tier points to min(floor x 2.0, ceiling x 0.85)
- cloud: freeze infrastructure values in assemblies
- cloud: hydrate terraform outputs for remote-backend cells
- cloud: implement the #2441 pricing-policy decisions (PR #3024)
- cloud: local-only session auth for the L3 verification stack
- cloud: make the sandbox cell serving
- cloud: permanently allow benchmark campaign self-review
- cloud: persist deliverable metadata in s3
- cloud: publish the live rate book to a public price artifact
- cloud: record and verify estate bootstrap approvals
- cloud: record convergence evidence and compare two runs
- cloud: warn loudly when a —live preflight run proves nothing
- control-plane: serve-image capability manifest and operator-opened tasks
- model-ops: add the docling English OCR artifact publish tool
Bug Fixes
Section titled “Bug Fixes”- ci: refuse an unusable assets_path before the gate
- ci: replace the batch ledger atomically and unwrap the summary prose
- clients: address terminal response review
- cloud: accept padded whisper transcripts
- cloud: accept real names, escape paths, close the commondir fail-open
- cloud: address deployment review findings
- cloud: admit provisional task measurement pairs
- cloud: admit same-apply IAM unknowns by provenance, not by fiat
- cloud: align catalog canary model contracts
- cloud: align catalog smoke timeout budget
- cloud: align document field canary semantics
- cloud: align phase 3 exit criteria and make skip tests prove their cause
- cloud: align PII acceptance taxonomy
- cloud: allow task candidates across routes
- cloud: allowlist read-only init by its full argument set
- cloud: apply the checkpoint whitelist to record, harden the git guard
- cloud: assert the full run built something and order the pair
- cloud: bind every resource to the verified provider, not only data sources
- cloud: bind read-only terraform access to the resolved cell’s root
- cloud: bind the published price table to the active billing authority
- cloud: bound Modal image build cache
- cloud: bound provenance to policy CONTENT, not just mentioned resources
- cloud: close checkpoint leak, env-independent repo guard, strict schema
- cloud: close fail-open branch chains in the estate preflight
- cloud: close fail-open paths in estate plan inspection
- cloud: close preflight fail-opens found by the mandated self-review
- cloud: close the fail-open paths found in review threads
- cloud: close the fail-opens the third self-review found
- cloud: close the holes the fourth self-review found in this round’s own permissive paths
- cloud: close the plan-inspection bypasses from the adversarial review
- cloud: close the two backend-classifier evasions from verification
- cloud: commit job results before settlement
- cloud: confine catalog subsets to sandbox
- cloud: constrain checkpoint field values, not just field names
- cloud: constrain screenshot bbox axes
- cloud: correct the apply-check rationale and scope-check claim
- cloud: correct the offload-smoke minted tuple annotation and notice an ignored —account
- cloud: correct the provider-key rule against ground truth, never echo after values
- cloud: declare the revoke 404 contract and pin its no-oracle body
- cloud: decouple availability from measurements
- cloud: defer gateway replica probe until serve
- cloud: derive the declared commercial margin from the price multiplier
- cloud: enforce the D1 boundary on the terraform primitive
- cloud: escape the last raw path, refuse an unreadable HEAD
- cloud: fail closed on an unreadable desired-state spec
- cloud: freeze deployment check results
- cloud: fullmatch the rate-book version before it becomes a filename
- cloud: give every checkpoint field a format, not a length cap
- cloud: guard the path parameters, pin refusals, drop the bare except
- cloud: harden document field fixture validation
- cloud: harden generation response handling
- cloud: harden the estate preflight probes per security review
- cloud: harden the point default and the artifact publish tool
- cloud: harden the preflight per the CodeRabbit review threads
- cloud: hold the D1 remote-backend boundary in live preflight too
- cloud: hold the priced generation window’s prefill mix to a review
- cloud: isolate deployment plan inputs
- cloud: keep production catalog topology strict
- cloud: keep the pricing-authority resolver import-safe off a checkout
- cloud: make shape handling total and stop rejecting the honest plan
- cloud: make the unpriced-storage zero loud
- cloud: parse s3 backend blocks brace-aware so nested blocks are read
- cloud: pin the AWS-edge coupling and refuse malformed input everywhere
- cloud: read not-owned terraform roots through a narrow allowlist
- cloud: read trust_remote_code from both adapter scopes everywhere
- cloud: recover externally cancelled live gateway candidate
- cloud: refuse impossible instants and stop the scope string overclaiming
- cloud: register migration 0079 and name the book on a merged sweep
- cloud: reject a run paired with itself and fix the stale exit bullet
- cloud: reject lineage cycles and parented full runs
- cloud: reject unusable SSM pagination tokens; pin cross-estate isolation
- cloud: repair chat catalog canary
- cloud: repair document field canary semantics
- cloud: require exact screenshot labels
- cloud: require statement constancy per security-relevant field
- cloud: restore cross-estate rejection on excluded entries, refuse composition
- cloud: restrict the instant pattern to ASCII digits
- cloud: retry transient job reads
- cloud: scan every leaf for foreign-estate references, not just name keys
- cloud: scope key mutations to an account the caller names
- cloud: snapshot untracked development roots
- cloud: stabilize catalog contract canary
- cloud: stabilize screenshot acceptance fixture
- cloud: stabilize screenshot contract canary
- cloud: stabilize whisper catalog canary
- cloud: stop arming gateway proxy auth without a public edge
- cloud: stop publishing SIE’s markup to customers
- cloud: tighten promotion evidence validation
- cloud: tolerate EC2 worker visibility lag
- cloud: track %{} template-directive spans in the backend classifier
- cloud: track template interpolations in the backend classifier
- cloud: unpin gateway path dependency
- cloud: validate mode, harden provider-key parsing, make two tests honest
- cloud: validate read entries and close self-review fail-opens
- cloud: validate the estate id where it enters, tighten approver
- cloud: verify materialized candidate bytes
- deps: provide the dependencies the bundle and the eval already declare
- model-ops: pin trusted digests and refuse unrecorded provenance
- release: unpin sidecar path dependency
- sdk: refresh stale job result refs
- server: echo queued extract item ids
- server: normalize Grounding DINO instructions
- server: preserve pinned models in serve config
- server: validate pinned model selection
- tools: allow SDK 0.7 releases
Code Refactoring
Section titled “Code Refactoring”- terraform: drop control-plane contract marker
v0.6.30 (2026-08-07)
Section titled “v0.6.30 (2026-08-07)”Highlights
Section titled “Highlights”- New capabilities: compile the deployment matrix from the estate registry
- Reliability and operations: drop the invalid queue key from the staging concurrency blocks
- Performance: accelerate Qwen thinking inference
Features
Section titled “Features”- cloud: compile the deployment matrix from the estate registry
Bug Fixes
Section titled “Bug Fixes”- ci: drop the invalid queue key from the staging concurrency blocks
Performance Improvements
Section titled “Performance Improvements”- models: accelerate Qwen thinking inference
v0.6.29 (2026-08-07)
Section titled “v0.6.29 (2026-08-07)”Highlights
Section titled “Highlights”- New capabilities: bind evidence and acceptance to estate identity; optimize Gemma thinking profiles
- Reliability and operations: preserve release tag during public sync verification; verify the acceptance OIDC identity token
Features
Section titled “Features”- cloud: bind evidence and acceptance to estate identity
- server: optimize Gemma thinking profiles
Bug Fixes
Section titled “Bug Fixes”- ci: preserve release tag during public sync verification
- cloud: verify the acceptance OIDC identity token
v0.6.28 (2026-08-07)
Section titled “v0.6.28 (2026-08-07)”Highlights
Section titled “Highlights”- New capabilities: derive resource names and paths from estate coordinates; gate public edge and vendor identity on estate capability; split managed gates into tier and estate capability; optimize non-thinking generation profiles
- Reliability and operations: classify entity canary failures; gate Stripe preflight before any provider request; keep vendor provider access tier-gated; refresh siglip diagnostic projection; remove stale console gateway URL and reject tampered cells
Features
Section titled “Features”- cloud: derive resource names and paths from estate coordinates
- cloud: gate public edge and vendor identity on estate capability
- cloud: split managed gates into tier and estate capability
- server: optimize non-thinking generation profiles
Bug Fixes
Section titled “Bug Fixes”- cloud: classify entity canary failures
- cloud: gate Stripe preflight before any provider request
- cloud: keep vendor provider access tier-gated
- cloud: refresh siglip diagnostic projection
- cloud: remove stale console gateway URL and reject tampered cells
- cloud: scope Modal denial floor to serving estates
- cloud: update gateway profile fixtures
- quality: extend PubTabNet task budget
- release: extract refresh artifact by id
- release: keep diagnostic suites out of git
- release: preserve refresh artifact across retries
- server: pin speculative draft launches
- server: stage drafts for local models
v0.6.27 (2026-08-06)
Section titled “v0.6.27 (2026-08-06)”Highlights
Section titled “Highlights”- New capabilities: add staging comparison diagnostics; pin benchmark campaigns to a verified main revision; add connector plan controls; expose recovery-bound connector repair; activate atomic connector run; activate explicit connector execute
- Reliability and operations: authenticate quality gate attempt markers; authenticate req14 quality gate markers; pin void authorization Python; bind retained recovery to candidate tasks; bind vendor mutations to estate authority
Features
Section titled “Features”- bench: add staging comparison diagnostics
- ci: pin benchmark campaigns to a verified main revision
- clients: add connector plan controls
- clients: expose recovery-bound connector repair
- cloud: activate atomic connector run
- cloud: activate explicit connector execute
- cloud: activate governed postgres connectors
- cloud: activate recovery-bound connector repair
- cloud: add atomic connector run admission
- cloud: add bounded connector recovery probes
- cloud: add bounded postgres source planning
- cloud: add component convergence contracts
- cloud: add connector execution journal client
- cloud: add connector plan guardrails
- cloud: add connector plan schema policy
- cloud: add crash-safe connector outer runner
- cloud: add direct connector inference
- cloud: admit executable connector plans
- cloud: attest connector workers before payload
- cloud: band the size-class tier book on (operation, parameters)
- cloud: band the tier book on (operation, parameters)
- cloud: bind connector executor dispatch
- cloud: bind connector plans to worker identity
- cloud: bind direct connector executions
- cloud: bound postgres connector execution context
- cloud: bridge durable connector execution dispatch
- cloud: commit the control-plane OpenAPI contract with a diff gate
- cloud: compile estate vendor coordinates
- cloud: compile optional estate vendor authority
- cloud: compose connector recovery journal
- cloud: compose dormant connector execution
- cloud: enable native console previews
- cloud: enable selective component deployment
- cloud: expose connector execution transitions
- cloud: expose durable connector plans
- cloud: fail closed on unconfigured vendor authority
- cloud: fence postgres connector publication
- cloud: govern connector plan manifests
- cloud: isolate SigLIP rate-grade diagnostics
- cloud: journal connector recovery phases
- cloud: land governed Postgres connector runtime
- cloud: persist connector batch lane authority
- cloud: persist connector execution authority
- cloud: persist durable connector plans
- cloud: preflight connector worker discovery
- cloud: prepare and fence connector execution
- cloud: prove executable postgres preflight
- cloud: provision connector runtime authorities
- cloud: provision isolated console previews
- cloud: reconcile connector publication recovery
- cloud: record the owner’s acceptance of the declared exposure
- cloud: replace catalog DSL with task implementations
- cloud: replay governed postgres manifests
- cloud: resolve estate coordinates at the provisioner boundary
- cloud: resolve terraform coordinates from the estate registry
- cloud: route generation through local ingest
- cloud: route managed generation through local ingest
- cloud: scaffold additional staging estate
- cloud: scaffold sandbox foundation
- cloud: settle connector staging before publication
- cloud: unify Modal environment authority in the estate registry
- gateway: expose immutable execution evidence
- generation: add 256k profile candidates
- generation: add hardware-specific context profiles
- generation: add strict Responses API parity
- model-ops: allowlist lightonai and OpenSearch-AI as trusted publishers
- model-ops: capture nested multimodal config scalars in onboarding evidence
- model-ops: enable multimodal image-text and document-retrieval discovery families
- model-ops: enable sparse-retrieval discovery family
- model-ops: re-task late-interaction discovery onto SCIDOCS retrieval evidence
- model-ops: trust JinaAI discovery publisher
- models: onboard lightonai/mLateOn
- models: onboard lightonai/mLateOn (scaffolded)
- models: seed mLateOn quality floors from governed measurement
- server: add bounded local-ingest generation client
- server: complete generative model serving
- sidecar: stream generation over local ingest
- terraform: register shared Scalr modules
Bug Fixes
Section titled “Bug Fixes”- bench: validate release evidence ingress early
- build: relock every standalone project that depends on sie-server
- ci: align standalone validation artifacts
- ci: authenticate quality gate attempt markers
- ci: authenticate req14 quality gate markers
- ci: clarify invalid model gate error
- ci: exempt registry library modules from the estate root list
- ci: fail closed for unlisted registry module roots
- ci: format campaign source tests and state the lowercase SHA rule
- ci: isolate connector verifier dependencies
- ci: pin void authorization Python
- ci: trust quality gate attempt cap source
- ci: validate quality gate void reasons
- clients: send connector repair idempotency header
- cloud: accept launch catalog schema v2
- cloud: accept operation-valid usage supersets
- cloud: address connector runtime review findings
- cloud: address Gateway identity review
- cloud: address staging rollback review
- cloud: align estate module contracts
- cloud: align preview posture lint config
- cloud: allow console preview bootstrap
- cloud: allow verified gateway final reuse
- cloud: batch the token-capped encode lanes and price two more launch cells
- cloud: bind connector intent to canonical worker
- cloud: bind every priced-window factor to non-caller evidence
- cloud: bind generation lanes to catalog hash
- cloud: bind retained benchmark owners
- cloud: bind retained recovery to candidate tasks
- cloud: bind task evidence to reviewed cases
- cloud: bind the billed report period to the priced window
- cloud: bind vendor mutations to estate authority
- cloud: bound Modal history reads separately
- cloud: bound Modal inventory reads separately
- cloud: close connector activation review gaps
- cloud: close connector review findings
- cloud: close connector review gaps
- cloud: close two vacuous guards in the tier banding logic
- cloud: declare the below-cost classes instead of hiding them
- cloud: default estate_id inside resolve_estate
- cloud: deny unused staging scheduler authority
- cloud: derive allocation shares from measured work, not fiat
- cloud: derive offered load from measured lane capacity
- cloud: disable staging cold-start reporting
- cloud: drop the exporter’s path workaround, annotate the document type
- cloud: enforce console preview posture
- cloud: fence selective Gateway dispatch by contract
- cloud: finish connector protocol review sweep
- cloud: fold the 2026-08-04 lane sweeps into the offered-load plan
- cloud: forward selective deployment mode
- cloud: gate connector activation on desired state
- cloud: handle local generation stream failures
- cloud: harden connector activation boundaries
- cloud: harden deployment mode forwarding
- cloud: harden edge backend ordering, cell parity, bucket names
- cloud: harden managed generation streaming
- cloud: harden Modal inventory validation
- cloud: harden preview OIDC validation
- cloud: harden registry identity validation and lock estate identity
- cloud: harden selective deployment reuse
- cloud: harden staging estate isolation
- cloud: harden staging foundation identity
- cloud: ignore local source caches
- cloud: isolate console preview git config
- cloud: isolate sidecar subprocess secrets
- cloud: isolate staging estate authorities
- cloud: make Vercel readback authoritative
- cloud: measure the enlarged fixtures and price 18 more launch cells
- cloud: merge main and add band-6 rerank capacities
- cloud: normalize component source modes
- cloud: omit read-only webhook rules
- cloud: omit response-only webhook rules
- cloud: order config image local layers
- cloud: pin gateway plan evidence
- cloud: pin sandbox Scalr modules at their first healthy versions
- cloud: preserve connector execution topology
- cloud: preserve positive sandbox activation target
- cloud: preserve retained worker provenance in release gates
- cloud: preserve valid rate book evidence
- cloud: price the measured window, not the billed report hour
- cloud: re-render siglip diagnostics lock after server source change
- cloud: reconcile connector activation with main
- cloud: record the measured decode length on the generation cells
- cloud: recover native production gateway
- cloud: recover retained candidates after newer merges
- cloud: refresh combined topology evidence
- cloud: refresh siglip diagnostic lock
- cloud: refresh task contract test fixtures
- cloud: refuse an unpublished class name instead of raising KeyError
- cloud: reject empty connector inference text
- cloud: reject JSON Terraform escapes
- cloud: remove unused type suppression
- cloud: report rejected reuse evidence
- cloud: require integer estate schema version
- cloud: reserve governed page ceilings
- cloud: restrict promotion authority to full estates
- cloud: retain legacy standalone tag gate
- cloud: scope promotion authority to estate cells
- cloud: stabilize component binary identity
- cloud: stabilize named-entity canary fixture
- cloud: state the audio_ms clip-length assumption in the quote
- cloud: tighten connector manifest cleanup
- cloud: tighten task contract evidence gates
- cloud: unblock reproducible Modal convergence
- cloud: use explicit connector protocol stubs
- cloud: use non-ambiguous visual distractor
- cloud: use representative vision smoke fixture
- cloud: validate retained migration without activation
- cloud: validate reused worker runtime CAS
- cloud: validate selected gateway in catalog smoke
- cloud: validate settled acceptance usage
- cloud: verify Modal service-user access
- colbert: honor transformers-5 rope_parameters when serving on the 4.x bundle
- generation: align direct backend output limits
- generation: defer resolved profile output limits
- generation: fit Gemma 256k KV cache on H100
- generation: handle family-specific reasoning boundaries
- generation: hide split Gemma thinking boundary
- generation: pair long-context qwen image bounds
- generation: preserve Gemma reasoning boundaries
- generation: preserve qwen vision bounds
- generation: reject overflowing runtime numerics
- generation: reserve Gemma long-context workspace
- generation: unblock cold model startup
- generation: use compatible Qwen 27B grammar backend
- generation: use compatible Qwen 35B grammar backend
- generation: validate high-context thinking profiles
- modal: isolate overlay FlashInfer cache
- modal: keep image yaml loading local
- modal: support strict GPU allocation
- modal: validate strict GPU selectors
- model-ops: an unmeasured expected_dim never confirms a multivector smoke
- model-ops: batch discovery evidence revision listing via one pinned tree fetch
- model-ops: bind trusted terminal outcomes
- model-ops: bound authority listing by model directories, not raw entries
- model-ops: calibrate late-interaction and doc-retrieval scanner blocks
- model-ops: close scaffold authorization gaps
- model-ops: complete scaffolds before Stage 4
- model-ops: exempt tooling stops from attempt cap
- model-ops: fail closed on duplicate search errors
- model-ops: harden scaffold validation boundaries
- model-ops: ignore fenced tooling markers
- model-ops: insert matrix rows once per models node so alias blocks ride their anchor
- model-ops: isolate scaffold candidate tests
- model-ops: namespace the config projection_dim claim and cap consumer-side nested ints
- model-ops: order attempt comments deterministically
- model-ops: preserve adapter auto authorization
- model-ops: scope duplicate onboarding guard
- model-ops: teach the Stage-4 smoke to probe multivector models
- models: complete mLateOn serving recipe
- quality-eval: harden runner repair
- quality-eval: quarantine unregistered runners
- quality-eval: shorten wake grace and map sglang bundles
- quality-eval: shorten wake grace and support sglang logical bundles
- quality-eval: use prepared scorer environment
- release: bound commit history batches
- release: deduplicate generated changelog notes
- release: finalize source before diagnostics
- release: refresh topology-bound diagnostics
- release: reject repeated release headings
- sdk: retry pre-execution SSE capacity errors
- server: align Gemma xgrammar dependency
- server: await standalone smoke teardown
- server: bound Qwen visual context
- server: constrain SGLang compatibility paths
- server: enforce Qwen fast image bounds
- server: keep Qwen3-VL reranker batches working on transformers 5.3
- server: pin docling to its artifact revision and stop OCR gating load
- server: preprocess native extract images
- server: preserve LightOn SGLang compatibility
- server: replace stale root playground with redirect to Swagger docs
- sidecar: harden local generation cancellation
- sidecar: reserve cancel request bytes
- terraform: address internal module review
- terraform: bound internal module lifecycle
- terraform: exclude worktree caches from remote config uploads
- terraform: harden Scalr module releases
- terraform: package internal Scalr module archive root
- terraform: package private Scalr registry modules
- terraform: package Scalr module archive envelopes
- terraform: register Scalr modules from plain monorepo roots
- terraform: reject null module registry provider
- terraform: relocate Scalr registry modules under deploy tree
- terraform: validate sandbox module with OpenTofu
v0.6.26 (2026-08-02)
Section titled “v0.6.26 (2026-08-02)”Highlights
Section titled “Highlights”- New capabilities: add H100 FP8 generation profiles; governed first-baseline mode + 16 vanilla floors for Qwen3-Embedding-8B and nomic v2 MoE
- Reliability and operations: add protected retained recovery dispatch; bind modal env before retained recovery; bind retained gateway recovery to failed release; bound retained recovery runtime; clarify retained gateway recovery
Features
Section titled “Features”- model-catalog: add H100 FP8 generation profiles
- quality-eval: governed first-baseline mode + 16 vanilla floors for Qwen3-Embedding-8B and nomic v2 MoE
Bug Fixes
Section titled “Bug Fixes”- bench: initialize buildx for worker image
- bench: quiesce gc during timed campaigns
- bench: validate non-root worker runtime
- cloud: add protected retained recovery dispatch
- cloud: address provision review findings
- cloud: align rollback preflight diagnostic
- cloud: allow full cold dispatch preflight budget
- cloud: audit extracted provision secrets
- cloud: bind modal env before retained recovery
- cloud: bind retained gateway recovery to failed release
- cloud: bind retained gateway to failed release
- cloud: bound retained recovery runtime
- cloud: clarify retained gateway recovery
- cloud: close provision review gaps
- cloud: converge config before cold gateway
- cloud: fence initial gateway absence recovery
- cloud: govern reconciliation cadence in releases
- cloud: guard bootstrap and benchmark worker runtime
- cloud: guard current-schema bootstrap seed
- cloud: harden connector smoke recovery
- cloud: harden extracted provision edges
- cloud: harden retained recovery diagnostics
- cloud: honor generation smoke cold starts
- cloud: isolate config app runtime import
- cloud: isolate connector smoke environments
- cloud: isolate sibling smoke environments
- cloud: keep connector diagnostic out of MVP release gate
- cloud: keep prod generation warm and scope spans
- cloud: make generation smoke single shot
- cloud: narrow connector loopback scope
- cloud: preserve accepted bundle promotion compatibility
- cloud: preserve recovery fence diagnostics
- cloud: recover first gateway to absence
- cloud: recover retained legacy gateway
- cloud: recover retained staging gateway to absence
- cloud: recover staging gateway to absence
- cloud: reject incomplete cold gateway plans
- cloud: retain production deployment evidence
- cloud: retain zero-floor generation attestations
- gateway: preserve grammar-safe FP8 profiles
- loadtest: prepare pinned perf-lab source before AWS
- model-ops: bind scaffold AUTO to trusted evidence
- model-ops: mirror task-equivalent standard blocks
- preserve cold-start reporter membership history
- quality-eval: preserve missing metric diagnostics
- quality: run comparison through locked project
- release: identify expired Modal attestation
- release: keep Modal diagnostic trim ASCII-only
- release: reconcile lagging Modal attestations
- release: tolerate completed Modal exec targets
- release: wait for refreshed PR head
- validate initial reporter lease atomically
v0.6.25 (2026-07-30)
Section titled “v0.6.25 (2026-07-30)”Highlights
Section titled “Highlights”- New capabilities: activate bootstrap-v2 pricing in prod; add —skip for recorded vendor-convergence skips; add cold-start live acceptance proof; add the mechanical launch-rate candidate promotion pipeline; add the size-class tier commercial policy to the rate book; admit rate-grade evidence in the launch catalog and coverage audit
- Reliability and operations: admit an image-witnessed authoritative zero at settlement; align runner authorization with smoke protocol; async-surface reservation parity — video guard, byte cap, tokenizer allowance; bound gateway cold-start recovery; cap identity convergence retry seam
- Performance: index the rotation-successor probe on api_keys
Features
Section titled “Features”- cloud: activate bootstrap-v2 pricing in prod
- cloud: add —skip for recorded vendor-convergence skips
- cloud: add cold-start live acceptance proof
- cloud: add the mechanical launch-rate candidate promotion pipeline
- cloud: add the size-class tier commercial policy to the rate book
- cloud: admit rate-grade evidence in the launch catalog and coverage audit
- cloud: batch/jobs generation via buffered single-terminal settlement
- cloud: bound and measure the sealed cold-start subsidy
- cloud: carry an authoritative-dimension claim set on UnitCounts
- cloud: carry the settled charge into realtime usage blocks
- cloud: commit reviewed launch allocations
- cloud: editable per-key spend limits + monthly auto-recharge cap
- cloud: editable per-key spend limits + monthly recharge cap
- cloud: GA audio on /v1/extract — book-driven whisper pricing
- cloud: gate production on a degraded staging acceptance
- cloud: generation input-token reservation bound + settle-contract pins
- cloud: gpu_second sealed billing regime — custom models become sellable
- cloud: install digest-bound generation routes
- cloud: install digest-bound governed generation routes
- cloud: lease dynamic cold-start reporters
- cloud: low-balance threshold email outbox + SES seam
- cloud: low-balance threshold emails via SES
- cloud: metering GA — waves 1–3
- cloud: move the license exclusion verdict to the registry resolution seam
- cloud: org-scoped usage export API — CSV + JSON
- cloud: persist cold-start subsidy evidence
- cloud: POST /v1/estimate dry-run cost endpoint
- cloud: POST /v1/estimate dry-run cost endpoint + SDKs
- cloud: production-bootstrap-v2 — prefill closed, sealed lanes priced
- cloud: promote launch-rate candidates to rate-grade evidence
- cloud: rate-grade size-class tier book assembly + operator-gated activation
- cloud: response-final settlement for streaming generation
- cloud: storage GiB-day accrual + book-driven register quote
- cloud: surface credits_charged + rate_book_version in usage
- cloud: unfence OpenAI-compat /v1/audio/transcriptions
- cloud: unify the settled-charge name across jobs, batches, and SDKs
- cloud: weekly/daily reconciliation with paging alerts and operator-gated LIVE mode
- cloud: weekly/daily reconciliation, attribution categories and the operator LIVE gate
- cloud: widen whole-job settlement to every contract dimension
- gateway: add optional governed generation routing seam
- model-ops: add ASR discovery family
- model-ops: add fail-closed discovery family registry
- model-ops: add multimodal discovery families
- model-ops: add object detection discovery
- model-ops: add OCR discovery families
- model-ops: add report-only generation discovery
- model-ops: add VQA and KIE discovery
- model-ops: discover guardrail candidates
- model-ops: discover sparse retrievers
- model-ops: discover structured extraction candidates
- model-ops: discover text classification candidates
- model-ops: discover text rerankers
- model-ops: read the executor’s trust clearance from the onboarding artifact
- model-ops: register guardrail discovery family
- model-ops: register structured extraction discovery
- model-ops: register text classification discovery
- model-ops: scan guardrail model candidates
- model-ops: scan structured extraction candidates
- model-ops: scan text classification candidates
- model-ops: support report-only family scans
- model-ops: wire the trust_remote_code scan into the AUTO gate
- models: onboard Qwen/Qwen3-Embedding-8B (scaffolded)
- sdk: estimate() on both SDKs for the /v1/estimate dry run
- server: bill sampled video frames as images
- telemetry: export the sealed cold-start metric remotely and chart it
Bug Fixes
Section titled “Bug Fixes”- benchmark: validate sorted deployment evidence
- catalog: bound GLiNER sequence lengths
- ci: reseal staging catalog evidence
- cloud: accept client-scoped WorkOS sessions
- cloud: accept Terraform canonical false in bootstrap guard
- cloud: accept WorkOS environment issuers
- cloud: address billing bootstrap review
- cloud: address CodeRabbit review — SES client bounds, send pacing, re-arm audit
- cloud: address CodeRabbit review on the /v1/estimate dry run
- cloud: address CodeRabbit review on the key-limit + recharge-cap PR
- cloud: address CodeRabbit review on the settled-charge surfaces
- cloud: address CodeRabbit review on the usage export
- cloud: address prerequisite gate review
- cloud: address the CodeRabbit review on the batch/jobs generation PR
- cloud: address the CodeRabbit round on the metering GA integration
- cloud: address the second CodeRabbit round on /v1/estimate
- cloud: admit an image-witnessed authoritative zero at settlement
- cloud: admit exact task-role tag bootstrap
- cloud: admit observed Modal US locations
- cloud: align benchmark smoke with rated profile
- cloud: align runner authorization with smoke protocol
- cloud: allow nested ECR repositories
- cloud: async-surface reservation parity — video guard, byte cap, tokenizer allowance
- cloud: attest benchmark runtime identity
- cloud: bind a rate_grade key to the price and hardware its envelope measured
- cloud: bind nightly health checks to staging
- cloud: bind paused forwarder candidate before prod cutover
- cloud: bind the charge-slot backstop to the stream, not the disconnect grace
- cloud: block the generation in/out split on a named measurement
- cloud: bound gateway cold-start recovery
- cloud: bound nightly evidence execution
- cloud: bound the sealed resolver cache and cap proxy response bodies
- cloud: cap identity convergence retry seam
- cloud: claim the hold before releasing it, and fold case like the registry
- cloud: classify benchmark evidence failures
- cloud: close cold recovery review gaps
- cloud: close cold-start evidence failure modes
- cloud: close final metering review faults
- cloud: close meter-only review gaps
- cloud: close metering GA review findings
- cloud: close rate evidence review gaps
- cloud: close reporter lease invariants
- cloud: close six gate findings on the /v1/estimate dry run
- cloud: close the batch-2a gate findings on editable limits + recharge cap
- cloud: close the CodeRabbit round on the candidate-promotion branch
- cloud: close the CodeRabbit round on the integration head
- cloud: close the CodeRabbit round on the reconciliation branch
- cloud: close the CodeRabbit round on the waves-1-3 integration head
- cloud: close the final CodeRabbit round on /v1/estimate
- cloud: close the gate findings on the resolution-seam verdict
- cloud: close the promotion gate’s operator-assertion escapes
- cloud: close the second CodeRabbit round on the credit surfaces
- cloud: CodeRabbit review — extract params must be explicit per line
- cloud: CodeRabbit review fixes — bounded sealed-class cache, doc reconciliation
- cloud: CodeRabbit review fixes — quote authority degradation, single-model projection
- cloud: CodeRabbit review fixes — sealed posture floor, generator provenance
- cloud: CodeRabbit round 2 — evidence doc, strict decode, single query walk
- cloud: CodeRabbit round 3 — exhaustive ceiling match, honest parse row
- cloud: CodeRabbit round on the control-plane prerequisite gate
- cloud: CodeRabbit round-2 — redact resolver errors, tighten class guard
- cloud: CodeRabbit round-6 — quote survives a dead pool, pin e2e clock
- cloud: CodeRabbit round-7 — atomic cache admission, sweep resilience
- cloud: compile exact generation customer routes
- cloud: compile exact generation routes
- cloud: correct console wallet path to /console/usage — verified against sie-web
- cloud: correct the race() return annotation in the snapshot test
- cloud: cover regular fleet cold starts
- cloud: cover regular Modal fleet cold starts
- cloud: decide license exclusion at the registry resolution seam
- cloud: declare generation input_tokens in the launch catalog
- cloud: default customer charging to disabled
- cloud: disable recharge schedule without teardown
- cloud: distinguish authoritative zero from missing count at settlement
- cloud: enforce the image witness on both sides of the rating boundary
- cloud: expose collector test configs to non-root
- cloud: fail a tier book whose price point has no exact Metronome decimal
- cloud: fail closed on incomplete preflight reads
- cloud: fail closed when a declared skip cannot be recorded
- cloud: fail loudly when test prerequisites vanish
- cloud: gate batch usage on settlement evidence, not a running counter
- cloud: gate estimates on resolved model identity
- cloud: gate fixes for the usage export — day-sliced scan, pool isolation, uniform masking
- cloud: harden cold-start acceptance proof
- cloud: harden dynamic cold-start membership
- cloud: harden staging release verification
- cloud: harden the multipart licensing gate and lock the route wiring
- cloud: index count-gated SES resources via one()/splat
- cloud: isolate benchmark evidence runtime
- cloud: log truncated usage-export streams; pin legacy-row null passthrough
- cloud: make extract canonical for Docling billing
- cloud: make generation rate evidence prefix-cache safe
- cloud: make test prerequisites and tracing deterministic
- cloud: make the jobs loopback estimate allowance-aware and prove the guard end to end
- cloud: make the jobs planner uphold the invariant settlement now enforces
- cloud: make the reconciliation bars statements the cost basis supports
- cloud: make the response cap operative, unbilled, and honest
- cloud: make the witnessed zero reach the wire and cover both rails
- cloud: merge candidate promotion and close #2628 review
- cloud: mirror the widened model canonicalization in the control plane
- cloud: observe the sealed cold-start metric; scope it to the prometheus path
- cloud: parse the multipart model under the route’s real body limit
- cloud: pin benchmark credential versions
- cloud: pin the sealed tariff evidence; harden the split and the posture tests
- cloud: pin the storage eligibility cutoff to UTC
- cloud: preflight input token rate bounds
- cloud: preserve cold-start evidence gaps
- cloud: preserve inactive prod rollback tasks
- cloud: price sealed-lane memory at the Sandbox rate, not standard
- cloud: propagate connector request identity
- cloud: publish a committed charge even when settlement then faults
- cloud: pull the authoritative invoice inside the single-flight lock
- cloud: recipient chain, re-arm edge, staleness bound, per-row commit
- cloud: reconcile what the batchgen merge hid from git
- cloud: record an admission outcome on every licensing rejection
- cloud: redact benchmark operator preflight
- cloud: reject a negative storage-accrual lookback
- cloud: reject an explicit storage-accrual day that is not closed
- cloud: reject expired recovery evidence
- cloud: reject null WorkOS client claims
- cloud: release the hold when the gateway cannot deliver the result
- cloud: renumber monthly-cap migration to 0040 — 0039 taken by sibling W3 branch
- cloud: renumber rotation-successor index migration to 0043 — 0042 taken by #2580
- cloud: report diagnostic cleanup failure
- cloud: report every live-preflight failure in one pass
- cloud: report malformed release pointers
- cloud: reseal release catalog dependents
- cloud: restore the spend-limit error contract on the W3 base
- cloud: retain exact benchmark deployment evidence
- cloud: retain exact benchmark deployment evidence\n\nRefs #2573
- cloud: retry benchmark identity convergence
- cloud: round-8 minors — env var name, authority restore, day-sweep guard
- cloud: route expected no-charge stream releases off the fault alerts
- cloud: route readiness failures through the gate without widening the opt-out
- cloud: route the async reservation through the shared video and token seams
- cloud: scope authoritative zero evidence
- cloud: scope rate preflight to exact authority
- cloud: scope the generation input bound to books that price input
- cloud: seal generation campaign artifact paths
- cloud: tell the truth about the video byte rail, and bound the tests that prove it
- cloud: tighten console origin validation, pace off the previous row
- cloud: tighten evidence diagnostic contract
- cloud: treat an explicit null usage as an empty block, not a refusal
- cloud: treat an explicitly null video as absent, not as a video
- cloud: type-gate the chat_template_kwargs allowlist
- cloud: unblock billing bootstrap migrations
- cloud: unseal console-free release subsets and offload-smoke
- cloud: unseal sealed release subsets, report every preflight failure, and gate degraded acceptance
- cloud: valid IAM tag charset + keep bootstrap secret-gate count at 2
- cloud: validate tier scheme inputs canonically
- cloud: wire the meter-only operator gate
- gateway: cap embeddings requests at 256 inputs
- gateway: harden governed generation dispatch
- gateway: preserve self-host grammar routing
- loadtest: use job token for GHCR pulls
- model-ops: accept GLiNER discovery models
- model-ops: address authorization review feedback
- model-ops: address evidence review feedback
- model-ops: authenticate the evidence producer and bind the identity to execution
- model-ops: bind CodeRabbit review completion
- model-ops: bind crash fallback repository
- model-ops: bind executor authorization
- model-ops: bind trusted completion status
- model-ops: close final trust gate review gaps
- model-ops: close malformed trust evidence gap
- model-ops: close three bypasses in the trust_remote_code gate
- model-ops: dispatch the Stage 4-5 gate completion explicitly
- model-ops: distinguish disabled discovery families
- model-ops: exclude OCR-only KIE models
- model-ops: explicitly dispatch discovery handoffs
- model-ops: gate readiness on bot reviews
- model-ops: harden bot review handoff
- model-ops: harden cold Modal image builds
- model-ops: harden completion status fallbacks
- model-ops: harden report-only scanning
- model-ops: harden trust gate provenance
- model-ops: keep a failed gate origin recoverable
- model-ops: narrow document OCR compatibility
- model-ops: never log gh’s stderr, which carries GH_TOKEN by construction
- model-ops: preserve completed review mutations
- model-ops: re-derive the clearance in matrix-block, bind evidence to the revision
- model-ops: re-grant contents:read to the dispatch job
- model-ops: recognize Copilot bot aliases
- model-ops: reconcile gate recovery reviews
- model-ops: refuse an unset RUNNER_TEMP instead of falling back into the workspace
- model-ops: refuse instead of crashing when the evidence fetch faults
- model-ops: refuse unusable evidence instead of falling back to blockers
- model-ops: reject ambiguous family matches
- model-ops: reject failed Copilot reviews
- model-ops: repair gate recovery replay
- model-ops: retry transient Modal build downloads
- model-ops: scan all loader-visible code
- model-ops: scan zero-shot classification models
- model-ops: suppress dynamic artifact failure logging
- model-ops: tolerate pending status retries
- model-ops: validate gate head before status writes
- model-ops: validate review completion evidence
- model-ops: wake the adapter executor from an agentless dispatch job
- model-ops: wake the adapter-exec executor by explicit dispatch
- quality-eval: retain models pending vanilla floors
- quality-eval: surface unfloored matrix rows
- release: bootstrap a fresh production database
- release: compare fresh proof types exactly
- release: retry malformed Modal attestations
- release: retry transient Modal attestation output
- release: verify empty replay target digest
- server: address CodeRabbit review on the video meter
- server: bill SigLIP padded text work
- server: bound isolation fan-out and close the NaN decode escapes
Performance Improvements
Section titled “Performance Improvements”- cloud: index the rotation-successor probe on api_keys
v0.6.24 (2026-07-26)
Section titled “v0.6.24 (2026-07-26)”Highlights
Section titled “Highlights”- New capabilities: add agent shell GPU fallbacks; add agent shell placement controls; add launch catalog contract canary; add Modal-native placement primitives; add native placement primitives; allow snapshotless native lanes
- Reliability and operations: align rate-book recovery checks; harden dispatch preflight gate; harden final generation gates; harden retained release recovery; preserve capacity diagnostics on timeout races
- Performance: cache MUVERA projection state
Features
Section titled “Features”- cloud: add agent shell GPU fallbacks
- cloud: add agent shell placement controls
- cloud: add launch catalog contract canary
- cloud: add Modal-native placement primitives
- cloud: add native placement primitives
- cloud: allow snapshotless native lanes
- cloud: attest realized placement identity
- cloud: establish broad-native release baseline
- cloud: govern prod legacy billing cutover
- cloud: preflight offline tokenizer dependencies
- server: pinnable LoRA adapter revisions in lora_paths
Bug Fixes
Section titled “Bug Fixes”- cloud: accept scheduled orphan-hold sweeps
- cloud: address native placement review
- cloud: address release review feedback
- cloud: align rate-book recovery checks
- cloud: align tokenizer snapshot dependencies
- cloud: allow absent generation in final smoke
- cloud: allow bounded prod platform bootstrap
- cloud: attest final gateway replacement
- cloud: bound snapshot lifecycle boots
- cloud: bound worker identity collection
- cloud: canonicalize generation GPU profiles
- cloud: close bounded prod bootstrap plan
- cloud: close launch review findings
- cloud: declare bootstrap alert container
- cloud: default orphan hold sweep ttl
- cloud: derive exact rate topology from broad release
- cloud: derive gateway gpu admission from plan
- cloud: diagnose bounded worker capacity waits
- cloud: drain legacy generation revisions
- cloud: enforce global canary request identity
- cloud: fail closed on tokenizer pin drift
- cloud: finalize gateway after generation deploy
- cloud: frame reusable candidate outputs
- cloud: gate billable release commit safely
- cloud: gate every gateway replica on dispatch readiness
- cloud: guard broad placement pending realized identity
- cloud: harden dispatch preflight gate
- cloud: harden final generation gates
- cloud: harden retained release recovery
- cloud: integrate launch remediation release gates
- cloud: keep slow batch leaders warm
- cloud: make Vercel deploy creation transport-safe
- cloud: package excluded gateway target
- cloud: package gateway from crate target
- cloud: preflight gateway generation dispatch
- cloud: preflight route and worker hash parity
- cloud: preserve capacity diagnostics on timeout races
- cloud: preserve launch route errors
- cloud: preserve snapshot cancellation cleanup
- cloud: reconcile launch remediation release gates
- cloud: recover ambiguous Vercel deploy creates
- cloud: recover bounded prod bootstrap outputs
- cloud: recover rate-book releases drain-first
- cloud: redact readiness diagnostics
- cloud: refresh staging catalog evidence digests
- cloud: rehearse immediate launch accounts
- cloud: reject partial readiness placement
- cloud: reserve token inputs from governed bounds
- cloud: respect Better Stack token scope
- cloud: restore authenticated gateway rollback health
- cloud: restrict rehearsal recovery fields
- cloud: serialize bootstrap state reads
- eval: make run-eval-group.sh and its tests portable to macOS
- eval: reject symlink escapes from the /tmp guard; reap the requirements temp dir
- loadtest: use valid AWS IAM purpose tag
- models: declare tokenizer dependencies
- quality: address remaining review findings
- quality: close nightly model gaps and add exact quarantine
- quality: gate bounded NVIDIA MUVERA checks
- quality: isolate parameterized FiNER evals
- quality: preserve floor metric contracts
- quality: record bounded concurrency-one targets
- quality: record faithful CMedQA floors
- quality: record faithful quality floors
- quality: reject non-finite qrel scores
- quality: restore bounded FiNER coverage
- quality: restore Mixedbread query expansion
- quality: stabilize scoped PR quality runs
- release: close prod bootstrap output contract
- release: redact bootstrap state read failures
- smoke: restore blocking-probe assertions and cold-touch retry coverage
- worker: copy sie_telemetry crate into worker image builds
Performance Improvements
Section titled “Performance Improvements”- server: cache MUVERA projection state
v0.6.23 (2026-07-24)
Section titled “v0.6.23 (2026-07-24)”Highlights
Section titled “Highlights”- New capabilities: accept native multimodal generation; support routed media measurement campaigns; add production bootstrap rate book; add pre-public task measurement candidates; bind launch catalog to commercial rate book; bind launch rates to commercial book
- Reliability and operations: harden generation quality sharding; harden multimodal batch validation; harden queue readiness for generation defaults; harden shard failure capture; restore governed retrieval floors
- Performance: bind bounded extraction evidence; bind current nonvision launch evidence; recover granite guardian queue evidence; reuse siglip preprocessing
Features
Section titled “Features”- api: accept native multimodal generation
- bench: support routed media measurement campaigns
- billing: add production bootstrap rate book
- catalog: add pre-public task measurement candidates
- catalog: bind launch catalog to commercial rate book
- catalog: bind launch rates to commercial book
- catalog: prepare launch serving evidence
- catalog: prepare non-public launch catalog checkpoint
- catalog: record current nonvision evidence
- catalog: record Snowflake and R3 evidence
- catalog: validate native multimodal recipes
- catalog: verify document vision recipes
- catalog: verify existing launch recipes
- catalog: verify GLiNER2 large launch evidence
- catalog: verify R3 pipeline on L4
- catalog: verify redaction task evidence
- catalog: verify Snowflake on preferred L4
- ci: complete req14 gates after lane runs
- ci: remove runner-held waits from official-recipe gates
- cloud: activate accounts and route launch events
- cloud: add bounded prod platform bootstrap
- cloud: add exact launch rate campaign suite
- cloud: add immutable managed release pipeline
- cloud: add launch account credit flow
- cloud: add launch rate campaign suite
- cloud: add Qwen task evidence campaign
- cloud: add secure staging catalog evidence dispatch
- cloud: add secure staging evidence dispatch
- cloud: archive Modal measurement releases
- cloud: assemble commercial rates from exact evidence
- cloud: authenticate Modal rate evidence
- cloud: bracket generation rate window
- cloud: compile isolated measurement topology
- cloud: derive commercial rates from exact resource seconds
- cloud: derive commercial rates from Modal billing
- cloud: execute document vision recipes
- cloud: execute remaining reviewed recipes
- cloud: route signup notifications directly to Slack
- cloud: verify BGE and GTE launch evidence
- deploy: add scalr-bootstrap root for internal terraform state
- deploy: run scalr-bootstrap under opentofu with live scalr backend
- deploy: scalr-held terraform state for internal roots
- deploy: split scalr workspaces into cli- and vcs-driven
- generation: type structured output contracts
- model-ops: split feasibility PROCEED-WITH-ADAPTER-AUTO vs human
- model-ops: trust_remote_code security pre-gate
- quality-eval: in-pipeline wall-clock stop with partial report
- sdk: add typed Responses client
Bug Fixes
Section titled “Bug Fixes”- 1841: stamp sealed metric units on the Modal collector, and make the CP filters cover what the suite reads
- adapters: serialize output_hidden_states forwards — close the #2144 recorder-race class
- adapters: serialize output_hidden_states forwards to close the #2144 recorder-race class
- api: address multimodal contract review
- api: close multimodal parity gaps
- bench: add faithful PyLate official recipe
- bench: address launch evidence review
- bench: align Any2Any dispatch semantics
- bench: align ColBERT reference validation
- bench: align pinned dataset metadata
- bench: bind generation evidence to runtime protocol
- bench: bind Modal measurement provenance
- bench: bind Qwen3.6 launch evidence to runtime identity
- bench: canonicalize campaign route identities
- bench: clean up failed server startups
- bench: close operation evidence review gaps
- bench: discard partial server results
- bench: emit strict quality JSON
- bench: fail closed on benchmark operation identity
- bench: forward only encode/score to EvalRunner in quality path
- bench: harden generation quality sharding
- bench: harden multimodal batch validation
- bench: harden queue readiness for generation defaults
- bench: harden shard failure capture
- bench: keep image types out of runtime imports
- bench: pin ragged PyLate reranking fix
- bench: pin trusted dataset loaders
- bench: preserve canonical FewRel dataset identity
- bench: preserve generation evidence imports
- bench: preserve invalid campaign archives
- bench: preserve profile benchmark identity
- bench: preserve safe Modal shard failures
- bench: reject extraction item errors
- bench: reject incompatible reference operations early
- bench: reject incomplete shard assembly
- bench: reject partial performance measurements
- bench: remove stale generation import
- bench: report empty perf evidence explicitly, not as ambiguous
- bench: restore governed retrieval floors
- bench: route profiled performance requests
- bench: satisfy review static analysis
- bench: start named profile routes
- bench: stop plural options poisoning choice extraction
- bench: type bounded image dispatch
- bench: validate multimodal encode batches
- bench: version routed rate candidates
- billing: bind Modal tags to pretraffic contract
- billing: bind zero-page job evidence
- billing: preserve zero-page job evidence
- billing: project observed units through reservation plans
- billing: restrict zero terminals to pages
- billing: verify multimodal fixture units
- catalog: bind named output evidence
- catalog: disable unsafe qwen speculative default
- catalog: harden native acceptance validation
- catalog: raise GTE Muvera repetitions
- catalog: route mxbai edge through native attention
- catalog: use pinned GTE query length
- catalog: validate named profile evidence
- ci: authenticate quality dependency sync
- ci: diff quality evals from current PR base
- ci: fail closed on mixed Modal output
- ci: isolate privileged req14 completion
- ci: isolate quality eval evidence
- ci: isolate quality runner credentials
- ci: let scalr-bootstrap track the opentofu pin
- ci: repair Qwen quality shard workflow
- ci: verify trusted req14 callback checkout
- cloud: accept canonical Modal volume mount
- cloud: accept Modal GCP provider identity
- cloud: address release workflow review follow-ups
- cloud: align launch evidence contracts
- cloud: align release lock namespace
- cloud: align release probes with scoped listing
- cloud: allow release image manifest readback
- cloud: allow runtime secret namespace
- cloud: allow scoped release key discovery
- cloud: allowlist release child environments
- cloud: approve permissive launch licenses
- cloud: attest LightOnOCR image compatibility
- cloud: attest Modal snapshot memory limits
- cloud: bind billing to executing Modal app
- cloud: bind billing to Modal environment
- cloud: bind commercial evidence to routed candidates
- cloud: bind ECR login to OIDC identity
- cloud: bind rate probes to archived model pins
- cloud: bind score probes to campaign batches
- cloud: bind topology plan semantics
- cloud: bootstrap purpose-bound Slack sinks
- cloud: bound and isolate release preflight reads
- cloud: classify ECS start failures first
- cloud: complete managed release operability follow-ups
- cloud: complete release operability audit
- cloud: dispatch exact catalog profiles
- cloud: dispatch quality shards in Modal batch mode
- cloud: fail closed in candidate image lookup
- cloud: fail closed on ECR lookup errors
- cloud: finish release child isolation
- cloud: finish release preflight hardening
- cloud: guard acceptance projection freshness
- cloud: guard catalog projection freshness
- cloud: harden commercial evidence joins
- cloud: harden launch identity and approval invariants
- cloud: harden launch rate campaign probes
- cloud: harden managed release promotion
- cloud: harden notification launch paths
- cloud: harden screenshot acceptance
- cloud: harden vision acceptance bounds
- cloud: include smart chat task evidence
- cloud: isolate launch credit approvals
- cloud: keep LightOn preflight CPU-safe
- cloud: make topology imports typecheck
- cloud: map static lanes to catalog pools
- cloud: measure both SigLIP towers
- cloud: pin batch worker execution identity
- cloud: pin dev worker identity
- cloud: pin launch measurement resources
- cloud: pin native worker execution identity
- cloud: pin staging batch memory limit
- cloud: preflight live release dependencies
- cloud: preflight Modal SGLang runtime
- cloud: preflight release task arguments
- cloud: preserve concurrent topology output
- cloud: preserve governed release tool paths
- cloud: preserve sidecar image contract
- cloud: prime catalog before gateway rollout
- cloud: probe exact campaign request batches
- cloud: reconcile commercial rates with launch profiles
- cloud: reconcile measurement topology with launch catalog
- cloud: refresh staging evidence pins
- cloud: reject duplicate generation shards
- cloud: reject malformed release metadata
- cloud: reject mismatched catalog evidence
- cloud: reject undecodable evidence
- cloud: reject unversioned S3 evidence
- cloud: require complete pii redaction proof
- cloud: reseal licensed launch release locks
- cloud: reseal production rate release locks
- cloud: restore CUDA toolkit to SGLang sidecars
- cloud: retain published measurement evidence
- cloud: retire legacy environment ECR state
- cloud: scope Better Stack credentials by product
- cloud: separate guardrail policy fixtures
- cloud: use lightweight generation evidence contract
- cloud: validate launch evidence integrity
- cloud: verify control-plane runtime dependencies
- cloud: verify Modal markers from files
- cloud: verify redact pii acceptance
- cloud: verify redact PII postprocessing
- deploy: pin internal scalr workspaces to opentofu
- gateway: validate schema keywords by context
- generation: close structured grammar review gaps
- grammar: bound pointer index parsing
- grammar: harden schema normalization
- loadtest: manage github-token in terraform + sync via ESO, pin sie-perf-lab ref
- loadtest: persist github-token in terraform, sync via ESO, pin perf-lab ref
- loadtest: remove illegal variable interpolation from output description
- model-ops: close review gaps in the AUTO verdict contract
- model-ops: require an explicit matrix_task key in auto_authoring
- profiles: honor routed outputs and retrieval recipes
- quality-eval: address Copilot review on the wall-clock stop
- quality-eval: remove shadow tree on exception paths, not just returns
- quality-eval: select governed official metric
- quality-eval: serve non-default smoke bundles from a pip —target shadow so the overlay reaches the bundle floor
- quality-eval: validate wall_budget_minutes inside run_pipeline
- release: authenticate gateway rollback probe
- release: keep sie-audio-prep internal to builds
- sdk: accept nullable grammar metadata
- sdk: keep grammar types on supported module path
- security: include aws edge in standing review
- security: pin hf_revision on the 10 unpinned trust_remote_code models
- server: apply trained Rotary ColBERT projection
- server: defer LightOn OCR compatibility patch
- server: hide unsupported visual muvera profiles
- server: mask ColBERT expansion keys in flash attention
- server: meter SigLIP tokens from attention masks
- server: pin launch guardrail policies
- server: preserve stream helper compatibility
- server: preserve trusted code revision on fallback
- server: restore ColBERTv2 retrieval recipe
- server: restore Jina ColBERT retrieval recipe
- server: serialize remaining tokenizer access
- sie_server: register remote-code snapshot dir on sys.path for ST-Router models
- sie-bench: isolate PyLate reference evaluation
- sie-bench: preserve expansion token attention
- smoke: fail loud when a bundle overlay undershoots its declared pins
- terraform: manage github-token container only, drop placeholder version
- tooling: derive overlay cache version from lock
- tooling: gate Modal quality shard fanout
- tooling: harden launch failure handling
- tooling: isolate Modal shell worktree environment
- tooling: make empty Modal status succeed
- tooling: mount tracked playbook fixtures on Modal
- tooling: namespace Modal JIT caches by ABI
- tooling: preserve Modal status errors
- tooling: preserve multiline Modal commands
- tooling: read Modal lock from mounted workspace
- tooling: resolve modal from task path
- tooling: scope Modal inventory conflicts
- tooling: validate Modal description types
- tooling: version Modal datasets cache
- vision: harden launch evidence
- vision: pad fused detection batches
Performance Improvements
Section titled “Performance Improvements”- catalog: bind bounded extraction evidence
- catalog: bind current nonvision launch evidence
- generation: recover granite guardian queue evidence
- vision: reuse siglip preprocessing
v0.6.22 (2026-07-22)
Section titled “v0.6.22 (2026-07-22)”Highlights
Section titled “Highlights”- New capabilities: bind exact rate measurement windows; shard OCR quality by page; support package revision evidence; attest worker execution identity; fan out quality shards on Modal; pin immutable Docling artifacts
- Reliability and operations: harden OCR shard provenance; harden package revision evidence; fail fast on managed gateway lock drift; attest OCR shard execution identity; clarify complete rate evidence
- Performance: concurrent OCR page dispatch so server batching actually engages
Features
Section titled “Features”- bench: bind exact rate measurement windows
- bench: shard OCR quality by page
- bench: support package revision evidence
- cloud: attest worker execution identity
- cloud: fan out quality shards on Modal
- server: pin immutable Docling artifacts
- server: support immutable Docling model artifacts
- tooling: attach named Modal shell secrets
Bug Fixes
Section titled “Bug Fixes”- bench: attest OCR shard execution identity
- bench: clarify complete rate evidence
- bench: clarify rate-window mismatch errors
- bench: harden OCR shard provenance
- bench: harden package revision evidence
- bench: mark serving basis non-commercial
- bench: require exact rate resource evidence
- ci: install sdk deps for catalog tests
- ci: queue release rerun when main moves
- ci: regenerate launch catalog on release PR
- ci: regenerate release PR when main moves
- cloud: bind benchmark evidence to deployed workers
- cloud: bind benchmark model identity
- cloud: bind quiesced ECS candidates before scaling
- cloud: bind reviewed catalog recipe executors
- cloud: bind worker resource attestation
- cloud: fail fast on managed gateway lock drift
- cloud: include compiler version dependency
- cloud: isolate catalog compiler dependencies
- cloud: meter detection catalog smoke images
- cloud: preserve deployment validation ordering
- cloud: preserve identity across billing handoff
- cloud: reject noninteger GPU identity counts
- cloud: remove member key spend ceiling
- cloud: require clean governed release sources
- cloud: validate integral worker identity fields
- cloud: validate managed gateway lock before release
- eval: constrain temporary artifact root
- eval: serve canonical profile routes
- tooling: bind release locks from clean head
- tooling: close Modal fanout lifecycle races
- tooling: fail closed on Modal volume sync
- tooling: isolate Modal uv cache
- tooling: isolate quality shards on Modal v2
- tooling: keep active modal shells alive
- tooling: preserve Modal fanout failures
- tooling: reject symlink fanout outputs
- tooling: resolve modal in detached shells
Performance Improvements
Section titled “Performance Improvements”- sie-bench: concurrent OCR page dispatch so server batching actually engages
v0.6.21 (2026-07-21)
Section titled “v0.6.21 (2026-07-21)”Highlights
Section titled “Highlights”- New capabilities: add rate-grade evidence contract; add reusable local queue stack; add upstream GLiNER NER reference; govern vision launch evidence; seed load-test request selection; SIESparseImageWrapper for served visual sparse retrieval
- Reliability and operations: capture generation timeout evidence; harden local queue stack lifecycle; drop stale metrics registry dependency; finalize billing money hardening; harden account lifecycle rehearsal
- Performance: record SPLADE 512-token run; record extraction L4 measurements; compact sparse output in native precision; fuse packed SPLADE pooling
Features
Section titled “Features”- bench: add rate-grade evidence contract
- bench: add reusable local queue stack
- bench: add upstream GLiNER NER reference
- bench: govern vision launch evidence
- bench: seed load-test request selection
- bench: SIESparseImageWrapper for served visual sparse retrieval
- billing: meter multimodal catalog units
- candle: add and optimize native SPLADE sparse encoding
- candle: add native SPLADE sparse encoding
- cloud: add exact qwen35 perf probe
- cloud: add fail-closed catalog acceptance gate #1943
- cloud: add fail-closed catalog acceptance runner
- cloud: add fast chat catalog acceptance
- cloud: add managed-edge abuse and tenant gates
- cloud: add OCR and speech acceptance probes
- cloud: add operator key administration
- cloud: add operator key administration and keyed usage
- cloud: add OWLv2 acceptance executor
- cloud: add OWLv2 catalog acceptance
- cloud: add staging account lifecycle rehearsal
- cloud: add staging rehearsal rate authority
- cloud: audit measured launch rate-book coverage
- cloud: audit measured rate book coverage
- cloud: bind verified DINO and Whisper evidence
- cloud: declare exact lane placement resources
- cloud: declare exact lane resources
- cloud: enforce conserved usage rating
- cloud: enforce exact managed catalog authority
- cloud: execute text acceptance recipes
- cloud: exercise DINO catalog acceptance
- cloud: expose operator billing health
- cloud: gate catalog tasks on evidence
- cloud: govern launch catalog topology
- cloud: govern vision launch catalog
- cloud: integrate control-plane launch readiness
- cloud: integrate exact request billing authority
- cloud: notify operators of new signups
- cloud: pack Whisper into shared launch pool
- cloud: record generation launch evidence
- cloud: record Qwen3.5 launch evidence
- cloud: rotate staging rehearsal rate book
- cloud: serve release-locked model catalog
- cloud: serve release-locked model catalog #1944
- cloud: settle exact pair and generation usage
- cloud: ship proprietary vision launch catalog
- cloud: validate extraction catalog recipes
- cloud: validate structured output acceptance
- cloud: verify qwen35 A100 model evidence
- cloud: wire privacy-safe console RUM
- cluster: opt-in bake cache write via SIE_BAKE_CACHE_WRITE
- core: expose authoritative request usage
- core: report authoritative score pair usage
- evidence: propagate worker execution identity
- extract: add pinned PII quality gate
- gateway: add strict Cohere rerank compatibility
- model-ops: feasibility accepts transformers-5-line bundle choice
- model-ops: scaffold re-renders cloud release locks with the model YAML
- models: onboard Snowflake/snowflake-arctic-embed-s (scaffolded)
- models: raise qwen36 launch context to 8k
- playbooks: upstream SparseEncoder reference arm for sparse-visual models
- quality-eval: wire the signature-DB loop — adjudicated-skip + verdict write-back
- reranking: aggregate launch runtime and release locks
- reranking: harden Qwen score runtime
- sdk: expose terminal billing metadata
- sdk: preserve terminal request metadata on errors
- serve visual-SPLADE via transformers5 (sparse-vision adapter + measurement path)
- server: bind package artifact identity
- server: bind package-backed model artifacts
- server: transformers514 bundle + SparseEncoder vision sparse adapter
- sie-bench: add generation quality sharding
- sie-bench: add generic generation quality sharding
- telemetry: centralize observability on OTLP
- vision: harden OSS runtime contracts
- vision: ship OSS launch runtime and evidence
Bug Fixes
Section titled “Bug Fixes”- audio: sandbox image builds the audio wheel opportunistically
- bench: align local generation lane routing
- bench: attest generation execution revision
- bench: bind broker artifact ARN
- bench: capture generation timeout evidence
- bench: clarify direct revision semantics
- bench: close execution identity provenance gaps
- bench: exercise generation queue sidecar
- bench: harden local queue stack lifecycle
- bench: isolate generation smoke stack lifecycle
- bench: isolate queue smoke ports
- bench: make fake stack no-op explicit
- bench: preserve ChemProt entity offsets
- bench: preserve exact reranker task identity
- bench: preserve NER task parameters
- bench: reject coerced execution identity
- bench: remove failure log content
- bench: render generation smoke help
- bench: resolve canonical AG News dataset
- bench: route canonical model profiles
- bench: select encode output families
- bench: separate model and machine profiles
- bench: serialize image rerank inputs
- bench: use SPIA field labels for PII eval
- bench: verify broker artifact bucket before reads
- bench: warm generation at measured concurrency
- billing: preserve legacy outbox drain
- billing: reject unexpected metronome rates
- billing: verify vendor-bound raw usage
- candle: bound CUDA compiler parallelism
- candle: include CUDA math constants for SPLADE
- candle: preserve tiny SPLADE log1p weights
- ci: bind cloud gateway cache to OSS source
- ci: install catalog compiler dependency
- ci: load sidecar image for release smoke tests
- ci: make cloud gateway checks deterministic
- ci: supersede obsolete pull request runs
- ci: validate release guard wait settings
- ci: wait for release check reruns
- cloud: accept numeric whisper fixture transcript
- cloud: address billing review findings
- cloud: address rate authority review
- cloud: align acceptance with native labels
- cloud: align rehearsal rates with measured audit
- cloud: benchmark qwen35 through native generate
- cloud: bind acceptance hash to source artifact
- cloud: bind exact lane cpu limits
- cloud: bind generation grammar pool
- cloud: bind structured schema semantics
- cloud: bound outbox poison isolation
- cloud: clarify idempotent deletion tests
- cloud: clean up metered rehearsal credits
- cloud: close console release preflight gaps
- cloud: close launch integration review gaps
- cloud: close launch review gaps
- cloud: close operator key launch gaps
- cloud: converge managed catalog exactly
- cloud: cover multimodal rehearsal units
- cloud: declare redact PII labels
- cloud: derive migration log stream
- cloud: drop stale metrics registry dependency
- cloud: enforce billing activation invariants
- cloud: finalize billing money hardening
- cloud: govern legacy billing cutover
- cloud: harden account lifecycle rehearsal
- cloud: harden billing money flows
- cloud: harden generation acceptance evidence
- cloud: harden launch rate evidence
- cloud: harden lifecycle rehearsal cleanup
- cloud: harden signup alert routing
- cloud: harden Stripe clawbacks and usage delivery
- cloud: include registry compiler dependency
- cloud: keep cleanup accounts suspended
- cloud: keep live secret values out of state
- cloud: keep proxy dry-run provenance current
- cloud: keep qwen35 measurements pending
- cloud: keep qwen35 performance pending
- cloud: match Better Stack canonical alert fields
- cloud: pin exact console release commits
- cloud: preserve charged structured failures
- cloud: preserve committed key recovery
- cloud: preserve custom catalog evidence
- cloud: preserve proxy token shape in dry-run
- cloud: rebase catalog acceptance gate
- cloud: reconcile exact metronome rating
- cloud: reconcile native vision label bindings
- cloud: refresh release lock artifact pins
- cloud: reject duplicate billing bindings
- cloud: reject empty key scopes
- cloud: reject malformed catalog exports
- cloud: reject unknown extraction recipes
- cloud: reject unsupported lane providers
- cloud: report gross Stripe funding
- cloud: require console release domain
- cloud: require ordered whisper transcript
- cloud: restore Metronome billing authority
- cloud: restore telemetry branch CI
- cloud: retain charged acceptance failures
- cloud: tighten billing terminal invariants
- cloud: tolerate inaccessible gateway candidates
- cloud: validate acceptance readiness reasons
- cloud: validate edge origin before secret sync
- cloud: validate lazy snapshot lanes
- cloud: validate modal cloud placement
- cloud: validate native label binding kinds
- cloud: validate rate tariff shape
- colpali: serialize model forwards — transformers output-recorder race
- colpali: stop GPU memory accumulation on Vidore evals
- control-plane: make wallet authority explicit
- core: preserve unit count wire order
- deploy: bound generation readiness process
- deploy: harden generation readiness gate
- deploy: honor local weight source precedence
- deploy: prove exact generation revision readiness
- deploy: resolve model staging weight sources
- eval: add faithful Qwen3-VL reference
- eval: bound multimodal reference scoring
- eval: bound reranker candidates by corpus
- eval: isolate VL reference dependencies
- eval: keep gate plan aligned with dispatch
- eval: preserve VL text coverage
- eval: reject empty reranker qrels
- evidence: bind immutable model identity
- evidence: close worker identity contract gaps
- evidence: harden worker identity boundaries
- extract: close review validation gaps
- extract: filter unconstrained relation endpoints
- keda: keep gateway coordination under autoscaling
- keda: simplify forward OTLP upgrade
- keda: support forward-only OTLP upgrade
- keda: verify lean forward upgrade
- keep audio prep release pin in sync
- loadtest: skip doomed bench-only scenarios when no release deployed
- metering: reject invalid score counts
- model-ops: align exact vision gate evidence
- model-ops: gate parses the trigger’s ‘Triggered by #N’ issue reference
- model-ops: void gate attempts that provably began no spend
- model: restore SPLADE quality sequence cap
- models: pin hf_revision for gte-Qwen2-7B remote-code load
- models: restore SPLADE 512-token capacity
- models: serve gte-Qwen2-7B via its own bidirectional forward
- playbooks: v2-native task extraction + timeout contract for the sparse arm
- quality-eval: identity-based signature dedup + branch hash (#2096 review)
- quality-eval: make runner wake parallel, fail-loud, and demand-accurate
- quality-eval: reset stale WakeRequestedAt against launch time
- quality-eval: tolerate YAML date-typed fields in signature identity (#2096 review)
- quality: tolerate missing gpu identity tooling
- refresh managed release locks
- release: refresh deployment locks during release regeneration
- release: run deployment lock refresh in project env
- reranking: address exact-head review
- reranking: emit authoritative fake usage
- reranking: emit fake score usage
- reranking: reject blank compatibility inputs
- server: empty bundle model_filter must advertise zero models, not all
- server: remove unused fake adapter import
- sie-bench: harden generation sharding evidence
- sie-tools: keep session failures auth-scoped
- sie-tools: preserve account lifecycle errors
- smoke: align OTel collector contract checks
- smoke: align OTLP metric assertions
- smoke: restore kind smoke on main
- splade: align telemetry and model docs
- splade: harden profile and sparse defaults
- tooling: fail fast before Modal dataplane runs
- tooling: stop modal shells by app id
- tools: preserve baked mise in modal shells
- vision: address review edge cases
- worker: propagate effective sequence caps
- worker: release OOM tracebacks in the no-recovery branch too (CodeRabbit)
Performance Improvements
Section titled “Performance Improvements”- benchmarks: record SPLADE 512-token run
- bench: record extraction L4 measurements
- candle: compact sparse output in native precision
- candle: fuse packed SPLADE pooling
- candle: pack SPLADE BERT inference
- vision: bulk-convert detection outputs
v0.6.20 (2026-07-18)
Section titled “v0.6.20 (2026-07-18)”Highlights
Section titled “Highlights”- New capabilities: add GLiGuard and fix extraction catalog behavior; add GLiGuard and fix LightOnOCR and GLiREL extraction; add cloud benchmark campaigns; add native R3-Skill evaluation; add staging cloud smoke path; add two-client staging rehearsal
- Reliability and operations: fence broker rollback recovery; harden cloud campaign evidence; harden cloud campaign readiness; harden live campaign preflight; harden managed campaign control plane
- Performance: cache exact bf16 gelu values; fuse ModernBERT exact GeGLU; optimize ModernBERT CUDA hot path; vectorize bf16 gelu lookup gate
Features
Section titled “Features”- add GLiGuard and fix extraction catalog behavior
- add GLiGuard and fix LightOnOCR and GLiREL extraction
- bench: add cloud benchmark campaigns
- bench: add native R3-Skill evaluation
- bench: add staging cloud smoke path
- bench: add two-client staging rehearsal
- bench: complete managed cloud campaigns
- bench: cover cloud task modalities
- bench: inspect diagnostic campaign archives
- bench: operationalize cloud measurement campaigns
- bench: operationalize managed SIE Cloud measurement campaigns
- ci: fake-stack-test job — Fake Engine regression suite + Cloud dispatcher loopback
- ci: fake-stack-test job + five Fake Engine regression tests + Cloud loopback
- cloud-gateway: serve job chunk results via signed capability URLs
- cloud-gateway: serve offload job results via signed capability URLs
- cloud: apex-domain cutover + console deploy model
- cloud: control-plane audit-read + org-settings + residency-read endpoints (console Tier-2)
- cloud: export worker-lane OTLP traces to the collector (lane-only, #1740)
- cloud: gateway OTLP trace+metric export — clean end-to-end telemetry spine
- cloud: govern Modal deployments with typed specs
- cloud: per-org alert-threshold toggles (0006 alert_prefs) gating _fire_alerts
- cloud: per-org console service keys via /internal/console-keys
- cloud: per-org console service keys via POST /internal/console-keys
- cloud: relocate the schedule runner in-VPC (EventBridge -> /internal/schedules/run-due)
- cloud: resolve fresh keys on-demand to kill the ~5-10s gateway 401 window
- cloud: surface the saved card’s brand/last4/exp (Stripe payment-method retrieve)
- cloud: WP4 billing-depth control-plane endpoints (invoices, usage, alert toggles, payment method)
- cloud: WP4 billing-depth endpoints — invoices, usage series, alert-config toggles
- control-plane: make auto-recharge fire end-to-end
- mac-smoke: additive fake phase 0 — weightless surface contract
- model-ops: add stage-5 quality+perf onboarding gate
- model-ops: add-model trigger — labeled issue starts the onboarding agent
- model-ops: stage-5 quality+perf onboarding gate — measure vs target, bounded fix loop, evidence packet
- models: add Tencent R3-Skill support
- models: onboard intfloat/multilingual-e5-small (scaffolded)
- quality-eval: #1810 experiment runner — parallel Modal sessions, GPU selection, budget backstop
- quality-eval: autonomous experiment runner — parallel Modal sessions, GPU selection, budget backstop (Req14 v2 3/4)
- quality-eval: end-to-end pipeline wiring + dedup/budget semantics (Req14 v2 4/4)
- quality-eval: wire triage→planner→runner→autofix into a bounded pipeline
- server: add deterministic weightless fake adapter family
- server: failure injection for the fake adapter family
- server: fake adapter family — deterministic weightless sie-fake models
- server: synthetic memory-tracker mode — deterministic pressure/eviction
- server: synthetic memory-tracker mode for deterministic eviction
- support Iso ModernColBERT on Candle
- tests: lost-result 120s characterization on the full queue topology
- tests: lost-result characterization on the full queue topology
Bug Fixes
Section titled “Bug Fixes”- bench: accept canonical UTF-8 broker artifacts
- bench: accept omitted optional request refs
- bench: align broker contract boundary nesting
- bench: align broker log retention
- bench: align campaign runtime dependencies
- bench: avoid imagebuilder template collision
- bench: bind ami runtime kernel evidence
- bench: bound R3 quality regression gate
- bench: bound staging smoke below service knee
- bench: close campaign review and ci gaps
- bench: close campaign review edge cases
- bench: expose safe broker failure diagnostics
- bench: externalize broker deployment settings
- bench: fence broker rollback recovery
- bench: harden cloud campaign evidence
- bench: harden cloud campaign readiness
- bench: harden live campaign preflight
- bench: harden managed campaign control plane
- bench: make SSM dispatch at most once
- bench: pass authorized request to orchestrator
- bench: pin runnable worker manifests
- bench: prepare campaign clients before start
- bench: preserve broker terminal state
- bench: preserve cloud campaign timing evidence
- bench: preserve Fargate scratch permissions
- bench: probe exact start control key
- bench: publish worker image to owned ecr path
- bench: reject malformed canonical artifacts
- bench: relocate runtime caches to scratch
- bench: remove stale quality import
- bench: retain ami assertion diagnostics
- bench: retry transient ECS role propagation
- bench: retry transient R3 score timeouts
- bench: roll task revisions without downtime
- bench: secure ephemeral identity creation
- bench: secure staging cloud validation
- bench: skip absent profile deletion
- bench: support repeated maintenance pauses
- bench: trust scheduler group for worker expiry
- bench: validate R3 evaluation inputs
- candle: close ModernBERT audit gaps
- ci: disable implicit cloud tool installs
- ci: pin nested mise runtime
- ci: verify cloud IaC toolchain
- cloud-gateway: complete loopback class-fix via one canonical helper; align Python mirror + docs (#1906 final review)
- cloud-gateway: IP-literal loopback check + recover stuck payload-volume reload (CodeRabbit #1906 r2)
- cloud-gateway: reject non-ASCII signature hex without panicking
- cloud-gateway: validate capability-URL origins, bound TTL, and fix object-store minting (CodeRabbit #1906)
- cloud: address CodeRabbit — same-item batch id pairing + metrics protocol env + /metrics deny log
- cloud: address CodeRabbit — serialize env-mutating metrics-protocol test with a module-local ENV_LOCK
- cloud: address follow-up review
- cloud: address Modal deployment review
- cloud: alert on zero-metering invoice drift; reject ambiguous usage_events apportionment
- cloud: align dev backend benchmark
- cloud: align vendor billing with ledger debits
- cloud: batch/offload jobs retry retryable per-item MODEL_LOADING instead of failing the chunk
- cloud: bound on-miss key resolver’s single-flight map + harden lookup
- cloud: CodeRabbit CP review — alert-configs PUT 404s, invoice-payload 502 guard, ty-ignore format
- cloud: console SIE_GATEWAY_URL is the public edge, not the raw Modal URL
- cloud: customer-scope the saved-card lookup — close cross-tenant card-PII leak
- cloud: deploy otel-collector + merge OTLP endpoint before lane snapshot so lane tracing activates
- cloud: derive outbox credits from the ledger debit and alert on reconciliation drift
- cloud: diagnose stale-checkout CP destroy + bridge VERCEL_API_TOKEN
- cloud: diagnose stale-checkout CP destroy + bridge VERCEL_API_TOKEN in cloud-provision
- cloud: don’t persist raw billing_email in audit_log detail (PII)
- cloud: final WP4 review — strict invoice currency, atomic alert toggle, UTC-naive parse, card guard
- cloud: gateway OTLP export — blocking reqwest client for span batch processor + HTTP/TLS metric transport
- cloud: gateway OTLP export — Tokio runtime for span batch processor + HTTP/TLS metric exporter
- cloud: harden console-key lifecycle
- cloud: harden whole-estate finalization
- cloud: hide console keys from /me/usage + /me revoke (security review)
- cloud: preserve billing alert and meter attribution
- cloud: refresh recovery handle on batch chunk re-spawn (avoid orphaned container + 409 spam)
- cloud: release the scheduler lock_conn on any pre-hand-off failure
- cloud: stage the gen-lane model in model-stage and env-qualify the unstaged-weights hint
- cloud: stop deploy from dirtying satellite uv.locks
- cloud: stop deploy from dirtying satellite uv.locks (clean-tree guard)
- cloud: WP4 review follow-ups — usage upper bound, invoice-fallback diagnosability, card exp contract
- control-plane: hand auto-recharge sweep off the EventBridge request path
- control-plane: harden auto-recharge collection
- control-plane: make recharge retries lease-safe
- fake-engine: address two-axis review findings across the stack
- fake-engine: CI loopback interop + CodeRabbit review findings
- gateway: attest deployed execution revision
- gateway: include unanimous error_code in all_items_failed 500 logs
- gateway: reject malformed queue-request bodies with a typed 400
- harden classification quality verification
- model-ops: apply CodeRabbit review findings on the stage-5 gate
- model-ops: apply review findings on the friction PR
- model-ops: deterministic escalation when the agent itself crashes
- model-ops: filter vanilla floor material to fleet metric families
- model-ops: harden add-model trigger per review
- model-ops: raise directly in _read_json per code-quality finding
- model-ops: read vanilla numbers from the verdict’s evidence field
- model-ops: rehearsal follow-ups — agent guardrail, record valves, auto-review, log-tail noise
- model-ops: rehearsal follow-ups — guardrails, record valves, log-tail noise
- model-ops: scaffolder crash on present-but-null sibling fields
- model-ops: seed-floor emits fleet metric families, not the raw mteb vector
- model-ops: stage-5 gate read vanilla numbers from verdict evidence, not probe measurements
- model-ops: tighten the agent-crash fallback per review
- ocr: address review lifecycle findings
- preserve parameterized quality gate targets
- quality-eval: accurate backstop messages for CPU playbooks (review S1)
- quality-eval: guard non-dict cluster signals in the pipeline (review)
- quality-eval: harden experiment runner per PR #1883 review (traversal, stale-verdict, rollback, build-clock)
- quality-eval: harden pipeline per PR #1891 review (report/verdict/workflow)
- quality-eval: reap process groups in rollback to avoid zombies (re-review)
- quality-eval: reject non-object verdict JSON in escalate_verdict (re-review)
- quality: address R3 review feedback
- quality: complete R3 faithful gate workflow
- resolve file-backed adapter impacts
- resolve onboarding bundle metadata
- resolve query-qualified quality targets
- reuse resolved uv in vanilla probe
- server: allow dependency-free model selections
- server: recommend bundle for full selection
- server: support cloud catalogs in bundle checks
- server: validate LightOnOCR bundle runtime
Performance Improvements
Section titled “Performance Improvements”- candle: cache exact bf16 gelu values
- candle: fuse ModernBERT exact GeGLU
- candle: optimize ModernBERT CUDA hot path
- candle: vectorize bf16 gelu lookup gate
- cloud: index usage_events_outbox (account_id, occurred_at) for the usage read (0007)
- ocr: serve generative models with SGLang
- ocr: serve generative OCR with SGLang continuous batching
v0.6.19 (2026-07-14)
Section titled “v0.6.19 (2026-07-14)”Highlights
Section titled “Highlights”- New capabilities: add catalog-smoke step — per-model serve_i6pn load verification; add env-specific Modal deploy provisioning; align prod-us manifest to latest i6pn fast-path topology; commit generated price book so the provisioner locates it; deploy-both IaC hardening — env-derived Modal env + env-targeted B6 smoke; event-driven fast-lane promotion + credit ramp
- Reliability and operations: drop the speculative mise install retry chain; harden deploy tooling — modal —json ANSI, async snapshot-commit race, Vercel sensitive-var upsert; harden lane_deploy driver — modal —json ANSI + async snapshot-commit race; retry server launch on startup death; scope the port sweep to our serve chain and harden the start guard
Features
Section titled “Features”- cloud: add catalog-smoke step — per-model serve_i6pn load verification
- cloud: add env-specific Modal deploy provisioning
- cloud: align prod-us manifest to latest i6pn fast-path topology
- cloud: commit generated price book so the provisioner locates it
- cloud: deploy-both IaC hardening — env-derived Modal env + env-targeted B6 smoke
- cloud: event-driven fast-lane promotion + credit ramp
- cloud: normalize the OSS terminal
failedreadiness state to MODEL_LOAD_FAILED on the Modal lane - cloud: productionize the i6pn realtime dispatch fast path end-to-end
- cloud: readiness-gated i6pn dispatch — thread loaded_models to the wire, gate cold workers
- model-ops: deterministic add-a-model scaffolder — verdict → YAML + matrix + target
- model-ops: deterministic add-a-model scaffolder + onboarding runbook
- model-ops: structured feasibility verdict — eval-model stage-2 gate
- model-ops: structured feasibility verdict for add-a-model stage 2
- quality-eval: add fix-proposal planner (tiered Fable-5) + experiment-plan schema
- quality-eval: bounded auto-fix — agent PRs for floor re-baselines, human merge gate (Req 14 stage 4)
- quality-eval: bounded auto-fix CLI for floor re-baseline
- quality-eval: fix-proposal planner (tiered Fable-5) + experiment-plan schema (Req14 v2 2/4)
- quality-eval: make the planner reason from signals + consume the LLM’s predicted evidence
- quality-eval: onboarding smoke playbook — served-route bring-up on Modal
- quality-eval: provenance-gated preflight + vanilla-lane floor template
- quality-eval: provenance-gated preflight + vanilla-lane floor template (Req14 v2 1/4)
Bug Fixes
Section titled “Bug Fixes”- ci: drop the speculative mise install retry chain
- ci: give the scheduled quality run its own concurrency group
- ci: stop rolling the un-retried rustup ensure in mac-smoke-nightly
- cloud: address catalog-smoke review — gate internal token, harden revoke + warm-up bound
- cloud: address CodeRabbit round-2 review on #1782
- cloud: address env deploy review comments
- cloud: address PR review comments
- cloud: bound wedged model loads + typed dead-letter on the lazy lane
- cloud: bound/adapt slab credit holds so a funded org stops self-wedging at 402
- cloud: catalog-smoke presents SIE_API_KEY when the gateway enforces managed keys
- cloud: don’t send
typeon Vercel env update (sensitive-var 400) - cloud: drop candle-only bge-large-en from the default L4 lane’s models_served
- cloud: gateway must not 401 valid keys during cold-replica key-snapshot warmup
- cloud: gen-lane warm floor + gateway-config secret merge in deploy-edge
- cloud: harden deploy tooling — modal —json ANSI, async snapshot-commit race, Vercel sensitive-var upsert
- cloud: harden lane_deploy driver — modal —json ANSI + async snapshot-commit race
- cloud: heartbeat TTL for i6pn discovery keys (corpse expiry on hard stop)
- cloud: narrow B6 local gateway bundle to the deployed models_served (#1771 hash parity)
- cloud: pg-connector smoke builds ModalLaneExecutor from SIE_LANE_APP_MAP
- cloud: pg-connector smoke builds ModalLaneExecutor from SIE_LANE_APP_MAP (#1782 missed caller)
- cloud: scope _modal_cli TERM=dumb to captured calls only (CodeRabbit #1807)
- cloud: skip B6 offload-batch in external-CP mode (unshareable payload store)
- cloud: stage refs/main alias so offline lanes resolve pinned-SHA models
- cloud: stop provisioning unbacked Stripe metered prices
- cloud: stop provisioning unbacked Stripe metered prices (prepaid-credits + Metronome rating is the model)
- cloud: tighten edge deploy conflict resolution
- cloud: Vercel env update sends value only (sensitive-var target 400)
- cloud: warm managed key before catalog-smoke sweep to beat the key-propagation race
- cloud: wire the default gen lane into the gateway census + app map
- gateway: bump yanked spin 0.10.0 to 0.10.1 in Cargo.lock
- mac-smoke nightly score-path contract + unknown-model 500s
- mac-smoke: call /v1/score with the slash-form model id
- mac-smoke: kill stale server by port between phases
- mac-smoke: retry server launch on startup death
- mac-smoke: scope the port sweep to our serve chain and harden the start guard
- model-ops: harden feasibility validator per review
- model-ops: harden scaffolder per review
- poc: pass explicit github_token to claude-code-action
- quality-eval: charset-gate vanilla_vector metric keys (review C1)
- quality-eval: classify malformed encode responses as REFUTED, not a crash
- quality-eval: controlled exit on planner output-write failure (CodeRabbit)
- quality-eval: harden autofix per CodeRabbit review round 2
- quality-eval: harden autofix preflight against output injection (review round 1)
- quality-eval: require string cell fields before charset validation
- server: colbert-small multivector encode returns one vector-set per input, worker-item order
- server: Florence-2 processor compat with transformers 4.57
- server: guard cross_encoder predict tokenization against the concurrent-metering race (class-fix for #1800)
- server: load tokenizer by pinned revision so offline chat-template rendering doesn’t need refs/main
- server: return 404 instead of 500 for unknown models on inference endpoints
- server: terminal Failed readiness + Florence-2 processor compat + pin launch-set revisions
- server: terminal Failed readiness state so the sidecar dead-letters permanently-failed model loads
v0.6.18 (2026-07-12)
Section titled “v0.6.18 (2026-07-12)”Highlights
Section titled “Highlights”- New capabilities: add flat→per-account Files-store migration; co-located AWS us-east-1 reverse-proxy edge for api.superlinked.com; consolidate data-plane deploy + manifest-drive proxy-auth/i6pn; namespace Files-store objects by account; opt-in Modal proxy-auth layer for config/lane/OTLP ingress; parameterize Modal environment + prod-us tf root by SIE env (§B2)
- Reliability and operations: fail fast on terraform plan errors before the destroy-check; bake SIE_MODAL_PROXY_AUTH into image env so the proxy-auth knob doesn’t diverge local↔remote; make proxy_auth_env authoritative in both directions; pin npm for TypeScript release publishing; aws-edge S3 backend uses native state locking (use_lockfile)
Features
Section titled “Features”- cloud: add flat→per-account Files-store migration
- cloud: co-located AWS us-east-1 reverse-proxy edge for api.superlinked.com
- cloud: consolidate data-plane deploy + manifest-drive proxy-auth/i6pn
- cloud: namespace Files-store objects by account
- cloud: opt-in Modal proxy-auth layer for config/lane/OTLP ingress
- cloud: parameterize Modal environment + prod-us tf root by SIE env (§B2)
- cloud: per-account Files-store namespacing + migration
- cloud: prod bring-up terraform prerequisites
- cloud: prod bring-up terraform prerequisites — prod-us auth-token outputs, aws-edge prod state
- cloud: prod enablement — aws-edge provisioner step, prod-us root, Modal-env parameterization
- cloud: prod-us terraform root — env-isolated control plane, separate state bucket
- cloud: reconcile OTLP transport + proxy-auth sender headers
- cloud: stand up the AWS edge via cloud-provision
- dispatcher: full IP-pin for object-store + postgres verify-full egress
- dispatcher: org-ownership assertion on upload:// ends — defense in depth behind the gateway check
- dispatcher: resolve-then-pin connector egress — SSRF TOCTOU backstop
Bug Fixes
Section titled “Bug Fixes”- ci: fail fast on terraform plan errors before the destroy-check
- ci: pin npm for TypeScript release publishing
- cloud: address CodeRabbit review on the prod-enablement PR
- cloud: aws-edge S3 backend uses native state locking (use_lockfile)
- cloud: bake SIE_MODAL_PROXY_AUTH into image env so the proxy-auth knob doesn’t diverge local↔remote
- cloud: collision-aware Files migration move
- cloud: correct console callback route + guard the CP terraform apply against silent destroys
- cloud: correct scheduler/slo-tuner DSN-secret comments + enforce non-ingress secret refs in deploy-posture audit
- cloud: default Metronome credit type to fiat USD-cents so low-balance alerts and grants provision
- cloud: gateway payload store must be a local path, not the s3 payload_store_url
- cloud: gateway payload store must be a local path, not the s3 payload_store_url (#1743 follow-through)
- cloud: gateway ships the cloud-storage backend; provisioner wires SIE_CONFIG_SERVICE_URL
- cloud: guard config-URL resolver against a malformed manifest
- cloud: harden Phase-3 CI + sibling-sweep per adversarial review (Refs #1757 #1758)
- cloud: make proxy_auth_env authoritative in both directions
- cloud: parameterize the deploy-env slug in managed Modal app names
- cloud: parameterize the OTLP collector app name with the deploy-env slug
- cloud: provision §7.6 prepaid credit-pack price + STRIPE_PRICE_ID
- cloud: redact all secret-bearing Config fields in Debug
- cloud: redact ModalProxyToken Debug + cover /gen handshake forwarding
- cloud: resolve gateway app name from gateway_app.APP_NAME, not ctx.env
- dispatcher: address CodeRabbit review on the #1742 egress pins
- dispatcher: harden connector egress pin against CodeRabbit findings
- stabilize tilt e2e regressions
- use pinned cargo-zigbuild in tilt builds
v0.6.17 (2026-07-09)
Section titled “v0.6.17 (2026-07-09)”Highlights
Section titled “Highlights”- New capabilities: add GTE RoPE embedding kernel; CLI WorkOS device-flow auth — retire the shared-token shortcut for customer commands; SIE Cloud managed service — AWS control plane, Modal data plane, provisioning + billing kernel; OSS amendments from the SIE Cloud build — engine, SDKs, gateway/sidecar seams; add deterministic regression-ledger detector for scheduled quality runs; add floor-archaeology playbook (CPU-only)
- Reliability and operations: retry model-dependent calls on MODEL_LOADING; harden regression-ledger live mode and audit input validation; harden multi-gpu worker routing; compile fused gated activation op; import CUDA pointer trait
Features
Section titled “Features”- candle: add GTE RoPE embedding kernel
- cloud: CLI WorkOS device-flow auth — retire the shared-token shortcut for customer commands
- cloud: SIE Cloud managed service — AWS control plane, Modal data plane, provisioning + billing kernel
- oss: OSS amendments from the SIE Cloud build — engine, SDKs, gateway/sidecar seams
- quality-eval: add deterministic regression-ledger detector for scheduled quality runs
- quality-eval: add floor-archaeology playbook (CPU-only)
- quality-eval: add Modal parity-probe playbook (diagnostic-only)
- quality-eval: add official-recipe, fp32-ab, config-sweep, determinism playbooks
- quality-eval: add report-only triage classifier + workflow for scheduled quality runs
- quality-eval: add validation-playbook harness + hypothesis/verdict contract
- quality-eval: report-only triage classifier + gated workflow for scheduled quality runs
- quality-eval: seed triage signature DB and short-circuit known clusters
- quality-eval: spread runner fleet across capacity pools with always-on floors
- quality-eval: validation playbooks — adversarial repro harness (Req 14 Proj 1, stage 3)
- server: support multi-gpu worker placement
- worker: route by child queue pressure
- worker: route multi-gpu sidecar children
- worker: support multi-gpu sidecar children
Bug Fixes
Section titled “Bug Fixes”- candle: compile fused gated activation op
- candle: import CUDA pointer trait
- mac-smoke: retry model-dependent calls on MODEL_LOADING
- quality-eval: address CodeRabbit + code-quality review comments
- quality-eval: address CodeRabbit review findings
- quality-eval: address triage review findings
- quality-eval: detect EBS cache volume via lsblk model string
- quality-eval: enforce budget cap on all eval playbooks + budget-stop artifacts
- quality-eval: fp32 ST arm via throwaway -st model, not a profile
- quality-eval: harden playbook verdict logic (review round 1)
- quality-eval: harden regression-ledger live mode and audit input validation
- quality-eval: harden triage input handling per CodeRabbit review
- quality-eval: keep register-runners.sh bash-3.2 compatible
- quality-eval: tolerate stop/start race in wake instance-stopped wait
- quality-triage: share data-bearing baseline discovery with triage live mode
- quality-triage: skip data-less scheduled runs in ledger baseline auto-discovery
- req6: align candle chart after rebase
- req6: harden multi-gpu worker review fixes
- req6: harden multi-gpu worker routing
- req6: resolve main rebase fallout
- req6: wire pressure-aware multi-gpu routing
- server: keep ipc ping alive on gpu health errors
- sidecar: bound queue admission through scheduler completion
- sidecar: refresh ready slots on cancel fanout failure
- sidecar: release cancelled scheduler pressure
- sidecar: wire local ingest through adapter pool
v0.6.16 (2026-07-07)
Section titled “v0.6.16 (2026-07-07)”Highlights
Section titled “Highlights”- New capabilities: opt-in template_as_prompt for ST prompt-length-aware templates; use_model_encode — delegate to checkpoint-native encode(); add Candle OOM pressure recovery; add Candle residency eviction; per-pool gpu_driver option + confidential H100 tracing example
- Reliability and operations: harden splade flash gate + add flash-vs-native parity test; splade-v3 floors were cosine-epoch artifacts — re-baseline to dot self-measures + harden flash gate; restore executable bit on mac-smoke mise task; faithful ModernBERT forward + hardened 1_Dense loading; restore e5 query/passage templates on the multilingual-e5-large sentence_transformer profile
Features
Section titled “Features”- adapters: opt-in template_as_prompt for ST prompt-length-aware templates
- adapters: use_model_encode — delegate to checkpoint-native encode()
- add Candle OOM pressure recovery
- add Candle residency eviction
- terraform/azure: per-pool gpu_driver option + confidential H100 tracing example
Bug Fixes
Section titled “Bug Fixes”- adapters: harden splade flash gate + add flash-vs-native parity test
- adapters: honor caller instruction in use_model_encode without a template
- address candle catalog review comments
- address Candle residency review comments
- align Candle idle eviction lifecycle
- align candle workers with catalog lane config
- bench: cap docling pubtabnet-html at 2000 pages to fit the job budget
- bench: splade-v3 floors were cosine-epoch artifacts — re-baseline to dot self-measures + harden flash gate
- bundles: pin flash-attn-4 prerelease so sglang 0.5.10.post1 resolves
- ci: restore executable bit on mac-smoke mise task
- colbert_modernbert_flash: apply trained ColBERT Dense projection (+ guard, tokenizer, re-baseline)
- colbert_modernbert_flash: faithful ModernBERT forward + hardened 1_Dense loading
- colbert: GTE muvera also at simhash 6 - CMedQA/MMarco evals OOM at dim 81920
- colbert: mxbai muvera at simhash 6 (FDE dim 20480) - runner dies at 81920
- colbert: retune GTE/mxbai muvera for the Dense-projected representation
- colbert: serve pylate Dense chains and align fallback with flash
- deploy: cap tester l4 lanes
- deploy: zero and cap tester L4 lanes
- deploy: zero AWS worker warm floor
- deploy: zero tester worker warm lanes
- gateway: fan out gpu-agnostic demand across pool profiles for scale-from-zero
- gateway: normalize explicit GPU demand labels
- helm: wake a deterministic owner lane on gpu-agnostic demand for multi-profile bundles
- models: NV-Embed-v2 — serve via checkpoint-native encode() (latent-attention recipe), NFCorpus 0.3715→0.4479
- models: restore e5 query/passage templates on the multilingual-e5-large sentence_transformer profile
- models: restore e5 templates on multilingual-e5-large ST profile
- models: serve NV-Embed-v2 via its native latent-attention ST path
- preserve Candle idle evictor failure logs
- quality-eval: fail fast when sie-server dies during readiness wait
- quality-eval: unbreak 7B sglang lanes (flash-attn-4 resolution) + fail-fast readiness + docling job budget
- restore kind smoke candle routing
v0.6.15 (2026-07-03)
Section titled “v0.6.15 (2026-07-03)”Highlights
Section titled “Highlights”- New capabilities: add Candle profile routing and catalog models; add candle runtime diagnostics; add candle xlm-roberta flash attention; add dev AMI bake script; add dev ec2 cleanup workflow; add dev EC2 launch script
- Reliability and operations: harden native worker routing; harden native worker runtime; restore sidecar model readiness; delete Python NATS pull loop from worker hot path; harden SIE dev AMI workflow
- Performance: add GEMM diagnostics switch; fuse xlm-r qkv projection; optimize bge-m3 xlm-r runtime path; enable candle reduced precision gemms
Features
Section titled “Features”- add Candle profile routing and catalog models
- add candle runtime diagnostics
- add candle xlm-roberta flash attention
- add dev AMI bake script
- add dev ec2 cleanup workflow
- add dev EC2 launch script
- add SIE API key rotation tooling
- candle: add native adapter runtime
- dev: add personal SIE CUDA/Rust AMI workflow
- document ssm ssh dev launcher
- gateway: count dropped stale NAKs from abandoned attempts
- helm: add opt-in OTLP tracing wiring + bundled collector to sie-cluster
- helm: add optional bundled Tempo tracing backend
- helm: add SIE tracing dashboard and tempo datasource uid
- helm: opt-in OTLP tracing wiring + bundled collector in sie-cluster
- helm: optional bundled Tempo tracing backend + Grafana datasource for sie-cluster
- improve candle replica concurrency
- improve dev AMI cache scalability
- observability: bound tracing shutdown + trim env inputs across all four runtimes
- observability: bound tracing shutdown and trim env inputs across all four runtimes
- observability: gate Rust tracing on SIE_TRACING_ENABLED
- observability: unify tracing enable switch across gateway, sidecar, and worker
- quality-eval: gate score-on-retrieval candidate cells once floor committed
- reattach dev cache volume
- server: land Apple Silicon as a primary device onto main
- server: land Apple Silicon as a primary device onto main (re-land of #1455)
- sie_gateway: emit gateway.publish span on non-streaming queue-publish path
- sie_server_rust: add OTLP distributed tracing
- sie_server_rust: add OTLP distributed tracing to the Rust native worker
- terraform: dedicated single-AZ node group for stateful observability
- tilt: enable tracing with bundled Tempo
- tracing: emit sidecar.dispatch span from sie_server_sidecar
- tracing: inject queue trace context for non-generation endpoints
- tracing: instrument sie_server_sidecar with OTLP (sidecar.dispatch span)
- tracing: propagate trace context through the non-streaming worker loop
- tune candle forward concurrency
- tune dev EC2 launch defaults
- wire dev AMI cache disk helpers
Bug Fixes
Section titled “Bug Fixes”- #1340: stella per-task instruction + re-baseline stale dense/colbert floors (blast radius was #1432)
- adapters: guard empty-document MaxSim batches; strengthen single-doc test
- address API key review feedback
- address API key rollback review
- address rust worker smoke review comments
- address sidecar PR review feedback
- align rust worker kind smoke wiring
- align rust worker warm smoke behavior
- align worker pool config hashes
- back up API key secret before apply
- candle: align cublaslt with candle cuda stream
- candle: align readiness with loaded model state
- candle: copy vendored cuda dependency in image
- candle: harden native worker routing
- candle: harden native worker runtime
- candle: import cublaslt cuda extension
- candle: keep catalog additions candle-only
- candle: publish loaded models in health
- candle: repair layer norm cuda lifetimes
- candle: restore sidecar model readiness
- candle: run bge-m3 profile in bf16
- candle: vendor cublaslt for cuda build
- candle: vendor layer norm cuda dependency
- clean stale Python queue ownership drift
- complete empty preformed worker requests
- compress dev ami user data
- config: defer cancellation after config commit starts
- config: handle idempotency in-flight races
- config: shield idempotency failure cleanup
- copy rust worker vendor deps before docker cook
- dashboard: harden run_status backfill per review feedback
- dashboard: keep cancelled nightlies off the Quality on main widget
- delete Python NATS pull loop from worker hot path
- dev: address dev AMI review hardening
- dev: harden SIE dev AMI workflow
- dev: persist Codex state on dev cache disk
- docs: drop otel.name from span-attribute table; tighten tracing dashboard test
- export profile bundle targets
- gateway: append forwarded embeddings headers to preserve multi-value tracestate
- gateway: cancel abandoned direct batch fallback work
- gateway: don’t report a succeeded SSE generation as inter_chunk_timeout on consumer lag
- gateway: drop stale NAKs from abandoned attempts in handle_nak
- gateway: forward traceparent/tracestate through /v1/embeddings rewrite
- gateway: harden batch direct fallback cancellation
- gateway: ignore pre-latch stale NAKs after republish
- gateway: ignore quick-xml <0.41 RustSec advisories in cargo-deny
- gateway: keep signalling SSE incompleteness on consumer lag; skip redundant cancel
- gateway: keep signalling SSE incompleteness on lag; only skip cancel when done
- gateway: record demand on backpressure so KEDA scales the lane
- gateway: route engine-pin errors through endpoint_error_response
- gateway: separate queue timeouts from model loading
- gateway: use real RUSTSEC id for second quick-xml advisory
- harden candle bge-m3 runtime bring-up
- harden dev EC2 launcher dry-run
- harden dev profile home fallback
- harden SIE API key tooling
- helm: make bundled collector downstream OTLP TLS configurable
- helm: set OTEL_EXPORTER_OTLP_PROTOCOL=grpc on traced pods
- helm: suppress bundled collector when explicit tracing endpoint set
- improve candle embedding runtime throughput
- install gateway pool config hashes
- install gpg agent for dev AMI bake
- keep dev AMI user data under limit
- log candle batch failure context
- models: load stock tokenizer for gte-Qwen2-7B-instruct
- models: restore per-task MTEB instruction for stella_en_v5
- models: serve embeddinggemma-300m via SentenceTransformerDenseAdapter
- models: serve gte-Qwen2-1.5B via faithful sentence_transformer adapter
- models: serve gte-Qwen2-1.5B via sentence_transformer adapter
- muvera: center tokens before SimHash to fix answerai @muvera Blackwell collapse
- muvera: drop same-dim count-sketch + retune colbert muvera configs
- muvera: drop same-dim count-sketch, retune colbert configs + floors
- persist dev cache volume
- persist dev kubeconfig
- playground: keyboard-accessible tabs + associated form labels
- polish rollback backup examples
- preserve candle compute cap diagnostics
- preserve control-plane bundle hash in sidecar
- quiet dev AMI tool installs
- remove candle replica fallback path
- remove Python worker pull-loop path
- repair candle ci routing checks
- repair candle rebase fallout
- repeat dev AMI bake markers
- require explicit API key rollback path
- resolve latest dev AMI base
- respect hf cache env in rust worker
- restore default candle cuda stream
- route kind rust worker build through docker task
- sdk: raise actionable error on encode result-count desync
- select rustls provider at binary startup
- select rustls provider for rust services
- server: /v1/embeddings maps OOM to 503 RESOURCE_EXHAUSTED, not 500
- server: address CI + review comments on the Apple Silicon re-land
- server: apply profile runtime options + append EOS on cluster encode path
- server: apply profile runtime options and append EOS on cluster encode path
- server: bound continuous-batch drain so a busy LoRA can’t starve siblings
- server: bound the continuous-batching drain so a busy LoRA can’t starve siblings
- server: cap LoRA drain to snapshot backlog so post-snapshot arrivals don’t overshoot
- server: make SGLang load eviction headroom-aware
- server: move load headroom behind adapter contract
- server: serialize hot reload registry mutations
- server: skip EOS on empty input; cover non-default profile on queue path
- server: stop mislabelling uint8 embeddings as “binary” on the queue path
- server: stop mislabelling uint8 embeddings as binary on the queue path
- server: unregister from MemoryManager even if adapter.unload() raises
- server: unregister model from MemoryManager even if adapter.unload() raises
- set home during dev AMI bake
- sidecar: bound scheduler bulk enqueue waves
- sidecar: keep only config hash wiring
- sidecar: start batch cancel before direct pulls
- sie_server_rust: only log OTLP init success when exporter built
- stabilize rust worker kind smoke
- terraform: address CodeRabbit review on obs node group
- terraform: auto-derive observability node group AMI from instance arch
- terraform: make observability node group ami_type/disk_type configurable
- tighten rust worker smoke diagnostics
- tilt: use single string for candle-helm-smoke cmd (Starlark has no implicit concat)
- tilt: wait for complete trace before asserting span chain
- tolerate missing zstd devel package
- tracing: assert items/metadata alignment before zip in scheduler batch
- tracing: attach inbound trace context for non-generation publishes
- tracing: parent sidecar.dispatch on first valid inbound context
- verify dev AMI bake via ssm
Performance Improvements
Section titled “Performance Improvements”- candle: add GEMM diagnostics switch
- candle: fuse xlm-r qkv projection
- candle: optimize bge-m3 xlm-r runtime path
- enable candle reduced precision gemms
- fuse candle xlm-roberta qkv projection
- gateway: build the HRW ring from the by_bundle index, not a whole-cluster scan
- improve candle worker concurrency and build cache
- restore candle xlm-roberta qkv layout
- server: vectorize bge_m3_flash CLS gather to drop N per-encode GPU syncs
- sidecar: bulk enqueue scheduler batches
Reverts
Section titled “Reverts”- restore gateway sidecar rustls defaults
v0.6.14 (2026-06-26)
Section titled “v0.6.14 (2026-06-26)”Highlights
Section titled “Highlights”- Reliability and operations: unblock ColBERT score-on-retrieval — score_pairs 500 + muvera dense-encode + gate co-serve; include sie-bench in image context
Bug Fixes
Section titled “Bug Fixes”- #1430: unblock ColBERT score-on-retrieval — score_pairs 500 + muvera dense-encode + gate co-serve
- docker: include sie-bench in image context
v0.6.13 (2026-06-25)
Section titled “v0.6.13 (2026-06-25)”Highlights
Section titled “Highlights”- New capabilities: publish sie-bench image to GHCR on release; per-tool GPU routing and env-overridable GLiNER models
- Reliability and operations: extend tester structured generation timeouts; retry failed payload cleanup deletes; align sie-test qwen27 lane routing; build sie-bench image CPU-only with headless opencv; tolerate cold-start provisioning errors in scale-from-zero loadtests
Features
Section titled “Features”- ci: publish sie-bench image to GHCR on release
- sie-mcp: per-tool GPU routing and env-overridable GLiNER models
Bug Fixes
Section titled “Bug Fixes”- align sie-test qwen27 lane routing
- bench: address CodeRabbit review (non-root image, clearer comment, checkout creds)
- bench: build sie-bench image CPU-only with headless opencv
- bench: tolerate cold-start provisioning errors in scale-from-zero loadtests
- extend tester structured generation timeouts
- gateway: dedupe offloaded payload cleanup keys
- gateway: retry failed payload cleanup deletes
- gateway: track exact offloaded payload keys
- loadtest: tolerate cold-start provisioning errors + publish sie-bench to GHCR
- persist sie-test qwen warm lanes
- scale sie-test spot lanes to zero
- server: drain removed profile variants during config apply
- server: preflight profile variant config updates
- server: preserve profile variants during config reconcile
- server: serialize async config registry updates
- server: tighten profile variant reconciliation
- sie-mcp: chunk GLiNER extract/redact for large docs; flag the other 4B
v0.6.12 (2026-06-24)
Section titled “v0.6.12 (2026-06-24)”Highlights
Section titled “Highlights”- New capabilities: enable gateway-API model pinning on sie-test; add per-profile quality floors for bge-m3 sparse/multivector
- Reliability and operations: exclude per-profile target.<profile>.json floors from result loaders; run docker task in project env; skip payload-store cleanup DELETEs for inline requests; guard label-window overflow + defer LEDGAR for v1.0 uni-encoders; gate subset-parameterized tasks; keep skipping candidate-gen dirs
- Performance: release MaxSim per-batch intermediates to cap peak memory
Features
Section titled “Features”- deploy: enable gateway-API model pinning on sie-test
- quality-eval: add per-profile quality floors for bge-m3 sparse/multivector
Bug Fixes
Section titled “Bug Fixes”- bench: exclude per-profile target.<profile>.json floors from result loaders
- ci: run docker task in project env
- deploy: address CodeRabbit review on sie-test pinning
- gateway: record offload before the store write (PR review)
- gateway: skip payload-store cleanup DELETEs for inline requests
- gliclass: guard label-window overflow + defer LEDGAR for v1.0 uni-encoders
- quality-eval: gate subset-parameterized tasks; keep skipping candidate-gen dirs
- sie_bench: batch-local query padding in MaxSim to fix #1461 OOM
Performance Improvements
Section titled “Performance Improvements”- sie_bench: release MaxSim per-batch intermediates to cap peak memory
v0.6.11 (2026-06-23)
Section titled “v0.6.11 (2026-06-23)”Highlights
Section titled “Highlights”- New capabilities: floor Florence-2-base on COCO detection (AP 0.2785); floor Florence-2-base on olmOCR-bench (accuracy 0.0824); fail fast when a bundle is placed on the wrong platform; add Gemma 4 generation models (E2B, E4B, 26B-A4B); first-class cuda13 platform for the gemma bundle; first-class cuda13 platform for the gemma bundle (build/CI/release/deploy)
- Reliability and operations: raise uv HTTP timeout for the gemma cu130 wheel downloads; harden logical pool queue routing; harden worker-direct batch fallback; return retryable model loading for generation; gate mteb task instruction to instruction-following models
Features
Section titled “Features”- bench: floor Florence-2-base on COCO detection (AP 0.2785)
- bench: floor Florence-2-base on olmOCR-bench (accuracy 0.0824)
- deploy: fail fast when a bundle is placed on the wrong platform
- server: add Gemma 4 generation models (E2B, E4B, 26B-A4B)
- server: first-class cuda13 platform for the gemma bundle
- server: first-class cuda13 platform for the gemma bundle (build/CI/release/deploy)
- server: keep pinned models loaded and exclude them from LRU eviction
- support logical pools backed by queue pools
- terraform: add marci-dev Azure A10 single-node smoke-test example
- terraform: Azure marci-dev A10 smoke-test example + node-subnet NSG LB fix
Bug Fixes
Section titled “Bug Fixes”- align retryable generation review feedback
- bench: gate mteb task instruction to instruction-following models
- bench: gate MTEB task instructions to instruction-following models
- bench: resolve profile output_types from adapter_options.runtime for bge-m3 sparse/multivector evals
- ci: raise uv HTTP timeout for the gemma cu130 wheel downloads
- ci: run the gemma cuda13 build smoke + cover it in release verify/stamp/warm
- disable payload store in kind smoke
- handle sidecar progress ack failures
- harden logical pool queue routing
- harden worker-direct batch fallback
- loadtest: disable payload store in nightly deploy
- merge runtime pool capacity changes
- preserve queue delivery budget
- return retryable model loading for generation
- server: correct stale cuda13/gemma platform references
- server: make docker bake/verify platform-aware (cuda13/gemma in mixed builds)
- server: make transformers version bundle-controlled
- terraform: allow public LoadBalancer/ingress inbound on AKS node subnet NSG
- terraform: clear internal mise references from public Azure module
- terraform: clear internal mise/tooling references from public AWS and GCP modules
- terraform: drop internal references from public Azure module README and variables
v0.6.10 (2026-06-22)
Section titled “v0.6.10 (2026-06-22)”Highlights
Section titled “Highlights”- New capabilities: floor measure-first dark models, retire donut-rvlcdip; provision and enable the payload store by default; fix <OD> bbox conversion, re-task base-ft/large to COCO detection
- Reliability and operations: align gliner label inputs; align gliner v2.5 conll protocol; improve generation failure observability; coalesce whitespace-only instruction to default; pool post-RMSNorm last_hidden_state
Features
Section titled “Features”- bench: floor measure-first dark models, retire donut-rvlcdip
- deploy: provision and enable the payload store by default
- florence2: fix <OD> bbox conversion, re-task base-ft/large to COCO detection
Bug Fixes
Section titled “Bug Fixes”- bench: align gliner label inputs
- bench: align gliner v2.5 conll protocol
- deploy: address payload-store review feedback
- improve generation failure observability
- qwen3_vl_embedding: coalesce whitespace-only instruction to default
- qwen3_vl_embedding: pool post-RMSNorm last_hidden_state
- sie_bench: coerce ClassLabel int names to str in classification loader
- sie_bench: repoint extract-kie FUNSD to live nielsr/funsd dataset
- sie-cluster: make the self-host deploy skill robust to its own runtime
- sie-server: make transformers5 bundle reachable via pip
- sync generation visibility api contract
v0.6.9 (2026-06-19)
Section titled “v0.6.9 (2026-06-19)”Highlights
Section titled “Highlights”- New capabilities: add per-pool pinned-model set to the pool config API; per-pool pinned-model set in the pool config API; support profile-qualified ids in per-pool pinned-model set; enable S3 payload store for >1MB work items
- Reliability and operations: bound pinned-model metric labels and add handler validation tests; scope worker preload models by lane; alias PEFT LoRA adapter names; expose model_cache_bucket_url as a root output
Features
Section titled “Features”- gateway: add per-pool pinned-model set to the pool config API
- gateway: per-pool pinned-model set in the pool config API
- gateway: support profile-qualified ids in per-pool pinned-model set
- tester-cluster: enable S3 payload store for >1MB work items
Bug Fixes
Section titled “Bug Fixes”- gateway: bound pinned-model metric labels and add handler validation tests
- scope worker preload models by lane
- server: alias PEFT LoRA adapter names
- tester-cluster: expose model_cache_bucket_url as a root output
v0.6.8 (2026-06-16)
Section titled “v0.6.8 (2026-06-16)”Highlights
Section titled “Highlights”- Reliability and operations: harden release artifact workflows; keep score traffic from idle-evicting rerankers; serialize worker use with unload state; share score media batch cost
Bug Fixes
Section titled “Bug Fixes”- ci: harden release artifact workflows
- server: keep score traffic from idle-evicting rerankers
- server: serialize worker use with unload state
- server: share score media batch cost
v0.6.7 (2026-06-16)
Section titled “v0.6.7 (2026-06-16)”Highlights
Section titled “Highlights”- New capabilities: answer_questions transient-QA tool (Req 12 #1309)
- Reliability and operations: guard model ready timeout against liveness budget; wire model ready timeout through helm; add Qwen3.6 27B 32K RTX serving profile; route grammar requests to a non-speculative profile (NEXTN bypasses Outlines FSM); make worker config reconciliation no-op safe
- Performance: extract shared vision patch-embed rebind helper; batch generate() across pages to close #601 L4 throughput gap
Features
Section titled “Features”- mcp: answer_questions transient-QA tool (Req 12 #1309)
Bug Fixes
Section titled “Bug Fixes”- add Qwen3.6 27B 32K RTX serving profile
- address qwen 27b review feedback
- generate: route grammar requests to a non-speculative profile (NEXTN bypasses Outlines FSM)
- guard model ready timeout against liveness budget
- make worker config reconciliation no-op safe
- route qwen 27b fp8 cold starts
- wire model ready timeout through helm
Performance Improvements
Section titled “Performance Improvements”- adapters: extract shared vision patch-embed rebind helper
- lighton_ocr: batch generate() across pages to close #601 L4 throughput gap
v0.6.6 (2026-06-14)
Section titled “v0.6.6 (2026-06-14)”Highlights
Section titled “Highlights”- Reliability and operations: align pool-scoped bundle hashes; avoid sticky missing bundle hashes; clarify missing profile inheritance; fail closed on missing bundle metadata; stabilize keda all-marker e2e
Bug Fixes
Section titled “Bug Fixes”- config: align pool-scoped bundle hashes
- config: avoid sticky missing bundle hashes
- config: clarify missing profile inheritance
- config: fail closed on missing bundle metadata
- tilt: stabilize keda all-marker e2e
v0.6.5 (2026-06-13)
Section titled “v0.6.5 (2026-06-13)”Highlights
Section titled “Highlights”- New capabilities: demand-side token-reduction benchmark (Req 12, #1311); add describe_image tool (caption + zero-shot tags); add describe_image tool (caption + zero-shot tags) — Req 12 #1310; cap describe_image payload size before cluster calls; claude.ai connector surface — OAuth bridge + skill ZIP (Req 12 #1312); sie_mcp edge with docs_to_markdown tool (Req 12 #1306)
- Reliability and operations: harden replace snapshot IPC; return retryable OpenAI provisioning errors; bind OAuth authorization codes to client_id; doctor classifies probe read-timeouts as cold, not unreachable; align GPU memory pressure defaults
- Performance: skip tag embedding when top_k <= 0
Features
Section titled “Features”- bench: demand-side token-reduction benchmark (Req 12, #1311)
- mcp: add describe_image tool (caption + zero-shot tags)
- mcp: add describe_image tool (caption + zero-shot tags) — Req 12 #1310
- mcp: cap describe_image payload size before cluster calls
- mcp: claude.ai connector surface — OAuth bridge + skill ZIP (Req 12 #1312)
- mcp: sie_mcp edge with docs_to_markdown tool (Req 12 #1306)
- mcp: structured extraction + structured generation tools (Req 12 #1308)
- mcp: wire measured token-reduction figures into savings metadata
- tools: add sie doctor — per-capability cluster diagnostics
- tools: Florence-2 fallback for image OCR
- tools: sie_tools — Claude Code context-offload client for managed clusters
Bug Fixes
Section titled “Bug Fixes”- align GPU memory pressure defaults
- config: detect bundle config hash drift
- config: fingerprint model pool ownership
- config: harden replace snapshot IPC
- config: replace drifted export snapshots
- deps: bump sidecar prometheus for protobuf advisory
- gateway: address provisioning review feedback
- gateway: align provisioning contract docs
- gateway: decode native media JSON bytes
- gateway: dereference structured output schema refs
- gateway: make provisioning non-2xx universally
- gateway: preserve ref sibling schema semantics
- gateway: return retryable OpenAI provisioning errors
- mcp: address review feedback on structured tools
- mcp: bind OAuth authorization codes to client_id
- mcp: blank-env fallback for model ids; honor SIE_MCP_IMAGE_TOP_K=0
- mcp: deep-copy committed token-reduction figures in build_metadata
- mcp: validate embedding shapes in _top_k_tags
- sdk: normalize score image payloads for wire transport
- sdk: normalize score images for wire transport
- server: guard readiness for removed configs
- server: honor pool-aware model configs
- server: render qwen3 vl reranker document images in user prompt
- server: render Qwen3-VL reranker document images in user prompt
- sie-cluster: add spot toleration to AKS worker pool
- tools: address doctor review feedback
- tools: doctor classifies probe read-timeouts as cold, not unreachable
- worker: keep SGLang loads off event loop
Performance Improvements
Section titled “Performance Improvements”- mcp: skip tag embedding when top_k <= 0
v0.6.4 (2026-06-11)
Section titled “v0.6.4 (2026-06-11)”Highlights
Section titled “Highlights”- New capabilities: add
grant_admin_to_creator(opt-in AAD-RBAC for caller); lock model-cache storage account to cluster VNet by default; install kubelogin and convert kubeconfig after AKS get-credentials; wire Azure provider tooling; add azure (AKS) terraform module; ship values-aks.yaml AKS overlay with the Azure module - Reliability and operations: harden storage_allowed_ip_ranges CIDR validation; harden release guarded merge checks; resubscribe stale NATS health stream; emit
az aks get-credentials --overwrite-existing; drop unreachable final_registry guard so ACR path can fire
Features
Section titled “Features”- azure-terraform: add
grant_admin_to_creator(opt-in AAD-RBAC for caller) - azure-terraform: lock model-cache storage account to cluster VNet by default
- cluster: install kubelogin and convert kubeconfig after AKS get-credentials
- cluster: wire Azure provider tooling
- deploy: add azure (AKS) terraform module
- helm: ship values-aks.yaml AKS overlay with the Azure module
Bug Fixes
Section titled “Bug Fixes”- azure-terraform: emit
az aks get-credentials --overwrite-existing - azure-terraform: harden storage_allowed_ip_ranges CIDR validation
- ci: harden release guarded merge checks
- cluster: address review feedback on Azure provider wiring
- cluster: drop unreachable final_registry guard so ACR path can fire
- cluster: set TF_VAR_* on Azure destroy path (same as create)
- deploy: revert system pool default to Standard_D4s_v3 (zoned everywhere)
- gateway: resubscribe stale NATS health stream
- sidecar: preserve msgpack work item payloads
v0.6.3 (2026-06-10)
Section titled “v0.6.3 (2026-06-10)”Highlights
Section titled “Highlights”- New capabilities: add azure blob payload store support; add server-side copy fast path for cloud weight sync; informational generation eval CI gate over committed floors; add vision (image) input to generate(); preserve text/image content-part ordering; vision (image) input for generate()
- Reliability and operations: harden cloud cache sync paths; clear HIGH Dependabot alerts (docling, rustls-webpki); ensure cloud weight sync creates local parents; evict stale gateway workers on shutdown; fall back to relay on S3/GCS server-side copy failure
- Performance: engage conformant image preprocessing for v1; engage conformant image preprocessing for v1 (1.8x)
Features
Section titled “Features”- add azure blob payload store support
- add server-side copy fast path for cloud weight sync
- bench: informational generation eval CI gate over committed floors
- generate: add vision (image) input to generate()
- generate: preserve text/image content-part ordering
- generate: vision (image) input for generate()
- support azure blob cluster cache
- tester-cluster: rtx6000 g7e.4xlarge + sglang preload + hf-token wiring
Bug Fixes
Section titled “Bug Fixes”- address azure cache review feedback
- address final cloud storage review issues
- bench: harden generation eval gate per review
- deps: clear HIGH Dependabot alerts (docling, rustls-webpki)
- ensure cloud weight sync creates local parents
- evict stale gateway workers on shutdown
- fall back to relay on S3/GCS server-side copy failure
- generate: address CodeRabbit review on vision input
- generate: address huronat review on vision input (F2-F8)
- generate: image-free content_parts field must not shadow layout
- generate: reject both-present image-bearing content layouts
- harden cloud cache sync paths
- loadtest-ci: self-heal orphaned cluster + stale lock in preflight
- nemo_colembed: trim left-padding rows from v1 conformant doc embeddings
- normalize local weight sync destination
- quality-adapter: gate v1 Vidore3 on English; finalize ?lang= plumbing
- skill: add bash language tag to hfCache —set fenced block (MD040)
- skill: move inline comments off shell continuation lines so the helm snippet pastes cleanly
- support cloud source weight sync
- tester-cluster: update rtx6000-spot machineType doc to g7e.4xlarge to match terraform
Performance Improvements
Section titled “Performance Improvements”- nemo_colembed: engage conformant image preprocessing for v1
- nemo_colembed: engage conformant image preprocessing for v1 (1.8x)
v0.6.2 (2026-06-08)
Section titled “v0.6.2 (2026-06-08)”Highlights
Section titled “Highlights”- New capabilities: defer sie-config NATS startup and honor log levels; refresh KEDA Tilt local dev branch; M4 dense encoders — mxbai-embed-large-v1, arctic-embed-l-v2.0, modernbert-embed-base; add daily guarded stable releases
- Reliability and operations: accept dense dim in qwen3 vl embedding adapter; preserve model query templates in mteb eval; scale single-profile bundles on gpu-agnostic demand; consolidate runtime ninja install; install ninja in cuda runtime
Features
Section titled “Features”- defer sie-config NATS startup and honor log levels
- dev: refresh KEDA Tilt local dev branch
- models: M4 dense encoders — mxbai-embed-large-v1, arctic-embed-l-v2.0, modernbert-embed-base
- release: add daily guarded stable releases
Bug Fixes
Section titled “Bug Fixes”- accept dense dim in qwen3 vl embedding adapter
- bench: preserve model query templates in mteb eval
- dev: address KEDA Tilt PR review
- helm: scale single-profile bundles on gpu-agnostic demand
- server: consolidate runtime ninja install
- server: install ninja in cuda runtime
- server: install ninja in CUDA SGLang runtime
- terraform: deny non-HTTPS access on state and quality-eval S3 buckets
v0.6.1 (2026-06-07)
Section titled “v0.6.1 (2026-06-07)”Highlights
Section titled “Highlights”- New capabilities: configure GPU disk sizing and generate smoke; support static queue pools
- Reliability and operations: fail fast on invalid static pool config; pin kind smoke workers to default queue pool; canonicalize static queue pool names; stabilize GPU disk Terraform test
Features
Section titled “Features”- configure GPU disk sizing and generate smoke
- gateway: support static queue pools
Bug Fixes
Section titled “Bug Fixes”- address GPU disk review comments
- ci: pin kind smoke workers to default queue pool
- gateway: canonicalize static queue pool names
- gateway: fail fast on invalid static pool config
- stabilize GPU disk Terraform test
v0.6.0 (2026-06-07)
Section titled “v0.6.0 (2026-06-07)”Highlights
Section titled “Highlights”- Breaking change: Queue work subjects and pool streams use the new sie.work.{pool}.{machine_profile}.{bundle}.{model} shape only; legacy subject filters are intentionally not preserved.; workers will subscribe to
sie.work.*.<poolName>instead ofsie.work.*.default. Deployed alone (without the matching gateway/sidecar update that publishes/filters on the new subject) this will break routing on every cluster. To preserve the old shared-queue behavior, setworkers.common.queuePool: "default"explicitly. - New capabilities: route work by queue pool lanes; default SIE_POOL to pool name (not “default”)
- Reliability and operations: harden queue lane admission; bump vitest 2.1.9 -> 4.1.0 (CVE-2026-47429); align lane defaults and tilt e2e; preserve worker-group queue defaults
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- gateway: Queue work subjects and pool streams use the new sie.work.{pool}.{machine_profile}.{bundle}.{model} shape only; legacy subject filters are intentionally not preserved.
- helm: workers will subscribe to
sie.work.*.<poolName>instead ofsie.work.*.default. Deployed alone (without the matching gateway/sidecar update that publishes/filters on the new subject) this will break routing on every cluster. To preserve the old shared-queue behavior, setworkers.common.queuePool: "default"explicitly.
Features
Section titled “Features”- gateway: route work by queue pool lanes
- helm: default SIE_POOL to pool name (not “default”)
Bug Fixes
Section titled “Bug Fixes”- deps: bump vitest 2.1.9 -> 4.1.0 (CVE-2026-47429)
- gateway: harden queue lane admission
- helm: align lane defaults and tilt e2e
- helm: preserve worker-group queue defaults
v0.5.0 (2026-06-04)
Section titled “v0.5.0 (2026-06-04)”Highlights
Section titled “Highlights”- Breaking change:
workers.pools.<name>.bundle(string),workers.pools.<name>.minReplicas,workers.pools.<name>.maxReplicas,workers.pools.<name>.extraEnv, andworkers.pools.<name>.imageBundleare replaced byworkers.pools.<name>.bundles.<bundle>.{minReplicas, maxReplicas, extraEnv, imageBundle, enabled}.workers.common.bundleis removed (no longer consumed). StatefulSet, ScaledObject, PDB, and image-prepull DaemonSet names change fromworker-{pool}toworker-{pool}-{bundle}, so in-place upgrades require deleting the old resources first. - New capabilities: agent-jobs text-gen readiness — code/SQL/tools/guard evals + Qwen3.6-27B + precision routing; transfer sie-cluster claude skill; P(unsafe) logprob threshold for CHECK POLICY precision; split worker pools into pool × bundles schema; surface code/sql/guard capabilities; resolve job aliases in configs/resolve; add sglang worker pool for generative models
- Reliability and operations: expose unauthenticated metrics scrape port; expose unauthenticated metrics scrape port safely for prom; preserve gateway metrics scrape labels; drop unsupported ebnf advertisement + restore guardian a100 guard threshold; fail-fast on missing Spider DBs + order-sensitive SQL exec accuracy
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- helm:
workers.pools.<name>.bundle(string),workers.pools.<name>.minReplicas,workers.pools.<name>.maxReplicas,workers.pools.<name>.extraEnv, andworkers.pools.<name>.imageBundleare replaced byworkers.pools.<name>.bundles.<bundle>.{minReplicas, maxReplicas, extraEnv, imageBundle, enabled}.workers.common.bundleis removed (no longer consumed). StatefulSet, ScaledObject, PDB, and image-prepull DaemonSet names change fromworker-{pool}toworker-{pool}-{bundle}, so in-place upgrades require deleting the old resources first.
Features
Section titled “Features”- agent-jobs text-gen readiness — code/SQL/tools/guard evals + Qwen3.6-27B + precision routing
- agents: transfer sie-cluster claude skill
- guard: P(unsafe) logprob threshold for CHECK POLICY precision
- helm: split worker pools into pool × bundles schema
- models: surface code/sql/guard capabilities; resolve job aliases in configs/resolve
- tester-cluster: add sglang worker pool for generative models
Bug Fixes
Section titled “Bug Fixes”- agents: address sie cluster review comments
- bench: fail-fast on missing Spider DBs + order-sensitive SQL exec accuracy
- gateway: expose unauthenticated metrics scrape port
- gateway: expose unauthenticated metrics scrape port safely for prom
- guard: reject multi-candidate sampling + keep logprobs consistent on rewrite
- guard: robust verdict thresholding, logprob hygiene, decoded-token logprobs
- helm: fail-fast on missing/invalid bundle replica bounds
- helm: preserve gateway metrics scrape labels
- helm: use sidecar binary for image pre-pull
- models: drop unsupported ebnf advertisement + restore guardian a100 guard threshold
- sie_server: honor params.instruction in Florence-2 extract
- tester-cluster: cap rtx6000 default bundle to avoid over-subscription
- tools: via-SIE EBNF response_format shape + request/preload model split
v0.4.2 (2026-06-03)
Section titled “v0.4.2 (2026-06-03)”Highlights
Section titled “Highlights”- New capabilities: 5-domain generation bench + via-sie quality matrix + gateway schema gaps; add e0-02 all-minilm time-share experiment; land coalesce_ms=5 + max_batch_requests=12 as Rust defaults; add —via-sie smoke path (route through sie_server); add min_tokens + system_prompt + temperature for G4 retry; close Qwen3.6-27B gap — min_tokens=10 + max=768 + ctx=4096
- Reliability and operations: set verbose=True on SIEServer so launch errors surface; document worker-sidecar metrics wiring; gate sidecar nats reconnect refresh; harden sidecar config recovery; budget loadtest barrier timeouts
- Performance: anchor min_batch_cost floor at max_batch_tokens // 4; tighten adaptive wait ceiling + revert gte-multilingual 32k; rebind vision Conv3d patch-embed to F.linear; raise max_batch_tokens 16k → 32k to stop IPC-batch shred
Features
Section titled “Features”- 5-domain generation bench + via-sie quality matrix + gateway schema gaps
- add e0-02 all-minilm time-share experiment
- batch_config: land coalesce_ms=5 + max_batch_requests=12 as Rust defaults
- bench-27b: add —via-sie smoke path (route through sie_server)
- bench-27b: add min_tokens + system_prompt + temperature for G4 retry
- bench-27b: close Qwen3.6-27B gap — min_tokens=10 + max=768 + ctx=4096
- bench-27b: launch full SIE stack (NATS+worker+gateway) for —via-sie
- bench+model: via-sie 4-task n=300 sweep + NEXTN smaller-draft on 27B
- bench: 0.6B via-sie validated; harness + 27B config gains
- bench: 5-shot CoT for CaseHOLD (item 5 — close 27B target gap)
- bench: fix Qwen3-0.6B GPQA (parrot bug) + 27B diagnostics; final matrix
- bench: improve perf eval output handling
- docling: accept image input + run on OCR-bench quality path
- gateway+worker: chat surface accepts min_tokens + chat_template_kwargs
- gateway: strengthen generation isolation guardrails
- latency: tighten FetchExpiryController defaults to 2/15/50
- model+bench: RTX-PRO-6000 FP8 profile for Qwen3.6-27B + 6000 validation
- model: bump Qwen3-0.6B serving context 1024→4096 for prod simple-task use
- models: add Marqo/marqo-fashionSigLIP (SigLIP open_clip, fashion image-text)
- ocr: docling accepts images + quality eval prefers documents
- reconcile live worker config in sidecar
- RTX PRO 6000 FP8 profile for Qwen3.6-27B + SIE-on-6000 generative benchmark matrix
- scheduler: load-aware pipeline_depth autotune (S14 follow-up)
- scheduler: production-parity defaults + serial pipeline (carveout p99 fix)
- scheduler: restore SIE_RUST_PIPELINE_DEPTH=2 default (deep-saturation fix)
- scheduler: SIE_PULL_QUANTUM_INCLUDE_QUEUE_MS for Py-main parity
- scheduler: SIE_RUST_WAVE_CADENCE env toggle (default on)
- scheduler: step adaptive controller once per wave (Python parity)
- sidecar: add worker config and pool admission reconciliation
- sidecar: wire generation direct dispatch
- sie_server: add MinerU2.5-Pro-2604-1.2B doc OCR adapter
- sie_server: carve out QueueExecutor + IPC types for Rust worker POC
- sie_server: integrate MinerU2.5-Pro-2604-1.2B doc OCR adapter
- sie_server: UDS msgpack IPC server for Rust worker sidecar
- sie_worker_rust: close parity gaps with Python pull loop + smoke test
- sie_worker_rust: scaffold Rust worker sidecar crate (Phase 1c)
- sie_worker_rust: wire end-to-end NATS -> IPC -> publish loop (Phase 1d)
- sie-bench: synchronize loadtest measurement start
- worker/rust: IPC connection pool — lift the sidecar’s last serialization bottleneck
- worker/rust: narrate the hot path — structured INFO, slow-RPC + heartbeat-streak WARNs, full error chains
- worker: introduce InferenceBackend trait + BackendRouter
- worker: native Candle BERT backend behind
candlefeature
Bug Fixes
Section titled “Bug Fixes”- accept dense_dim in dense adapters
- adapters: replace Qwen3-VL vision Conv3d patch-embed with matmul
- adapters: route Qwen3-VL VLMs through flash attention (Vidore3 throughput)
- address pr review quality issues
- bench-27b: drop bundle from SIEServer (sie-server rejects bundle+models combo)
- bench-27b: set verbose=True on SIEServer so launch errors surface
- bench-27b: skip chat_template_kwargs on via-sie (gateway rejects unsupported field)
- bench-27b: wait for sie-server /healthz (not /health)
- bench: bump casehold/gpqa max_tokens to 2048 (CoT truncation)
- bench: let via-SIE smoke serve a profile-variant model end-to-end
- bench: resolve CPU deps for quality server
- catalog: include eval-matrix tasks so dispatch filter accepts them
- ci: address analyzer findings and stale queue test
- ci: avoid nested mise in integration fixture
- ci: keep sidecar out of warm cache
- ci: refresh gateway openapi contract
- correct e0 vm runbook paths
- deploy: add sidecar registry resources
- deploy: address server sidecar review feedback
- deploy: align server sidecar naming
- deploy: align server sidecar naming and kind preload smoke
- deploy: align tilt sidecar image naming
- deploy: document worker-sidecar metrics wiring
- deploy: keep sidecar on GHCR by default
- deploy: normalize server sidecar naming
- deploy: publish server sidecar image
- deploy: rename sidecar container to worker-sidecar
- deploy: wire SIE server sidecar for kind smoke
- deploy: wire worker sidecar image across kind and cloud
- gate sidecar nats reconnect refresh
- gateway+server: queue is the only mode — kill direct-mode cruft
- gateway: suppress H9 first-chunk-fallback on single-worker pools
- harden sidecar config recovery
- impact-map: keep profiles distinct when adapter_options differ
- keep generation machinery off default queue path
- loader: wire profile runtime.default_sampling into the adapter
- modal: report actual GPU on remote, not stale env-default
- model: bump Qwen3.6-27B default/h100 mem_fraction_static 0.85 → 0.92
- orchestrator: thread CLI -p profile through to client.extract
- preserve worker batch identity and publish image
- product: update design audit for topical docs
- quality_eval: take results-bearing JSON envelope in load_eval_json
- quality: batch3 of CodeQL findings + bench KIE bug
- quality: batch3 of CodeQL findings + bench KIE root-cause
- quality: batches 1+2 of CodeQL quality findings
- quality: close CodeQL quality-tab findings
- quality: drop redundant inline imports in donut + registry
- quality: repair adapter eval harness regressions
- remove e0 preflight httpx dependency
- require rust sidecar for queue workers
- review: 0.6B ctx test 1024->4096, loader except logs, README gaps resolved
- review: recompute 27B target delta_vs_baseline for the 2048 scores
- run directory creation
- run e0 vm scripts via uv
- scheduler: autotune signal — observed_p50/target_p50 ratio
- scope bundle config hash cache per registry
- security: bump astro to ^6.4.2 for website
- security: bump gateway deps to patched versions
- security: bump product/gtm Python lockfiles
- security: bump product/gtm/content/slides npm transitives
- security: bump root pnpm deps + add overrides for transitives
- security: bump root Python deps to patched versions
- security: bump sie_dashboard npm deps to patched versions
- security: bump sie_ts_sdk standalone pnpm transitives
- security: bump sst to ^4 to drop vulnerable aws-sdk v2
- security: cap vite at ^6 + add Node engines to website
- security: close ~190 Dependabot alerts across 9 manifests
- security: sanitize one-pager template with DOMPurify
- security: use Reflect.construct for WebSocket headers shim
- sie_bench: send SIE profile via X-SIE-MACHINE-PROFILE header
- sie_bench: use rapidfuzz for OmniDocBench edit distance
- sie_server: clear CUDA cache on uncovered VLM paths + drop private sem _value access
- sie_server: VLM cache clears on uncovered paths + drop private sem _value access
- sie-bench: budget loadtest barrier timeouts
- slow sidecar nats consumer reconcile
- smoke: launch sie_server worker with -b sglang, not -m <model>
- smoke: preload the target model in via-sie worker
- test: restore donut helper call contract
- worker-sidecar: harden queue carveout contracts
- worker/rust: one long-lived pull stream — kill 30s ack_wait stall
- worker/rust: re-copy src after cargo chef cook so real build isn’t a stub
- worker/rust: set CUDA_COMPUTE_CAP at build time (default 89, L4)
- worker/rust: stop shipping the cargo-chef stub binary as the real build
- worker: harden Candle backend + align dispatcher error contract
- worker: harden payload store + error paths; surface silent success bugs
- worker: SGLang adapter accepts min_new_tokens kwarg + 27B via-sie validated
Performance Improvements
Section titled “Performance Improvements”- adaptive: anchor min_batch_cost floor at max_batch_tokens // 4
- batching: tighten adaptive wait ceiling + revert gte-multilingual 32k
- glm_ocr: rebind vision Conv3d patch-embed to F.linear
- gte-multilingual-base: raise max_batch_tokens 16k → 32k to stop IPC-batch shred
- mineru_vl: O(L) incremental no-repeat-ngram for greedy decode
- ocr: swap pure-Python Levenshtein DP for rapidfuzz
- rope_flash: vectorize CLS/mean pooling, eliminate per-item .item() sync
- server: FP16 on GPU, coalesce sized for IPC bursts, starvation self-heal
Reverts
Section titled “Reverts”- restore adaptive batching defaults to 15/50ms
- scheduler: drop depth autotune (signal didn’t pan out in S17)
v0.4.1 (2026-05-28)
Section titled “v0.4.1 (2026-05-28)”Highlights
Section titled “Highlights”- New capabilities: add Qwen3.6-27B model + migrate to CUDA 12.9
- Reliability and operations: isolate generation direct dispatch from shared queues; resolve 18 open CodeQL alerts; use SHA256 (not SHA1) for actor_id log tag; colocate tests under infra/, update sync contract
Features
Section titled “Features”- server: add Qwen3.6-27B model + migrate to CUDA 12.9
Bug Fixes
Section titled “Bug Fixes”- isolate generation direct dispatch from shared queues
- security: resolve 18 open CodeQL alerts
- security: use SHA256 (not SHA1) for actor_id log tag
- terraform-sync: colocate tests under infra/, update sync contract
Reverts
Section titled “Reverts”- security: drop advanced CodeQL setup
v0.4.0 (2026-05-27)
Section titled “v0.4.0 (2026-05-27)”Highlights
Section titled “Highlights”- Breaking change: fail-closed authentication (default-deny)
- New capabilities: generation quality-gate scoring core (roadmap §5, trust-critical); generation-quality regression gate over the existing scorers; add cohere measurements for us-east-1; add openai measurements for us-east-1; add voyage measurements for us-east-1; regex/EBNF response_format + developer role (roadmap 1.7)
- Reliability and operations: forward provision_timeout_s in SIEImageTextWrapper.encode; raise image-task eval timeouts to fix Flickr30k nightly; exclude favicon + OG image from auth middleware; keep public surfaces vague about what’s behind auth; revert NextAuth function-form, use try/catch on Resource
- Performance: warm one Lambda, bump timeout, narrow S3 verdict fetch
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- gateway: fail-closed authentication (default-deny)
Features
Section titled “Features”- bench: generation quality-gate scoring core (roadmap §5, trust-critical)
- bench: generation-quality regression gate over the existing scorers
- benchmarks: add cohere measurements for us-east-1
- benchmarks: add openai measurements for us-east-1
- benchmarks: add voyage measurements for us-east-1
- chat: regex/EBNF response_format + developer role (roadmap 1.7)
- dashboard: add executive quality summary widget on landing
- dashboard: add hover-tooltip on ‘to verify’ explaining WARN
- dashboard: brand alignment foundation - palette, fonts, header
- dashboard: brand favicon, opengraph image, light-mode hover fix
- dashboard: brand foundation - palette, fonts, header logo
- dashboard: brand surfaces - retokenise landing page + widget
- dashboard: public preview shell on / so Slack unfurls work
- dashboard: retokenise badges - brand-elevated neutrals, DM Mono labels
- dashboard: retokenise landing surfaces onto brand palette
- dashboard: retokenise loadtest pages onto brand surfaces
- dashboard: retokenise quality pages onto brand surfaces
- dashboard: retokenise shared components and sign-in onto brand palette
- dashboard: treat WARN as passing-with-verify, shade gauge amber
- gateway,sdk,server: add native generate endpoint with improved admission control and validation
- gateway,sdk,server: add OpenAI-compatible chat completions with streaming and sampling extensions
- gateway,server: add multi-turn tool-use support with OpenAI-compatible message format
- gateway: /v1/completions (legacy OpenAI Completions, raw-prompt)
- gateway: /v1/completions streaming (text_completion SSE)
- gateway: /v1/generate accepts seed/logprobs/logit_bias/n/best_of/lora_adapter (M8)
- gateway: /v1/responses (OpenAI Responses API, MVP)
- gateway: /v1/responses structured array input (conversation history)
- gateway+worker: per-choice OpenAI streaming for n>1 (H4, H5, M4)
- gateway: accept OpenAI multimodal content-parts; reject images (no VL model)
- gateway: add routing salt + byte-preserving key mode (M11)
- gateway: advertise lora_adapters on /v1/models + pre-validate unknown names
- gateway: fail-closed authentication (default-deny)
- gateway: meaningful system_fingerprint on chat responses (roadmap 1.3/§5)
- gateway: refactor streaming and routing with improved error handling and metrics
- gateway: register /v1/moderations as explicit 501 (roadmap 1.8, phase 3)
- gateway: serve a rendered API reference at /docs (Redoc)
- gateway: unify /v1/embeddings on the OpenAI error envelope (roadmap 1.4)
- generation: best_of — over-generate + rank by logprob, return top n
- generation: complete M4 req2 generation primitive with streaming, structured outputs, and routing
- generation: multi-candidate n>1 (non-streaming) end-to-end (roadmap 1.5)
- generation: multi-LoRA serving (one base, N adapters, per-request) (roadmap 6.2)
- generation: ship generate() primitive — Qwen3.5-4B + NEXTN/MTP + xgrammar, adapter perf at parity with raw SGLang
- generation: streaming n>1 — per-candidate SSE interleave
- helm/sie-cluster: bundle cert-manager + trust-manager (opt-in) with self-signed TLS mode
- openapi: add tool_calls support to chat completion schema
- python-sdk: expose typed params for chat n/logprobs/lora_adapter/etc (M7)
- quality-eval: add heartbeat logging and improve long-running process observability
- routing: cache-aware (prefix-hash) routing (roadmap §6.3)
- sie_bench: add Cohere as a first-class eval source
- sie_bench: add Cohere multimodal embeddings
- sie_bench: add Cohere rerank backend for native MTEB rerank tasks
- sie_bench: add OpenAI Embeddings as a first-class eval source
- sie_bench: add Voyage provider source plumbing
- sie_bench: implement Voyage text embedding runner
- sie_dashboard: add /quality/compare to diff two quality runs
- terraform: add default_tags Project=sie/Cluster on all AWS providers
- terraform: add on-demand RTX 6000 baseline pool to tester-cluster
- terraform: add uniform project=sie label across all GCP clusters
- terraform: idle-stop and on-demand wake for quality-eval runner fleet
- terraform: scale quality-eval fleet to 5+5, smart-wake, 4h timeout
- terraform: wake-runners retries + watchdog queued-jobs backstop
- tester-cluster: add on-demand L4 worker pool + capacityType node pins
- ts-sdk: handle 202 provisioning in chatCompletions + expose missing fields (H1+M6)
- worker: SGLang owns grammar; worker preflight opt-in only (H8, ADR-0002)
- worker: wire mixed-pool fairness scheduler into the pull-loop (opt-in)
- worker: WorkClassScheduler core for mixed-pool fairness (roadmap §6.1)
Bug Fixes
Section titled “Bug Fixes”- bench: declare olmocr[bench] dep for OCR-bench quality eval
- bench: drop —with-deps from playwright install (no sudo on g7e)
- bench: forward provision_timeout_s in SIEImageTextWrapper.encode
- bench: install playwright chromium for OCR-bench KaTeX rendering
- bench: playwright install —with-deps for OCR-bench chromium
- bench: raise image-task eval timeouts to fix Flickr30k nightly
- bench: respect similarity() inputs in ColBERT/ColPali wrappers
- bench: run olmocr.bench.tests off orchestrator’s asyncio loop
- centralize worker_id subject normalization (M5)
- chart: pin image-prepull DaemonSets to GPU nodes
- ci: copy assets/ into the gateway Docker build (redoc bundle)
- dashboard: close remaining ‘SIE Dashboard’ leaks in <title> and og:image:alt
- dashboard: drop edge runtime on opengraph-image for OpenNext
- dashboard: exclude favicon + OG image from auth middleware
- dashboard: keep public surfaces vague about what’s behind auth
- dashboard: pick healthy daily via coverage + health gates
- dashboard: require >=50 pairs on main-run fallback in nightly picker
- dashboard: revert NextAuth function-form, use try/catch on Resource
- dashboard: short tab title for authed users, neutral for unauth
- dev: set explicit auth opt-in for local gateway launchers (post fail-closed)
- gateway,sdk,server: add wire-level validation, improve resource cleanup, and enhance observability across request lifecycle
- gateway,sdk,server: prevent metric cardinality DoS and fix non-idempotent retry logic
- gateway,sdk,server: strengthen validation and eliminate silent failures across request lifecycle
- gateway,sdk,server: validate numeric fields and improve error handling
- gateway: add NATS config trusted producers helm override
- gateway: document ModelCapabilities in OpenAPI + refresh on profile delta-update
- gateway: generation timeouts bypass legacy request-timeout ceiling (H7)
- gateway: scope LoRA adapter capabilities per profile (M10)
- gateway: strict allow-list + 400 contract on /v1/completions (H3)
- gateway: strict allow-list + 400 contract on /v1/responses (H2)
- gateway: tighten chat sampler/token-cap + tool-history validation (M1, M13)
- gateway: trust chart-rendered sie-config pod name for NATS deltas
- generation: cancel tombstone prevents first-chunk fallback double-execution (H9)
- generation: LoRA lora_path is a top-level /generate field, not a sampling param
- generation: tighten lossy tool-control flags (M14)
- grammar: resolve tokenizer adapter for Outlines processor factories and remove anchors from regex patterns
- helm/sie-cluster: guard validateTls probe against deployments with nil labels
- helm/sie-cluster: guard validateTls probe against nil deployment labels
- helm/sie-cluster: include cert-manager mode in presence-check gate
- helm/sie-cluster: label-based cert-manager detection + bidirectional runtime check
- helm/sie-cluster: make self-signed root-CA namespace configurable
- helm/sie-cluster: one-step bundled cert-manager install with self-signed TLS
- helm/sie-cluster: probe cert-manager controllers cluster-wide
- helm/sie-cluster: regenerate Chart.lock with synced digest
- helm: trim and drop empty entries in ingress.hosts
- modal: exclude cargo target/ from sandbox image mount
- pytorch_embedding: accept and forward revision kwarg
- quality_eval: tolerate stdout noise around eval JSON envelope
- quality-eval: handle paginated jobs API and align last two filters
- release-docker,warm-cache: address review findings
- server,gateway: add GPU-aware health probes to detect and recover from wedged CUDA contexts
- server,gateway: GPU-aware health probes to detect & recover from wedged CUDA contexts
- sie_bench: register OPENAI_SOURCE so —save-targets openai actually saves
- sie_dashboard: address compare-page review nits
- sie_dashboard: label compare log links with run IDs
- sie_server: base64-decode JSON image inputs
- sie_server: enforce media bytes contract at every consumer
- sie_server: install cv2 system libs for docling extract
- streaming: no-silent-drop on chunk-queue backpressure (H6)
- terraform: detect in-flight workflow runs via explicit status query
- terraform: drop redundant Project overrides in quality-eval-l4
- terraform: drop watchdog_idle_minutes default to 5, ignore PIP drift
- terraform: grant ec2:DescribeInstanceStatus to wake role
- terraform: per-runner idle-stop via GitHub Actions runners API
- terraform: require positive activity observation before idle-stop
- terraform: scope quality-eval IAM on Role tag instead of Project
- terraform: seed GPU node-group desired_size from min_size
- terraform: watchdog backstop covers queued-status runs
- terraform: watchdog grants HANG_MINUTES grace from LaunchTime
- terraform: watchdog ignores LastBusyAt older than LaunchTime
- tester-cluster: pin worker pool nodeSelectors to gpu-type as well
Performance Improvements
Section titled “Performance Improvements”- dashboard: warm one Lambda, bump timeout, narrow S3 verdict fetch
Reverts
Section titled “Reverts”- release-docker: drop matrix consolidation, keep deps push retry
v0.3.4 (2026-05-14)
Section titled “v0.3.4 (2026-05-14)”Highlights
Section titled “Highlights”- New capabilities: default payload store to model-cache bucket /payloads; typed InputTooLongError for extract 400 INPUT_TOO_LONG
- Reliability and operations: bump dev-g6-spot to g6.2xlarge; bump dev-g6-spot to g6.2xlarge so default worker pool fits; default workers to shared queue pool; pin opencv-python-headless to drop X11 runtime deps; tolerate config conflicts in bootstrap, gate on sie-config
Features
Section titled “Features”- infra: default payload store to model-cache bucket /payloads
- sdks: typed InputTooLongError for extract 400 INPUT_TOO_LONG
Bug Fixes
Section titled “Bug Fixes”- aws-example: bump dev-g6-spot to g6.2xlarge
- aws-example: bump dev-g6-spot to g6.2xlarge so default worker pool fits
- chart: default workers to shared queue pool
- cluster.py,aws.py: address review suggestions 1 & 2
- deps: pin opencv-python-headless to drop X11 runtime deps
- gateway: tolerate config conflicts in bootstrap, gate on sie-config
- gateway: tolerate config conflicts in bootstrap, gate on sie-config ready
- sdk: widen sie-sdk requires-python to >=3.12
- terraform-aws: default ECR creation off, prefix repo names with project_name
- terraform-aws: trim slashes from ecr_repository_prefix
- terraform-google-sie: wait for identity pool before binding WI
- terraform: relax required_version from ~> 1.14.3 to >= 1.14
v0.3.3 (2026-05-13)
Section titled “v0.3.3 (2026-05-13)”Highlights
Section titled “Highlights”- New capabilities: add ColQwen3 + Nemotron ColEmbed v2 visual doc retrieval; add text classification task support; add post-download load timeout with stall-based download bounds; add scope-able workflow_dispatch with model/profile/task filters; add new INPUT_TOO_LONG ErrorCode; enforce overflow_policy in gliclass adapter
- Reliability and operations: surface empty matrix and add measurement-mode for unbaselined adapters; emit task_class in quality-adapter JSON output; score detection eval predictions from result[“objects”]; annotate empty-diff path that bypasses impact_map; capture real exit code from impact_map in resolve-impact.sh
Features
Section titled “Features”- adapters: add ColQwen3 + Nemotron ColEmbed v2 visual doc retrieval
- extraction: add text classification task support
- model-loader: add post-download load timeout with stall-based download bounds
- quality-adapter: add scope-able workflow_dispatch with model/profile/task filters
- server: add new INPUT_TOO_LONG ErrorCode
- server: enforce overflow_policy in gliclass adapter
- server: route INPUT_TOO_LONG to HTTP 400 in extract API
- server: validate overflow_policy in resolve_runtime_options
Bug Fixes
Section titled “Bug Fixes”- bench: emit task_class in quality-adapter JSON output
- bench: score detection eval predictions from result[“objects”]
- ci: annotate empty-diff path that bypasses impact_map
- ci: capture real exit code from impact_map in resolve-impact.sh
- ci: collapse adapter-equivalent profiles in quality-adapter matrix
- ci: pin mise to 2026.5.5 in loadtest workflows
- ci: surface empty matrix and add measurement-mode for unbaselined adapters
- docker: install libspatialindex-c6 in worker images
- gliclass: catch IndexError empty-tensor crash as InputTooLongError
- gliclass: raise InputTooLongError from argmax-empty backstop
- probe-chart: make 3-OLD/2-NEW sample asymmetry explicit in title
- quality-adapter: namespace quality_eval tests to avoid conftest collision
- quality-adapter: split adapter_paths on commas before —changed-dirs
- quality-adapter: split Pair column so pair_key | stops breaking the table
- quality-adapter: split Pair column to stop pair_key | breaking the table
v0.3.2 (2026-05-08)
Section titled “v0.3.2 (2026-05-08)”Highlights
Section titled “Highlights”- New capabilities: default score_pairs() in BaseAdapter + baseline reranking targets; bump cold-start schema to v6 with deserialize/warmup split; per-model perf concurrency defaults for OCR adapters; adapter-triggered quality eval on persistent L4 runner; make destroy conditional on workflow_dispatch input; nightly loadtest pipeline + baseline recorder
- Reliability and operations: raise loadtest job timeout to GH Actions ceiling (360 min / 6h); clarify experimental NATS health mode; lang tag on fenced block; fail fast on invalid scenarios; surface error/no_results rows in MD; harden parse_label against unexpected filenames; adapt OCR perf Item shape to model.inputs and fail loudly on errors
Features
Section titled “Features”- adapters: default score_pairs() in BaseAdapter + baseline reranking targets
- bench,charts: bump cold-start schema to v6 with deserialize/warmup split
- bench: per-model perf concurrency defaults for OCR adapters
- ci: adapter-triggered quality eval on persistent L4 runner
- ci: make destroy conditional on workflow_dispatch input
- ci: nightly loadtest pipeline + baseline recorder
- colbert: add score_pairs support and expand model coverage
- dashboard: add status and kind filters to runs list
- dashboard: introduce run-group concept (run = 3 scenarios)
- dashboard: loadtest results dashboard (Next.js + SST + DynamoDB)
- dashboard: render every metric in the perf-lab archive
- dashboard: scaffold loadtest dashboard (Next.js + SST)
- dashboard: status and kind filters on runs list
- dashboard: track run_status; gh-API one-time backfill
- docling: add ocr profile defaulting do_ocr=true
- gateway: expose OpenAPI contract
- gateway: unify API errors and align probe contracts
- helm: expose probes value trees for worker/gateway/config
- helm: tighten startup/readiness probes for faster pod-ready
- helm: TLS termination via cert-manager + BYO matrix docs
- helm: wire probe templates to values trees
- infra: opt-in S3 cluster model cache
- ltfr: cache-vs-no-cache compare chart, with-cache run data, and 8 single-mode chart refresh
- matrix: add task_class stamping to eval measurements
- server,bench: split deserialize/warmup in cold-start instrumentation (v6)
- server: cap torch CPU threads at worker startup
- sie_server: per-stage timing markers in lifespan for engine_boot attribution
- sie_server: split adapter.warmup() out of load() with cold-start log markers
- tools: bump cold-start bench to v5 with scenario flag
- tools: LTFR per-scenario bench tooling + results (issue #652)
- tools: ltfr-bench orchestrator (issue #652)
Bug Fixes
Section titled “Bug Fixes”- bench: adapt OCR perf Item shape to model.inputs and fail loudly on errors
- bench: address CodeRabbit review on PR #779
- bench: correctly detect v6 split presence in flattened runs[]
- bench: derive emitted gpu_load_s from v6 deserialize+warmup split when available
- chart: pass —cluster-cache to sie-server and correct populate command in docs
- charts: vertical legend so ‘image pull + container init’ and ‘node prov’ aren’t clipped
- ci+terraform: three deterministic root causes for loadtest pipeline
- ci: address CodeRabbit findings on quality-adapter PR
- ci: address CodeRabbit’s second-pass review on quality-adapter
- ci: address CodeRabbit’s third-pass review on quality-adapter
- ci: auto-clear stale terraform state lock from prior runner crashes
- ci: drop double cuda12 suffix + force codebuild for missing images
- ci: forensic dump on argo failure + LB-ENI release before destroy
- ci: gate stale-lock clear behind force_unlock input + pass —aws-region to destroy
- ci: override registry/gpu-selector/tolerations + Python heredoc
- ci: parse markdown bench output → result.json synthesis
- ci: pass WORKFLOW env to run_scenarios.sh in loadtest.yml
- ci: preflight env-var check in run_scenarios + finalize scripts
- ci: provision GH PAT secret + in-cluster github-token before bootstrap
- ci: raise loadtest job timeout to GH Actions ceiling (360 min / 6h)
- ci: read bench-config from local clone, not raw.githubusercontent.com
- ci: right-size bench pod + worker pod resources for cluster shape
- cluster: move orphan-LB sweep into cmd_destroy, drop parallel script
- cluster: use project_name (not example name) for orphan-LB VPC tag
- cluster: use project_name for orphan-LB VPC tag lookup
- dashboard,ci: keep run_status consistent between S3 and DynamoDB
- dashboard,ci: wire real Prometheus matrix shape + extend headlines
- dashboard: drop time-based legacy run grouping (was unsafe)
- dashboard: GPU util shown as 0-100 (was being multiplied by 100 again)
- dashboard: include duration_seconds in run-meta.json (was DynamoDB-only)
- dashboard: normalize array-shaped searchParams before .trim()
- deps: bump plotly to >=6.1.1 for kaleido compat
- docker: stub bundles/ and models/ in deps stage
- docling: cache DocumentConverter per (device, ocr_enabled)
- docling: mark adapter unloaded in unload()
- docling: thread device through PdfPipelineOptions accelerator_options
- gateway: address latest coderabbit contract notes
- gateway: address PR review for NATS health mode
- gateway: address probe and SDK review findings
- gateway: align CreatePoolRequest OpenAPI with runtime validation
- gateway: clarify experimental NATS health mode
- gateway: close remaining review contract gaps
- gateway: preserve embeddings timing headers
- gateway: preserve scale-from-zero request path
- gateway: reject unsupported embeddings token arrays
- helm: validate ACME server and privateKeySecretRef in validateTls
- infra: grant kms:Decrypt to workers when model cache uses SSE-KMS
- infra: normalize whitespace-only model cache string inputs
- infra: treat empty model_cache_kms_key_id as unset
- infra: use flat lifecycle key for s3-bucket module v5
- loadtest-ci: force-delete orphan elbv2 LBs before terraform destroy
- loadtest-ci: force-delete orphan LBs and stop swallowing destroy failures
- loadtest-ci: poll workflow phase instead of argo submit —wait
- loadtest-ci: poll workflow phase instead of relying on argo —wait
- ltfr-bench,notes: lang tag on fenced block; fail fast on invalid scenarios; surface error/no_results rows in MD
- ltfr-bench: hoist imports to top; guard payload.results shape
- ltfr-bench: mark request_failed rows in scenario-a/b MD tables
- ltfr-bench: preserve failure context in aggregated rows; add request_failed status
- ltfr-bench: treat no_results cells as failures in exit code
- ltfr-charts: strip legend clip-path so labels render full width
- ltfr: tighten UID/timestamp guards in capture_image_pull_events
- multi_pod_cold_start: raise on ASG terminate fail; UID-filter pull events; isolate scenario-c pod
- paddleocr_vl: pass use_cache=True to generate
- paddleocr_vl: pass use_cache=True to generate to enable KV-cache
- review: tighten score_pairs options handling and query text validation
- sie-server: include model and bundle directories in wheel distribution
- terraform: detect HF-cache EBS by NVMe model + size, not by Linux name
- terraform: set resolve_conflicts_on_* = OVERWRITE on EKS addons
- tools: drop module-level docstrings (AGENTS.md rule)
- tools: guard fig_per_cell_table aggregation against empty results
- tools: guard mean() against empty engine_boot_s in aggregate()
- tools: harden parse_label against unexpected filenames
- tools: mark cold-start-bench.py executable (EXE001)
- tools: remove module docstring from cold_start_charts.py (repo rule)
v0.3.1 (2026-04-29)
Section titled “v0.3.1 (2026-04-29)”Highlights
Section titled “Highlights”- New capabilities: add OmniDocBench OCR quality loader; support /v1/score with dense/sparse/colbert/hybrid modes; add Marqo/marqo-ecommerce-embeddings-B via open_clip backend
- Reliability and operations: add terminal failed state to model registry (sie-test#85)
Features
Section titled “Features”- bench: add OmniDocBench OCR quality loader
- bge-m3: support /v1/score with dense/sparse/colbert/hybrid modes
- siglip: add Marqo/marqo-ecommerce-embeddings-B via open_clip backend
Bug Fixes
Section titled “Bug Fixes”- server: add terminal failed state to model registry (sie-test#85)
v0.3.0 (2026-04-29)
Section titled “v0.3.0 (2026-04-29)”Highlights
Section titled “Highlights”- Breaking change: openapi.json is now a committed artifact that must be regenerated and committed when API changes are made
- New capabilities: add GLM-OCR adapter; add Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B multimodal adapters; add GLiNER2 and GLiNER-bi adapters; add Qwen3-Reranker-0.6B and 4B causal LM reranker support; add SigLIP 2 base-patch16-224 vision-language encoder; add minimal cache weights snapshot command for offline deployments
- Reliability and operations: retry only transient connection errors under wait_for_capacity; surface unrouteable models loudly and helm-repo-add on pristine hosts; emit identical NATS payload to bundle and _all subjects; surface mixed-profile unrouteable models and keep snapshot consistent on writes; add retry logic for deadsnakes PPA to handle Launchpad outages
- Performance: cache JPEG-encoded corpus images across queries; lazily JPEG-encode corpus images on first use; cache SDK version parse, integer audit latency, UUIDv7; cut hot-path allocations, fuse numpy decode, tighten backpressure
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- openapi: openapi.json is now a committed artifact that must be regenerated and committed when API changes are made
Features
Section titled “Features”- adapters: add GLM-OCR adapter
- adapters: add Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B multimodal adapters
- add GLiNER2 and GLiNER-bi adapters
- add Qwen3-Reranker-0.6B and 4B causal LM reranker support
- add SigLIP 2 base-patch16-224 vision-language encoder
- admin: add minimal cache weights snapshot command for offline deployments
- bench: honor SIE_BENCH_SERVER_READY_TIMEOUT in eval orchestrator
- ci: nightly loadtest gate against dedicated EKS cluster
- ci: nightly loadtest gate, ephemeral cluster per run
- extract: add Docling adapter for PDF/DOCX/HTML extraction
- extract: add Docling adapter for PDF/DOCX/HTML parsing
- extract: plumb document items and structured
dataresults - observability: add Prometheus metrics to sie-config and expand sie-gateway coverage
- observability: Prometheus metrics for sie-config and sie-gateway
- oom: implement defensive exception fan-out and improve recovery metrics
- oom: improve error semantics and budget exhaustion detection
- openapi: add static spec export and validation
- router: import Rust gateway source tree
- server: add reactive OOM recovery and proactive idle eviction
- social: daily social content pipeline with 5-source drafts + engagement
- types: add
documentinput modality across SDKs, server, and metadata
Bug Fixes
Section titled “Bug Fixes”- adapters: add input validation guards for empty/failed visual inputs
- adapters: address review findings for Qwen3-VL adapters
- adapters: clarify video placeholder, validate token IDs, fix torch_dtype key
- add client-side hour filter to search_x_posts (was date-level only)
- address CodeRabbit review feedback
- address follow-up PR review nits
- address remaining CodeRabbit feedback (round 2)
- address review findings — negative truncation guard, score() options, constant dedup
- bench: show correct unit labels for MP/s throughput in —print-gap
- bundles: declare Qwen3-VL adapters in default bundle
- ci: use Blacksmith runner in CI
- client: retry only transient connection errors under wait_for_capacity
- cluster: address PR #701 review comments
- cluster: correct kubectl flag combo and reorder LB sweep before helm uninstall
- cluster: helm uninstall before terraform destroy to clean up AWS LB leftovers
- cluster: unblock end-to-end
mise run cluster create --build - config,cluster: surface unrouteable models loudly and helm-repo-add on pristine hosts
- config: emit identical NATS payload to bundle and _all subjects
- config: surface mixed-profile unrouteable models and keep snapshot consistent on writes
- docker: add retry logic for deadsnakes PPA to handle Launchpad outages
- docker: propagate failure when all add-apt-repository retries exhausted
- docling: per-task converter, hf_revision guard, callable typing (CodeRabbit)
- docs: Update packages/sie_server/Dockerfile.cuda11
- fail closed on missing/unparsable timestamps in lookback filter
- gateway,config,sdk: resiliency, concurrency, and cross-service hash parity
- gateway,config: address PR review — 404 for unknown models, 202 on default routing, full YAML propagation
- gateway,config: harden auth, trusted NATS producers, and recovery path; drop gateway HA default
- gateway,sdk: map upstream timeouts to 503+MODEL_LOADING for SDK retry
- gateway: add GET /v1/models/{model} detail route
- gateway: address PR #716 review feedback
- gateway: align /v1/models error and list shapes
- gateway: drop double-counted REQUEST_COUNT / REQUEST_LATENCY emit
- gateway: emit X-SIE-Error-Code header on model-loading 503
- gateway: keep record_request async to match main’s call shape
- gateway: make sie-config single source of truth for bundles with live resync
- gateway: normalize model ids in NATS work subjects + docs/tooling/ha cleanup
- gateway: pre-instantiate request/demand metric families on startup
- gateway: prioritize epoch-rewind branch; harden no-thrash test; correct arch-guide on ephemeral restart
- guard score() and score_pairs() against empty input lists
- helm: default clusterRouting to “queue” on import-sie-router-rust
- helm: enable NATS + JetStream by default to match queue clusterRouting
- helm: fail fast when gateway has no bundle source
- kind-smoke: add —no-pool-isolation for static clusters + contract-drift fixes
- kind-smoke: address bot review feedback
- kind-smoke: enable configStore and harden config/gateway tests
- kind-smoke: enable JetStream on test NATS and drop duplicate subchart
- kind-smoke: wire sie-config image and helm overrides into kind cluster fixture
- kind-smoke: wire sie-config image into kind cluster fixture
- observability: address PR review blockers on metrics PR
- sdk: cluster cache prefix probe uses list, not head
- sdk: cluster cache prefix probe uses list, not head (Refs #732, #654)
- sdk: has_children filters folder-marker objects (Refs #732, #654)
- sdk: preserve caller-supplied document format over inferred (CodeRabbit)
- sdk: retry mid-flight transport disconnects, not just timeouts
- sdk: retry on connection errors and generic 503s
- sie_config: address PR review feedback
- terraform/aws: set 100GB root volume on cpu node group to avoid DiskPressure
- terraform/gcp: undo router→gateway rename on GCP Cloud Router + NAT
- tests: include sie-config in expected missing-image list
- tests: restore docker gateway smoke test after router rename
- tmux-scripts: improve robustness of session parsing and argument handling
- types: adapt to ty 0.0.32 stricter ignore handling
- use searchTerms for X tweet-scraper actor (was searchQueries)
Performance Improvements
Section titled “Performance Improvements”- bench: cache JPEG-encoded corpus images across queries
- bench: lazily JPEG-encode corpus images on first use
- docker: add —link + move ARG BUNDLE to eliminate cross-bundle layer noise
- docker: normalize mtimes so shared venv layer is dedupable
- docker: reorder stages for maximum BuildKit cache reuse
- docker: split worker venv into shared + bundle-specific layers
- gateway: cache SDK version parse, integer audit latency, UUIDv7
- gateway: cut hot-path allocations, fuse numpy decode, tighten backpressure
- gateway: fuse msgpack_numpy decode into the response path
- gateway: move score-endpoint unwrap instead of cloning
- gateway: pass msgpack items through as rmpv::Value
- gateway: publish work items concurrently + borrow shared fields
- gateway: tighten cold-pool backpressure + cheaper QPS counter
- gateway: trim per-request work on the inference hot path
v0.2.0 (2026-04-17)
Section titled “v0.2.0 (2026-04-17)”Highlights
Section titled “Highlights”- Breaking change: Removed
--modelCLI args from worker startup; useSIE_PRELOAD_MODELSenv var or--preloadflag instead - New capabilities: add ModernBERT flash dense embedding support with fallback mechanism; add OCR quality benchmarks (olmOCR-bench); add OCR quality benchmarks with olmOCR-bench; add pages/sec throughput metric for OCR perf eval; add perf metrics to OCR eval pipeline; also report query throughput in mpix/s for image queries
- Reliability and operations: add missing NATS Helm repo to release workflow; don’t set NODE_AUTH_TOKEN for OIDC npm publishes; harden affinity spill with bounds check, clamp, and debug log; make rejected requests visible to KEDA scaling metrics; remove redundant tokenizer validation and unused template parameter
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- workers: Removed
--modelCLI args from worker startup; useSIE_PRELOAD_MODELSenv var or--preloadflag instead
Features
Section titled “Features”- adapters: add ModernBERT flash dense embedding support with fallback mechanism
- bench: add OCR quality benchmarks (olmOCR-bench)
- bench: add OCR quality benchmarks with olmOCR-bench
- bench: add pages/sec throughput metric for OCR perf eval
- bench: add perf metrics to OCR eval pipeline
- bench: also report query throughput in mpix/s for image queries
- benchmarks: add MTEB NFCorpus evaluation results for ModernBERT-based embedders
- bench: report vision corpus throughput in mpix/s
- bench: report vision corpus throughput in mpix/s instead of items/s
- deps: migrate from pynvml to nvidia-ml-py package
- haystack: add haystack_integrations namespace aliases
- haystack: add namespace-convention aliases
- observability: add anonymous usage telemetry
- sdk: add max_concurrency param to SIEAsyncClient to prevent connection pool exhaustion
- server: add lightonai/LightOnOCR-2-1B OCR adapter with next bundle
- workers: implement model preloading at startup to reduce first-request latency
Bug Fixes
Section titled “Bug Fixes”- adapters: remove redundant tokenizer validation and unused template parameter
- address PR review — panel title, namespace variable
- bench: handle unloaded images in pixel count computation
- bench: use concurrent async requests for OCR perf eval
- bench: validate image entries before computing pixel counts
- bench: validate pixel counts before using them for image corpus throughput
- build: downgrade dockerfile syntax version to 1 for broader compatibility
- ci: add missing NATS Helm repo to release workflow
- ci: don’t set NODE_AUTH_TOKEN for OIDC npm publishes
- dashboard: queue routing dashboard accuracy and usability
- docs: add update date to portfolio header
- docs: correct PR reference in reranker reclassification note
- docs: populate reranker data and simplify table header
- docs: update stale model counts after reranker reclassification
- haystack: rename namespace alias to sie
- install uv via curl instead of COPY —from ghcr.io
- preload smoke test checks model.loaded instead of nonexistent workers field
- readme: heading format
- release: add LanceDB integrations to release-please config
- router: add overflow spill to break model affinity deadlock
- router: harden affinity spill with bounds check, clamp, and debug log
- router: make rejected requests visible to KEDA scaling metrics
- sie-bench: account for in-flight drain in throughput calculation
- sie-bench: use union wall-clock for multiprocess throughput merge
- tester-cluster: patient KEDA scale-down for worker pools
v0.1.10 (2026-04-09)
Section titled “v0.1.10 (2026-04-09)”Highlights
Section titled “Highlights”- New capabilities: add async, chunking, and streaming to Weaviate document enricher; improve DLQ routing and score response handling; implement Config Management API with NATS-based distribution and review fixes; add LanceDB integration (Python + TypeScript); queue routing dashboard + NATS exporter + router image tag; queue routing dashboard, NATS prom exporter, router image tag
- Reliability and operations: correct cluster routing condition, stream max_age units, and reconnect state ordering; add recreate strategy for router deployment when nats config restore is enabled; restore Chart.yaml deps from main, keep appVersion v-prefix; queue routing dashboard PromQL for NATS wait; configurable NATS fetch budget, Helm-wired queue params
- Performance: decouple scanner and SIE batch sizes in enrich_table; stream enrich_table batch-by-batch instead of full materialization; use Lance scanner for column projection in enrich_table; bypass FastAPI for hot proxy paths via raw ASGI middleware
Features
Section titled “Features”- add async, chunking, and streaming to Weaviate document enricher
- dlq,pull-loop: improve DLQ routing and score response handling
- implement Config Management API with NATS-based distribution and review fixes
- integrations: add LanceDB integration (Python + TypeScript)
- observability: queue routing dashboard + NATS exporter + router image tag
- observability: queue routing dashboard, NATS prom exporter, router image tag
- sdk: add get_model() and configure LanceDB release workflows
- terraform: add AWS eval-eu EKS cluster with multi-GPU support
- terraform: add evaluation cluster setup for AWS with multi-GPU support and updated configurations
- terraform: add node labels, adjust pool sizes for tester cluster
Bug Fixes
Section titled “Bug Fixes”- config,queue,nats: correct cluster routing condition, stream max_age units, and reconnect state ordering
- handle BytesIO images in LlamaIndex and validate Weaviate classify config
- helm: add recreate strategy for router deployment when nats config restore is enabled
- helm: restore Chart.yaml deps from main, keep appVersion v-prefix
- helm: use generic release-please updater for appVersion
- helm: use generic updater for both Chart.yaml version fields
- helm: use l4-spot/rtx6000-spot naming convention for spot profiles
- integrations: address CodeRabbit review findings for LanceDB PR
- observability: queue routing dashboard PromQL for NATS wait
- queue-routing: configurable NATS fetch budget, Helm-wired queue params
- queue-routing: resolve bugs, add configurable NATS params, fix score wire format
- queue-routing: score response format and DLQ fallback routing key
- release: use NPM_TOKEN for initial sie-lancedb publish
- router: use “scores” key in queue-mode score responses
- terraform: add GPU subnet coverage validation
- terraform: relax AZ validation and clarify defaults
- terraform: review fixes for tester cluster infra
- terraform: switch tester-cluster to us-east-2
- terraform: Switch tester-cluster to us-east-2 and update deployment docs
- terraform: validate gpu_node_groups for duplicate and reserved names
- test: add buildx builder pause recovery and improve build error diagnostics
- update adapter tests and address code review feedback
- use OCI registry URI for helm chart in README
Performance Improvements
Section titled “Performance Improvements”- lancedb: decouple scanner and SIE batch sizes in enrich_table
- lancedb: stream enrich_table batch-by-batch instead of full materialization
- lancedb: use Lance scanner for column projection in enrich_table
- router: bypass FastAPI for hot proxy paths via raw ASGI middleware
- router: reduce thread pool pressure by inlining small deserialization
- router: remove msgpack_numpy global patch and BaseHTTPMiddleware
- router: replace stdlib json with orjson for 3-10x faster serialization
- sdk+router: lazy msgpack_numpy.patch and pure ASGI middleware
v0.1.9 (2026-04-02)
Section titled “v0.1.9 (2026-04-02)”Highlights
Section titled “Highlights”- Reliability and operations: increase docker smoke test timeouts and add retry; include $platform in worker image tag format; revert pool names to machine profile names; remove —provenance flag (requires public repo)
Bug Fixes
Section titled “Bug Fixes”- helm: include $platform in worker image tag format
- helm: revert pool names to machine profile names
- increase docker smoke test timeouts and add retry
- remove —provenance flag (requires public repo)
v0.1.8 (2026-04-01)
Section titled “v0.1.8 (2026-04-01)”Highlights
Section titled “Highlights”- Reliability and operations: add sie-qdrant and sie-weaviate to release-please config; point sync-terraform default repos to production; remove —provenance from npm publish for private repo; correct image.tag comment to reflect actual format; remove duplicate platform suffix from worker image tag
Bug Fixes
Section titled “Bug Fixes”- add sie-qdrant and sie-weaviate to release-please config
- ci: point sync-terraform default repos to production
- ci: remove —provenance from npm publish for private repo
- helm: correct image.tag comment to reflect actual format
- helm: remove duplicate platform suffix from worker image tag
- remove internal-only references from COMPATIBILITY.md
v0.1.7 (2026-04-01)
Section titled “v0.1.7 (2026-04-01)”Highlights
Section titled “Highlights”- New capabilities: add profiling script for sparse encoding hot path; add GitHub Actions workflow to sync Terraform modules to registry repos; apply QoL improvements from PR #484 review comments; switch default GPU from g5 (A10G) to g6 (L4); add rerank/score support to TEI runner; implement configurable document length limits and custom prefix token registration
- Reliability and operations: restore triggering ref for source checkout; restore quality by enabling causal attention and QK-normalization; restore dev-l4-spot zones to us-central1 for GPU availability; check /metrics endpoint in test_prometheus_metrics_exist; add per-attempt timeout to lease renewal fetch
- Performance: optimize MoE expert dispatch with sorted-expert routing; batch MaxSim scoring across documents on GPU; batch sparse aggregation with segment_reduce and fuse relu; batch split_embeddings + validate ColBERT performance
Features
Section titled “Features”- adapters: add profiling script for sparse encoding hot path
- add GitHub Actions workflow to sync Terraform modules to registry repos
- apply QoL improvements from PR #484 review comments
- aws: switch default GPU from g5 (A10G) to g6 (L4)
- bench: add rerank/score support to TEI runner
- colbert: implement configurable document length limits and custom prefix token registration
- deploy: move namespace, SA, and HF token secret management to Helm chart
- deploy: prepare Terraform modules for public registry publishing
- deploy: rewrite example module sources to registry references
- deploy: rewrite Helm and internal references for public release
- deploy: two-artifact model — GCP Terraform infra-only, batteries-included Helm chart
- docker: add —docker-platform flag to docker build task
- extend create_pool API/SDK with minimum_worker_count and bundle
- helm: add batteries-included sub-chart dependencies to sie-cluster
- helm: add image pre-pull DaemonSet for GPU worker pools
- helm: add step to build Helm chart dependencies in Kind smoke tests
- helm: default router to image-embedded model configs
- helm: enable image pre-pull DaemonSet by default
- helm: port health gates from Terraform to Helm post-install hooks
- helm: remove prometheus alias, bump to v0.2.0, standardize chart
- infra: add Modal GPU sandbox for remote benchmark execution
- infra: add rollout warning and explicit image_type for GCFS
- infra: enable GCFS image streaming on GPU node pools
- infra: set min_node_count=1 on L4 spot GPU node pools
- integrations: add Qdrant integration
- integrations: add Qdrant integration with native sparse vector support
- integrations: add Weaviate v4 integration with Go module spec
- multiprocess loadtest + SDK aiohttp migration
- sdk: add version negotiation headers between SDK and server
- sdk: default wait_for_capacity=True and timeout=900s
- sdk: version negotiation header (SDK ↔ server)
- sie-bench: add dataset/input_type fields for mTEB corpus inputs
- sie-bench: built-in multiprocess loadtest mode
- skills: add eval-model skill for HF model assessment
- skills: add eval-model skill for HF model integration assessment
- sync Terraform modules to registry repos
- tei-runner: add /embed_sparse support for sparse models
- tei: add /embed_sparse support and auto-detect pooling mode
- terraform/aws: restore cluster autoscaler helm release to infra module
- terraform/aws: strip k8s resources, restructure as infra module with examples
- terraform: add cluster name and artifact registry variables; update node pool configuration
- terraform: add EBS CSI driver, NVIDIA device plugin, default StorageClass
- terraform: strip gcp k8s/ layer; examples use infra-only module
- tools: add ColBERT query vs document profiling script
- tools: add dense P50 latency profiling script
Bug Fixes
Section titled “Bug Fixes”- adapters: sort IDF unique_ids to satisfy SparseVector contract
- add missing production example to tf validate; fix tempfile leak; remove module docstring
- address PR #478 review feedback
- address PR review — GPU alert formula, kubectl parsing, CI path filter
- address review feedback for npm publish
- address review findings - race prevention, cleanup, lighter checkout
- alloy: add stage.cri{} before stage.json to unwrap CRI log envelopes
- alloy: explicitly set configMap name and key for sub-chart wiring
- alloy: scope pod discovery to current node via field selector
- bench: complete g5 to g6 migration in AWS eval configs and GPU mapping
- benchmarks: use TEI /embed_all for ColBERT multi-vector models
- bench: skip loading candidates_model for single-model servers
- chart: update home URL and Helm install command in README
- CI compatibility and consistent env var usage
- CI compatibility for sync-terraform workflow
- ci: add contents: read permission to publish-pypi-oidc job
- ci: add helm repo add + dep build to kind-smoke workflow
- ci: g5 refactored to g6 already
- cluster: build concrete helm command in status from infra_outputs
- cluster: guard helm/kubectl post-create log when outputs are empty
- deploy: clean terraform init artifacts before push
- deploy: correct smoke test TypedDict access and helm dry-run args
- deploy: correct StatefulSet rollout semantics, PDB scope, and KEDA pause
- deploy: remove dangling kubernetes_namespace_v1.sie references from health_gates.tf
- deploy: restore triggering ref for source checkout
- deploy: update default destination repos for GCP and AWS modules
- deploy: use triggering ref for source checkout in sync-terraform
- disable LoRA adapter layers after loading to prevent quality corruption
- docs: clarify optional image push in AWS and GCP README files
- docs: update Helm chart path in AWS and GCP README files
- fix integration test
- helm,hook: deploy/helm/sie-cluster/templates/hooks/prometheus-ready-test.yaml
- helm: add before-hook-creation to Job delete policies; document count==0 expectation
- helm: address coderabbit findings on health gate hooks
- helm: address non-blocking review findings from PR #336
- helm: address review findings in batteries-included sub-chart PR
- helm: address reviewer suggestions for health gate hooks
- helm: address second-pass review findings
- helm: aggregate buckets by le in p95 latency alert
- helm: bump chart version to 0.1.1 (patch, not minor)
- helm: clarify prometheusAddress comment — ignored when sub-chart is installed
- helm: correct kube-prometheus-stack semver constraint and remove hardcoded grafana password
- helm: correct misleading validation comment in router-deployment.yaml
- helm: don’t emit ScaledObject CRDs unless KEDA is confirmed present
- helm: downscope KEDA RBAC to Role/RoleBinding; remove runtime apk installs
- helm: fix loki service URL and extract alloy config to file
- helm: fix three blocking review issues in sub-chart dependencies
- helm: improve temporary values file handling in helm_template function
- helm: improve, simplify, and modularize sie-cluster chart
- helm: move ‘app.kubernetes.io/part-of’ label to selector labels for consistency
- helm: remove autoscaling.enabled from values-aws.yaml
- helm: render KEDA ScaledObjects via post-install hook to avoid CRD chicken-and-egg
- helm: replace hardcoded namespace in provisioning alert rules
- helm: require non-empty hfToken.value when hfToken.create is true
- helm: sub-chart naming, Loki compactor, event exporter ECR, Grafana folders
- helm: use autoscaling.prometheusAddress in prometheus hook; remove stub health_gates.tf
- helm: use full FQDN for Prometheus service in KEDA and health gates
- helm: use router.service.port in NOTES.txt instead of hardcoded 8080
- infra: update min_node_count default in top-level GCP module
- normalize SDK version warned-set key to major.minor
- pool error types, add pool/progress test coverage
- profiling: add flash variant registry, device validation, top-level import
- profiling: sync GPU before tensor timing, move script to tools/
- profiling: use in-place relu_ to match production code path
- qwen3: restore quality by enabling causal attention and QK-normalization
- readme: correct helm chart path
- readme: correct helm install command
- release: track all package versions via release-please extra-files
- release: track TS SDK version.ts via release-please
- replace corrupted bge-m3 NanoFiQA2018 target + set bfloat16 precision
- review items
- router: increase pool lease TTL to survive rolling upgrades
- router: resolve default pool GPU for scale-up when gpu/pool omitted
- router: use effective_pool instead of pool_name for default pool GPU extraction
- sdk: defer aiohttp session creation to fix “no running event loop” in SIEAsyncClient
- sie_bench: improve —print-gap report accuracy and readability
- tei-runner: validate /embed_all returns per-token embeddings
- tei-runner: validate output_type in TEIRunner init
- terraform/aws: add full -backend-config flags to production init command
- terraform/aws: add precondition asserting >=2 GPU-capable AZs exist
- terraform/aws: address review findings post-restructure
- terraform/aws: correct helm chart path in dev-g5-spot example comment
- terraform/aws: filter VPC AZs to only zones offering the GPU instance type
- terraform/aws: fix invalid splat on instance type offerings locations
- terraform/aws: remove provider aws block from child module
- terraform/aws: use var.project_name in VPC subnet cluster tags
- terraform: add validation for GPU node pool zones to ensure they match the configured region
- terraform: restore dev-l4-spot zones to us-central1 for GPU availability
- terraform: update GPU instance type description for clarity and add dev-g6-spot example
- terraform: update stale k8s module references in comments
- terraform: upgrade AWS modules and fix deprecations
- test: check /metrics endpoint in test_prometheus_metrics_exist
- test: update EKS tests from g5 to g6 after GPU instance type change
- ts-sdk: add per-attempt timeout to lease renewal fetch
- use prepack instead of prepublishOnly
- validate minimum_worker_count input and soften docstrings
Performance Improvements
Section titled “Performance Improvements”- adapter: optimize MoE expert dispatch with sorted-expert routing
- adapters: batch MaxSim scoring across documents on GPU
- adapters: batch sparse aggregation with segment_reduce and fuse relu
- adapters: batch split_embeddings + validate ColBERT performance
- adapters: batch split_embeddings in ColBERT adapters
- adapters: eliminate GPU overhead from IDF query encode path
- florence2: greedy decoding for OCR (-23% P50)
- florence2: switch OCR configs from beam search to greedy decoding
- server,bench: add batch coalescing, query warmup, and benchmark stability improvements
- server: dispatch immediately when worker is idle
- server: optimize BertFlashAdapter inference path (+35% corpus throughput)
- server: reduce batch wait timeout 10ms \u2192 2ms for lower Doc P50
v0.1.6 (2026-03-12)
Section titled “v0.1.6 (2026-03-12)”Highlights
Section titled “Highlights”- Breaking change: remove
florence2andglinerstandalone bundles — extraction adapters (gliner, glirel, gliclass) are now included in thedefaultbundle - New capabilities: add native MTEB reranking task support with MRR metric; encode-dense matrix eval — 3 models × 8 tasks; add date-prefixed versioning for chronological filename ordering; reorganize and expand model size lookup table with alphabetical ordering; add detailed perf metrics, metric filter, and threshold selector; add marimo benchmark dashboard notebook
- Reliability and operations: restore /var/cache/apt mounts, keep /var/lib/apt removed; add HF_TOKEN auth and config kwargs and fix stella models; add dense projection support to Qwen2FlashAdapter; apply query_template from runtime options in SentenceTransformerDenseAdapter; replace pip install with uv add in docker error messages
- Performance: vectorize GTE sparse encode path; vectorize tokenization and packing for gte-multilingual-base; vectorize tokenization and packing to reduce throughput gap; switch Qwen/GTE models to flash attention adapter
Breaking Changes
Section titled “Breaking Changes”- bundles: remove
florence2andglinerstandalone bundles — extraction adapters (gliner, glirel, gliclass) are now included in thedefaultbundle
Features
Section titled “Features”- bench: add native MTEB reranking task support with MRR metric
- bench: encode-dense matrix eval — 3 models × 8 tasks
- benchmarks: add date-prefixed versioning for chronological filename ordering
- benchmarks: reorganize and expand model size lookup table with alphabetical ordering
- benchview: add detailed perf metrics, metric filter, and threshold selector
- benchview: add marimo benchmark dashboard notebook
- benchview: add perf metric selector to Model Size tab
- benchview: detailed perf metrics, metric filter, threshold selector
- router,bench,sdk: improve throughput with inflight tracking, batching, and connection pooling
- server: typed request parsing with msgspec
- sie_server: add gliner, glirel, and gliclass extraction dependencies to the default bundle
Bug Fixes
Section titled “Bug Fixes”- adapter: add dense projection support to Qwen2FlashAdapter
- adapter: apply query_template from runtime options in SentenceTransformerDenseAdapter
- apply CodeRabbit auto-fixes
- bench: replace pip install with uv add in docker error messages
- benchview: add missing statistics import and use _median helper
- bundles: include sglang bundle in default cluster and eval-matrix configs
- client: update websocket header parameter name from extra_headers to additional_headers
- colbert: enable native mode fallback for non-CUDA devices and add Matryoshka truncation
- deps: cap timm upper bound and fix lazy handler init
- docker: clear stale apt lists before update to prevent 404s
- docker: remove all apt cache mounts from Dockerfiles
- docker: remove no-op /var/lib/apt cache mount from apt RUN blocks
- docker: restore /var/cache/apt mounts, keep /var/lib/apt removed
- helm: increase CPU worker pool memory limits for expanded default bundle
- model: add missing query_template to stella_en_400M_v5
- models: switch all-MiniLM-L6-v2 to SentenceTransformerDenseAdapter
- multilingual-e5-large-instruct: use instruct query template, NFCorpus 0.3521 → 0.3567
- replace invalid HTML entities in SVG with XML numeric entities
- rope_flash: clear cached _rope_dummy on unload and use torch.cat for packing
- router: also resolve pool-derived GPU names to spot variants
- router: resolve bare GPU types to spot variants for KEDA scaling
- sdk: resolve sync/async client inconsistencies in score() and encode()
- server: centralize request validation to prevent 500s from malformed items
- server: use BertFlashAdapter for e5-small-v2, resolve e5 perf anomalies
- server: use BertFlashAdapter for intfloat/e5-small-v2 and remove stale benchmarks
- set compute_precision to bfloat16 for stella_en_1.5B_v5
- sie_server: add HF_TOKEN auth and config kwargs and fix stella models
- sie_server: resolve BGE-M3 linear weights loading for HF model IDs and fix test fixtures
- sie_server: support NV-Embed-v2 with PyTorch embedding adapter
- splade: align special token filtering and guard empty batches
- test: always rebuild Docker images to pick up code changes
- typecheck: move ty type checker from mise tool to uv dependency
Performance Improvements
Section titled “Performance Improvements”- adapters: vectorize GTE sparse encode path
- rope_flash: vectorize tokenization and packing for gte-multilingual-base
- rope_flash: vectorize tokenization and packing to reduce throughput gap
- server: switch Qwen/GTE models to flash attention adapter
- splade: vectorize tokenization and sparse aggregation (1.5x throughput)
- splade: vectorize tokenization and sparse aggregation in SPLADEFlashAdapter
v0.1.5 (2026-02-27)
Section titled “v0.1.5 (2026-02-27)”Highlights
Section titled “Highlights”- New capabilities: add GLiNER v2.5 model configs; stream request bodies through proxy instead of buffering; add classification model configs for GLiClass-large and cross-encoder NLI
- Reliability and operations: release pipeline cache collision and smoke test timeout; strip content-length header from streamed proxy responses
- Performance: stream response body to eliminate bytes.join bottleneck
Features
Section titled “Features”- models: add GLiNER v2.5 model configs
- router: stream request bodies through proxy instead of buffering
- sie_server: add classification model configs for GLiClass-large and cross-encoder NLI
Bug Fixes
Section titled “Bug Fixes”- release pipeline cache collision and smoke test timeout
- router: strip content-length header from streamed proxy responses
Performance Improvements
Section titled “Performance Improvements”- router: stream response body to eliminate bytes.join bottleneck
v0.1.4 (2026-02-27)
Section titled “v0.1.4 (2026-02-27)”Highlights
Section titled “Highlights”- Reliability and operations: revert sharing=locked, add cache-read-only for build step
Bug Fixes
Section titled “Bug Fixes”- revert sharing=locked, add cache-read-only for build step
v0.1.3 (2026-02-26)
Section titled “v0.1.3 (2026-02-26)”Highlights
Section titled “Highlights”- Reliability and operations: revert token lifetime extension, re-auth before push instead; revert token lifetime, re-auth before push
Bug Fixes
Section titled “Bug Fixes”- revert token lifetime extension, re-auth before push instead
- revert token lifetime, re-auth before push
v0.1.2 (2026-02-26)
Section titled “v0.1.2 (2026-02-26)”Highlights
Section titled “Highlights”- Reliability and operations: release image builds failing from GCP token expiry
Bug Fixes
Section titled “Bug Fixes”- release image builds failing from GCP token expiry
v0.1.1 (2026-02-26)
Section titled “v0.1.1 (2026-02-26)”Highlights
Section titled “Highlights”- Reliability and operations: update bundle definitions to replace legacy and gte-qwen2 with gliner
Bug Fixes
Section titled “Bug Fixes”- bundles: update bundle definitions to replace legacy and gte-qwen2 with gliner
v0.1.0 (2026-02-26)
Section titled “v0.1.0 (2026-02-26)”Highlights
Section titled “Highlights”- Breaking change: HTTP 409 dependency conflict responses are removed from all API endpoints; the DEPENDENCY_CONFLICT error code no longer exists; .beads/ issue tracking data removed from repository
- New capabilities: add X-SIE-Worker response header for per-worker metrics tracking; add encode-image-text measurements to benchmarks dir; add encode-multivector perf measurements; add encode-multivector performance measurements; add encode-visual-document perf measurements; add encode-visual-document performance measurements
- Reliability and operations: increase helm install timeout from 10m to 15m; add trailing empty line to gitignore; align release-images workflow with docker task flags; register GLiClass and DeBERTa models in bundles; build and deploy gliner bundle in Kind smoke tests
- Performance: add connection pooling load test results (Feb 24); pool httpx client and add X-SIE-Worker header in router proxy; pool httpx client in router proxy to eliminate per-request TCP overhead; move transformers imports to module level
⚠ BREAKING CHANGES
Section titled “⚠ BREAKING CHANGES”- deps: HTTP 409 dependency conflict responses are removed from all API endpoints; the DEPENDENCY_CONFLICT error code no longer exists
- .beads/ issue tracking data removed from repository
- deps: model config files no longer support the
dependenciesfield
Features
Section titled “Features”- add X-SIE-Worker response header for per-worker metrics tracking
- benchmarks: add encode-image-text measurements to benchmarks dir
- benchmarks: add encode-multivector perf measurements
- benchmarks: add encode-multivector performance measurements
- benchmarks: add encode-visual-document perf measurements
- benchmarks: add encode-visual-document performance measurements
- benchmarks: add extract-detection L4-SPOT performance measurements
- benchmarks: add extract-kie-docvqa measurements to benchmarks dir
- benchmarks: add extract-relation L4-SPOT performance measurement
- benchmarks: add score-colbert perf measurements
- benchmarks: add score-colbert performance measurements
- models: add encode-image-text measurements
- models: add extract-detection measurements
- models: add extract-kie-docvqa measurements
- models: add extract-relation measurements
- router: add structured audit logging for API requests
Bug Fixes
Section titled “Bug Fixes”- .claude: add trailing empty line to gitignore
- align release-images workflow with docker task flags
- bundles: register GLiClass and DeBERTa models in bundles
- ci: build and deploy gliner bundle in Kind smoke tests
- colbert: remove CUDA requirement and improve device compatibility
- eval: read ‘sie_id’ instead of ‘name’ from model configs in runner
- extract: use dict access for Entity TypedDict in sort
- gliner: relax stale transformers<4.52 pin
- increase helm install timeout from 10m to 15m
- reduce cpu-gliner resource requests for Kind CI
- router: read ‘sie_id’ instead of ‘name’ from model configs
- server: migrate NLI adapter to classifications and improve API consistency
- server: migrate nli_classification adapter and improve type annotations
- server: populate classifications instead of entities in GLiClass adapter
- use manifest mode for release-please and reset to v0.0.0
- use nested .gitignore for .claude/ directory
Performance Improvements
Section titled “Performance Improvements”- add connection pooling load test results (Feb 24)
- pool httpx client and add X-SIE-Worker header in router proxy
- pool httpx client in router proxy to eliminate per-request TCP overhead
- pytorch-embedding: move transformers imports to module level
- server: use uvloop as default event loop for uvicorn
Reverts
Section titled “Reverts”- keep CONTRIBUTING.md clone URLs pointing to sie.git
Miscellaneous Chores
Section titled “Miscellaneous Chores”- remove beads, agent prompts, mypy refs; consolidate ty config
Code Refactoring
Section titled “Code Refactoring”- deps: move adapter dependencies from per-adapter pyproject.toml to bundle YAML
- deps: remove model-level dependencies feature