Skip to content
Why did we open-source our inference engine? Read the post

Release Notes

Latest version: v0.7.1 (2026-08-09).

  • New capabilities: serve docling OCR with the English recogniser
  • Reliability and operations: meter executed GLiNER token window
  • server: serve docling OCR with the English recogniser
  • server: meter executed GLiNER token window
  • Breaking change: aws-control-plane no longer accepts component_contract_sha256 or injects its environment marker.
  • New capabilities: add fail-closed estate-bootstrap saved-plan inspection; add frozen assembly apply seam; add minimal promotion eligibility guard; admit estates by declared capability and add the sandbox desired state; allow benchmark campaign self-review; assert apply-free and same-scope second convergence runs
  • Reliability and operations: align catalog smoke timeout budget; apply the checkpoint whitelist to record, harden the git guard; bind the published price table to the active billing authority; harden document field fixture validation; harden generation response handling
  • terraform: aws-control-plane no longer accepts component_contract_sha256 or injects its environment marker.
  • cloud: add fail-closed estate-bootstrap saved-plan inspection
  • cloud: add frozen assembly apply seam
  • cloud: add minimal promotion eligibility guard
  • cloud: admit estates by declared capability and add the sandbox desired state
  • cloud: allow benchmark campaign self-review
  • cloud: assert apply-free and same-scope second convergence runs
  • cloud: batch the priced rate-grade cells behind one approval
  • cloud: bill audio_ms at a 1,150 ms minimum duration
  • cloud: bind the inspection verdict to the exact plan digest
  • cloud: declare the terraform backend kind per estate cell
  • cloud: deepen the read-only estate bootstrap preflight
  • cloud: default tier points to min(floor x 2.0, ceiling x 0.85)
  • cloud: freeze infrastructure values in assemblies
  • cloud: hydrate terraform outputs for remote-backend cells
  • cloud: implement the #2441 pricing-policy decisions (PR #3024)
  • cloud: local-only session auth for the L3 verification stack
  • cloud: make the sandbox cell serving
  • cloud: permanently allow benchmark campaign self-review
  • cloud: persist deliverable metadata in s3
  • cloud: publish the live rate book to a public price artifact
  • cloud: record and verify estate bootstrap approvals
  • cloud: record convergence evidence and compare two runs
  • cloud: warn loudly when a —live preflight run proves nothing
  • control-plane: serve-image capability manifest and operator-opened tasks
  • model-ops: add the docling English OCR artifact publish tool
  • ci: refuse an unusable assets_path before the gate
  • ci: replace the batch ledger atomically and unwrap the summary prose
  • clients: address terminal response review
  • cloud: accept padded whisper transcripts
  • cloud: accept real names, escape paths, close the commondir fail-open
  • cloud: address deployment review findings
  • cloud: admit provisional task measurement pairs
  • cloud: admit same-apply IAM unknowns by provenance, not by fiat
  • cloud: align catalog canary model contracts
  • cloud: align catalog smoke timeout budget
  • cloud: align document field canary semantics
  • cloud: align phase 3 exit criteria and make skip tests prove their cause
  • cloud: align PII acceptance taxonomy
  • cloud: allow task candidates across routes
  • cloud: allowlist read-only init by its full argument set
  • cloud: apply the checkpoint whitelist to record, harden the git guard
  • cloud: assert the full run built something and order the pair
  • cloud: bind every resource to the verified provider, not only data sources
  • cloud: bind read-only terraform access to the resolved cell’s root
  • cloud: bind the published price table to the active billing authority
  • cloud: bound Modal image build cache
  • cloud: bound provenance to policy CONTENT, not just mentioned resources
  • cloud: close checkpoint leak, env-independent repo guard, strict schema
  • cloud: close fail-open branch chains in the estate preflight
  • cloud: close fail-open paths in estate plan inspection
  • cloud: close preflight fail-opens found by the mandated self-review
  • cloud: close the fail-open paths found in review threads
  • cloud: close the fail-opens the third self-review found
  • cloud: close the holes the fourth self-review found in this round’s own permissive paths
  • cloud: close the plan-inspection bypasses from the adversarial review
  • cloud: close the two backend-classifier evasions from verification
  • cloud: commit job results before settlement
  • cloud: confine catalog subsets to sandbox
  • cloud: constrain checkpoint field values, not just field names
  • cloud: constrain screenshot bbox axes
  • cloud: correct the apply-check rationale and scope-check claim
  • cloud: correct the offload-smoke minted tuple annotation and notice an ignored —account
  • cloud: correct the provider-key rule against ground truth, never echo after values
  • cloud: declare the revoke 404 contract and pin its no-oracle body
  • cloud: decouple availability from measurements
  • cloud: defer gateway replica probe until serve
  • cloud: derive the declared commercial margin from the price multiplier
  • cloud: enforce the D1 boundary on the terraform primitive
  • cloud: escape the last raw path, refuse an unreadable HEAD
  • cloud: fail closed on an unreadable desired-state spec
  • cloud: freeze deployment check results
  • cloud: fullmatch the rate-book version before it becomes a filename
  • cloud: give every checkpoint field a format, not a length cap
  • cloud: guard the path parameters, pin refusals, drop the bare except
  • cloud: harden document field fixture validation
  • cloud: harden generation response handling
  • cloud: harden the estate preflight probes per security review
  • cloud: harden the point default and the artifact publish tool
  • cloud: harden the preflight per the CodeRabbit review threads
  • cloud: hold the D1 remote-backend boundary in live preflight too
  • cloud: hold the priced generation window’s prefill mix to a review
  • cloud: isolate deployment plan inputs
  • cloud: keep production catalog topology strict
  • cloud: keep the pricing-authority resolver import-safe off a checkout
  • cloud: make shape handling total and stop rejecting the honest plan
  • cloud: make the unpriced-storage zero loud
  • cloud: parse s3 backend blocks brace-aware so nested blocks are read
  • cloud: pin the AWS-edge coupling and refuse malformed input everywhere
  • cloud: read not-owned terraform roots through a narrow allowlist
  • cloud: read trust_remote_code from both adapter scopes everywhere
  • cloud: recover externally cancelled live gateway candidate
  • cloud: refuse impossible instants and stop the scope string overclaiming
  • cloud: register migration 0079 and name the book on a merged sweep
  • cloud: reject a run paired with itself and fix the stale exit bullet
  • cloud: reject lineage cycles and parented full runs
  • cloud: reject unusable SSM pagination tokens; pin cross-estate isolation
  • cloud: repair chat catalog canary
  • cloud: repair document field canary semantics
  • cloud: require exact screenshot labels
  • cloud: require statement constancy per security-relevant field
  • cloud: restore cross-estate rejection on excluded entries, refuse composition
  • cloud: restrict the instant pattern to ASCII digits
  • cloud: retry transient job reads
  • cloud: scan every leaf for foreign-estate references, not just name keys
  • cloud: scope key mutations to an account the caller names
  • cloud: snapshot untracked development roots
  • cloud: stabilize catalog contract canary
  • cloud: stabilize screenshot acceptance fixture
  • cloud: stabilize screenshot contract canary
  • cloud: stabilize whisper catalog canary
  • cloud: stop arming gateway proxy auth without a public edge
  • cloud: stop publishing SIE’s markup to customers
  • cloud: tighten promotion evidence validation
  • cloud: tolerate EC2 worker visibility lag
  • cloud: track %{} template-directive spans in the backend classifier
  • cloud: track template interpolations in the backend classifier
  • cloud: unpin gateway path dependency
  • cloud: validate mode, harden provider-key parsing, make two tests honest
  • cloud: validate read entries and close self-review fail-opens
  • cloud: validate the estate id where it enters, tighten approver
  • cloud: verify materialized candidate bytes
  • deps: provide the dependencies the bundle and the eval already declare
  • model-ops: pin trusted digests and refuse unrecorded provenance
  • release: unpin sidecar path dependency
  • sdk: refresh stale job result refs
  • server: echo queued extract item ids
  • server: normalize Grounding DINO instructions
  • server: preserve pinned models in serve config
  • server: validate pinned model selection
  • tools: allow SDK 0.7 releases
  • terraform: drop control-plane contract marker
  • New capabilities: compile the deployment matrix from the estate registry
  • Reliability and operations: drop the invalid queue key from the staging concurrency blocks
  • Performance: accelerate Qwen thinking inference
  • cloud: compile the deployment matrix from the estate registry
  • ci: drop the invalid queue key from the staging concurrency blocks
  • models: accelerate Qwen thinking inference
  • New capabilities: bind evidence and acceptance to estate identity; optimize Gemma thinking profiles
  • Reliability and operations: preserve release tag during public sync verification; verify the acceptance OIDC identity token
  • cloud: bind evidence and acceptance to estate identity
  • server: optimize Gemma thinking profiles
  • ci: preserve release tag during public sync verification
  • cloud: verify the acceptance OIDC identity token
  • New capabilities: derive resource names and paths from estate coordinates; gate public edge and vendor identity on estate capability; split managed gates into tier and estate capability; optimize non-thinking generation profiles
  • Reliability and operations: classify entity canary failures; gate Stripe preflight before any provider request; keep vendor provider access tier-gated; refresh siglip diagnostic projection; remove stale console gateway URL and reject tampered cells
  • cloud: derive resource names and paths from estate coordinates
  • cloud: gate public edge and vendor identity on estate capability
  • cloud: split managed gates into tier and estate capability
  • server: optimize non-thinking generation profiles
  • cloud: classify entity canary failures
  • cloud: gate Stripe preflight before any provider request
  • cloud: keep vendor provider access tier-gated
  • cloud: refresh siglip diagnostic projection
  • cloud: remove stale console gateway URL and reject tampered cells
  • cloud: scope Modal denial floor to serving estates
  • cloud: update gateway profile fixtures
  • quality: extend PubTabNet task budget
  • release: extract refresh artifact by id
  • release: keep diagnostic suites out of git
  • release: preserve refresh artifact across retries
  • server: pin speculative draft launches
  • server: stage drafts for local models
  • New capabilities: add staging comparison diagnostics; pin benchmark campaigns to a verified main revision; add connector plan controls; expose recovery-bound connector repair; activate atomic connector run; activate explicit connector execute
  • Reliability and operations: authenticate quality gate attempt markers; authenticate req14 quality gate markers; pin void authorization Python; bind retained recovery to candidate tasks; bind vendor mutations to estate authority
  • bench: add staging comparison diagnostics
  • ci: pin benchmark campaigns to a verified main revision
  • clients: add connector plan controls
  • clients: expose recovery-bound connector repair
  • cloud: activate atomic connector run
  • cloud: activate explicit connector execute
  • cloud: activate governed postgres connectors
  • cloud: activate recovery-bound connector repair
  • cloud: add atomic connector run admission
  • cloud: add bounded connector recovery probes
  • cloud: add bounded postgres source planning
  • cloud: add component convergence contracts
  • cloud: add connector execution journal client
  • cloud: add connector plan guardrails
  • cloud: add connector plan schema policy
  • cloud: add crash-safe connector outer runner
  • cloud: add direct connector inference
  • cloud: admit executable connector plans
  • cloud: attest connector workers before payload
  • cloud: band the size-class tier book on (operation, parameters)
  • cloud: band the tier book on (operation, parameters)
  • cloud: bind connector executor dispatch
  • cloud: bind connector plans to worker identity
  • cloud: bind direct connector executions
  • cloud: bound postgres connector execution context
  • cloud: bridge durable connector execution dispatch
  • cloud: commit the control-plane OpenAPI contract with a diff gate
  • cloud: compile estate vendor coordinates
  • cloud: compile optional estate vendor authority
  • cloud: compose connector recovery journal
  • cloud: compose dormant connector execution
  • cloud: enable native console previews
  • cloud: enable selective component deployment
  • cloud: expose connector execution transitions
  • cloud: expose durable connector plans
  • cloud: fail closed on unconfigured vendor authority
  • cloud: fence postgres connector publication
  • cloud: govern connector plan manifests
  • cloud: isolate SigLIP rate-grade diagnostics
  • cloud: journal connector recovery phases
  • cloud: land governed Postgres connector runtime
  • cloud: persist connector batch lane authority
  • cloud: persist connector execution authority
  • cloud: persist durable connector plans
  • cloud: preflight connector worker discovery
  • cloud: prepare and fence connector execution
  • cloud: prove executable postgres preflight
  • cloud: provision connector runtime authorities
  • cloud: provision isolated console previews
  • cloud: reconcile connector publication recovery
  • cloud: record the owner’s acceptance of the declared exposure
  • cloud: replace catalog DSL with task implementations
  • cloud: replay governed postgres manifests
  • cloud: resolve estate coordinates at the provisioner boundary
  • cloud: resolve terraform coordinates from the estate registry
  • cloud: route generation through local ingest
  • cloud: route managed generation through local ingest
  • cloud: scaffold additional staging estate
  • cloud: scaffold sandbox foundation
  • cloud: settle connector staging before publication
  • cloud: unify Modal environment authority in the estate registry
  • gateway: expose immutable execution evidence
  • generation: add 256k profile candidates
  • generation: add hardware-specific context profiles
  • generation: add strict Responses API parity
  • model-ops: allowlist lightonai and OpenSearch-AI as trusted publishers
  • model-ops: capture nested multimodal config scalars in onboarding evidence
  • model-ops: enable multimodal image-text and document-retrieval discovery families
  • model-ops: enable sparse-retrieval discovery family
  • model-ops: re-task late-interaction discovery onto SCIDOCS retrieval evidence
  • model-ops: trust JinaAI discovery publisher
  • models: onboard lightonai/mLateOn
  • models: onboard lightonai/mLateOn (scaffolded)
  • models: seed mLateOn quality floors from governed measurement
  • server: add bounded local-ingest generation client
  • server: complete generative model serving
  • sidecar: stream generation over local ingest
  • terraform: register shared Scalr modules
  • bench: validate release evidence ingress early
  • build: relock every standalone project that depends on sie-server
  • ci: align standalone validation artifacts
  • ci: authenticate quality gate attempt markers
  • ci: authenticate req14 quality gate markers
  • ci: clarify invalid model gate error
  • ci: exempt registry library modules from the estate root list
  • ci: fail closed for unlisted registry module roots
  • ci: format campaign source tests and state the lowercase SHA rule
  • ci: isolate connector verifier dependencies
  • ci: pin void authorization Python
  • ci: trust quality gate attempt cap source
  • ci: validate quality gate void reasons
  • clients: send connector repair idempotency header
  • cloud: accept launch catalog schema v2
  • cloud: accept operation-valid usage supersets
  • cloud: address connector runtime review findings
  • cloud: address Gateway identity review
  • cloud: address staging rollback review
  • cloud: align estate module contracts
  • cloud: align preview posture lint config
  • cloud: allow console preview bootstrap
  • cloud: allow verified gateway final reuse
  • cloud: batch the token-capped encode lanes and price two more launch cells
  • cloud: bind connector intent to canonical worker
  • cloud: bind every priced-window factor to non-caller evidence
  • cloud: bind generation lanes to catalog hash
  • cloud: bind retained benchmark owners
  • cloud: bind retained recovery to candidate tasks
  • cloud: bind task evidence to reviewed cases
  • cloud: bind the billed report period to the priced window
  • cloud: bind vendor mutations to estate authority
  • cloud: bound Modal history reads separately
  • cloud: bound Modal inventory reads separately
  • cloud: close connector activation review gaps
  • cloud: close connector review findings
  • cloud: close connector review gaps
  • cloud: close two vacuous guards in the tier banding logic
  • cloud: declare the below-cost classes instead of hiding them
  • cloud: default estate_id inside resolve_estate
  • cloud: deny unused staging scheduler authority
  • cloud: derive allocation shares from measured work, not fiat
  • cloud: derive offered load from measured lane capacity
  • cloud: disable staging cold-start reporting
  • cloud: drop the exporter’s path workaround, annotate the document type
  • cloud: enforce console preview posture
  • cloud: fence selective Gateway dispatch by contract
  • cloud: finish connector protocol review sweep
  • cloud: fold the 2026-08-04 lane sweeps into the offered-load plan
  • cloud: forward selective deployment mode
  • cloud: gate connector activation on desired state
  • cloud: handle local generation stream failures
  • cloud: harden connector activation boundaries
  • cloud: harden deployment mode forwarding
  • cloud: harden edge backend ordering, cell parity, bucket names
  • cloud: harden managed generation streaming
  • cloud: harden Modal inventory validation
  • cloud: harden preview OIDC validation
  • cloud: harden registry identity validation and lock estate identity
  • cloud: harden selective deployment reuse
  • cloud: harden staging estate isolation
  • cloud: harden staging foundation identity
  • cloud: ignore local source caches
  • cloud: isolate console preview git config
  • cloud: isolate sidecar subprocess secrets
  • cloud: isolate staging estate authorities
  • cloud: make Vercel readback authoritative
  • cloud: measure the enlarged fixtures and price 18 more launch cells
  • cloud: merge main and add band-6 rerank capacities
  • cloud: normalize component source modes
  • cloud: omit read-only webhook rules
  • cloud: omit response-only webhook rules
  • cloud: order config image local layers
  • cloud: pin gateway plan evidence
  • cloud: pin sandbox Scalr modules at their first healthy versions
  • cloud: preserve connector execution topology
  • cloud: preserve positive sandbox activation target
  • cloud: preserve retained worker provenance in release gates
  • cloud: preserve valid rate book evidence
  • cloud: price the measured window, not the billed report hour
  • cloud: re-render siglip diagnostics lock after server source change
  • cloud: reconcile connector activation with main
  • cloud: record the measured decode length on the generation cells
  • cloud: recover native production gateway
  • cloud: recover retained candidates after newer merges
  • cloud: refresh combined topology evidence
  • cloud: refresh siglip diagnostic lock
  • cloud: refresh task contract test fixtures
  • cloud: refuse an unpublished class name instead of raising KeyError
  • cloud: reject empty connector inference text
  • cloud: reject JSON Terraform escapes
  • cloud: remove unused type suppression
  • cloud: report rejected reuse evidence
  • cloud: require integer estate schema version
  • cloud: reserve governed page ceilings
  • cloud: restrict promotion authority to full estates
  • cloud: retain legacy standalone tag gate
  • cloud: scope promotion authority to estate cells
  • cloud: stabilize component binary identity
  • cloud: stabilize named-entity canary fixture
  • cloud: state the audio_ms clip-length assumption in the quote
  • cloud: tighten connector manifest cleanup
  • cloud: tighten task contract evidence gates
  • cloud: unblock reproducible Modal convergence
  • cloud: use explicit connector protocol stubs
  • cloud: use non-ambiguous visual distractor
  • cloud: use representative vision smoke fixture
  • cloud: validate retained migration without activation
  • cloud: validate reused worker runtime CAS
  • cloud: validate selected gateway in catalog smoke
  • cloud: validate settled acceptance usage
  • cloud: verify Modal service-user access
  • colbert: honor transformers-5 rope_parameters when serving on the 4.x bundle
  • generation: align direct backend output limits
  • generation: defer resolved profile output limits
  • generation: fit Gemma 256k KV cache on H100
  • generation: handle family-specific reasoning boundaries
  • generation: hide split Gemma thinking boundary
  • generation: pair long-context qwen image bounds
  • generation: preserve Gemma reasoning boundaries
  • generation: preserve qwen vision bounds
  • generation: reject overflowing runtime numerics
  • generation: reserve Gemma long-context workspace
  • generation: unblock cold model startup
  • generation: use compatible Qwen 27B grammar backend
  • generation: use compatible Qwen 35B grammar backend
  • generation: validate high-context thinking profiles
  • modal: isolate overlay FlashInfer cache
  • modal: keep image yaml loading local
  • modal: support strict GPU allocation
  • modal: validate strict GPU selectors
  • model-ops: an unmeasured expected_dim never confirms a multivector smoke
  • model-ops: batch discovery evidence revision listing via one pinned tree fetch
  • model-ops: bind trusted terminal outcomes
  • model-ops: bound authority listing by model directories, not raw entries
  • model-ops: calibrate late-interaction and doc-retrieval scanner blocks
  • model-ops: close scaffold authorization gaps
  • model-ops: complete scaffolds before Stage 4
  • model-ops: exempt tooling stops from attempt cap
  • model-ops: fail closed on duplicate search errors
  • model-ops: harden scaffold validation boundaries
  • model-ops: ignore fenced tooling markers
  • model-ops: insert matrix rows once per models node so alias blocks ride their anchor
  • model-ops: isolate scaffold candidate tests
  • model-ops: namespace the config projection_dim claim and cap consumer-side nested ints
  • model-ops: order attempt comments deterministically
  • model-ops: preserve adapter auto authorization
  • model-ops: scope duplicate onboarding guard
  • model-ops: teach the Stage-4 smoke to probe multivector models
  • models: complete mLateOn serving recipe
  • quality-eval: harden runner repair
  • quality-eval: quarantine unregistered runners
  • quality-eval: shorten wake grace and map sglang bundles
  • quality-eval: shorten wake grace and support sglang logical bundles
  • quality-eval: use prepared scorer environment
  • release: bound commit history batches
  • release: deduplicate generated changelog notes
  • release: finalize source before diagnostics
  • release: refresh topology-bound diagnostics
  • release: reject repeated release headings
  • sdk: retry pre-execution SSE capacity errors
  • server: align Gemma xgrammar dependency
  • server: await standalone smoke teardown
  • server: bound Qwen visual context
  • server: constrain SGLang compatibility paths
  • server: enforce Qwen fast image bounds
  • server: keep Qwen3-VL reranker batches working on transformers 5.3
  • server: pin docling to its artifact revision and stop OCR gating load
  • server: preprocess native extract images
  • server: preserve LightOn SGLang compatibility
  • server: replace stale root playground with redirect to Swagger docs
  • sidecar: harden local generation cancellation
  • sidecar: reserve cancel request bytes
  • terraform: address internal module review
  • terraform: bound internal module lifecycle
  • terraform: exclude worktree caches from remote config uploads
  • terraform: harden Scalr module releases
  • terraform: package internal Scalr module archive root
  • terraform: package private Scalr registry modules
  • terraform: package Scalr module archive envelopes
  • terraform: register Scalr modules from plain monorepo roots
  • terraform: reject null module registry provider
  • terraform: relocate Scalr registry modules under deploy tree
  • terraform: validate sandbox module with OpenTofu
  • New capabilities: add H100 FP8 generation profiles; governed first-baseline mode + 16 vanilla floors for Qwen3-Embedding-8B and nomic v2 MoE
  • Reliability and operations: add protected retained recovery dispatch; bind modal env before retained recovery; bind retained gateway recovery to failed release; bound retained recovery runtime; clarify retained gateway recovery
  • model-catalog: add H100 FP8 generation profiles
  • quality-eval: governed first-baseline mode + 16 vanilla floors for Qwen3-Embedding-8B and nomic v2 MoE
  • bench: initialize buildx for worker image
  • bench: quiesce gc during timed campaigns
  • bench: validate non-root worker runtime
  • cloud: add protected retained recovery dispatch
  • cloud: address provision review findings
  • cloud: align rollback preflight diagnostic
  • cloud: allow full cold dispatch preflight budget
  • cloud: audit extracted provision secrets
  • cloud: bind modal env before retained recovery
  • cloud: bind retained gateway recovery to failed release
  • cloud: bind retained gateway to failed release
  • cloud: bound retained recovery runtime
  • cloud: clarify retained gateway recovery
  • cloud: close provision review gaps
  • cloud: converge config before cold gateway
  • cloud: fence initial gateway absence recovery
  • cloud: govern reconciliation cadence in releases
  • cloud: guard bootstrap and benchmark worker runtime
  • cloud: guard current-schema bootstrap seed
  • cloud: harden connector smoke recovery
  • cloud: harden extracted provision edges
  • cloud: harden retained recovery diagnostics
  • cloud: honor generation smoke cold starts
  • cloud: isolate config app runtime import
  • cloud: isolate connector smoke environments
  • cloud: isolate sibling smoke environments
  • cloud: keep connector diagnostic out of MVP release gate
  • cloud: keep prod generation warm and scope spans
  • cloud: make generation smoke single shot
  • cloud: narrow connector loopback scope
  • cloud: preserve accepted bundle promotion compatibility
  • cloud: preserve recovery fence diagnostics
  • cloud: recover first gateway to absence
  • cloud: recover retained legacy gateway
  • cloud: recover retained staging gateway to absence
  • cloud: recover staging gateway to absence
  • cloud: reject incomplete cold gateway plans
  • cloud: retain production deployment evidence
  • cloud: retain zero-floor generation attestations
  • gateway: preserve grammar-safe FP8 profiles
  • loadtest: prepare pinned perf-lab source before AWS
  • model-ops: bind scaffold AUTO to trusted evidence
  • model-ops: mirror task-equivalent standard blocks
  • preserve cold-start reporter membership history
  • quality-eval: preserve missing metric diagnostics
  • quality: run comparison through locked project
  • release: identify expired Modal attestation
  • release: keep Modal diagnostic trim ASCII-only
  • release: reconcile lagging Modal attestations
  • release: tolerate completed Modal exec targets
  • release: wait for refreshed PR head
  • validate initial reporter lease atomically
  • New capabilities: activate bootstrap-v2 pricing in prod; add —skip for recorded vendor-convergence skips; add cold-start live acceptance proof; add the mechanical launch-rate candidate promotion pipeline; add the size-class tier commercial policy to the rate book; admit rate-grade evidence in the launch catalog and coverage audit
  • Reliability and operations: admit an image-witnessed authoritative zero at settlement; align runner authorization with smoke protocol; async-surface reservation parity — video guard, byte cap, tokenizer allowance; bound gateway cold-start recovery; cap identity convergence retry seam
  • Performance: index the rotation-successor probe on api_keys
  • cloud: activate bootstrap-v2 pricing in prod
  • cloud: add —skip for recorded vendor-convergence skips
  • cloud: add cold-start live acceptance proof
  • cloud: add the mechanical launch-rate candidate promotion pipeline
  • cloud: add the size-class tier commercial policy to the rate book
  • cloud: admit rate-grade evidence in the launch catalog and coverage audit
  • cloud: batch/jobs generation via buffered single-terminal settlement
  • cloud: bound and measure the sealed cold-start subsidy
  • cloud: carry an authoritative-dimension claim set on UnitCounts
  • cloud: carry the settled charge into realtime usage blocks
  • cloud: commit reviewed launch allocations
  • cloud: editable per-key spend limits + monthly auto-recharge cap
  • cloud: editable per-key spend limits + monthly recharge cap
  • cloud: GA audio on /v1/extract — book-driven whisper pricing
  • cloud: gate production on a degraded staging acceptance
  • cloud: generation input-token reservation bound + settle-contract pins
  • cloud: gpu_second sealed billing regime — custom models become sellable
  • cloud: install digest-bound generation routes
  • cloud: install digest-bound governed generation routes
  • cloud: lease dynamic cold-start reporters
  • cloud: low-balance threshold email outbox + SES seam
  • cloud: low-balance threshold emails via SES
  • cloud: metering GA — waves 1–3
  • cloud: move the license exclusion verdict to the registry resolution seam
  • cloud: org-scoped usage export API — CSV + JSON
  • cloud: persist cold-start subsidy evidence
  • cloud: POST /v1/estimate dry-run cost endpoint
  • cloud: POST /v1/estimate dry-run cost endpoint + SDKs
  • cloud: production-bootstrap-v2 — prefill closed, sealed lanes priced
  • cloud: promote launch-rate candidates to rate-grade evidence
  • cloud: rate-grade size-class tier book assembly + operator-gated activation
  • cloud: response-final settlement for streaming generation
  • cloud: storage GiB-day accrual + book-driven register quote
  • cloud: surface credits_charged + rate_book_version in usage
  • cloud: unfence OpenAI-compat /v1/audio/transcriptions
  • cloud: unify the settled-charge name across jobs, batches, and SDKs
  • cloud: weekly/daily reconciliation with paging alerts and operator-gated LIVE mode
  • cloud: weekly/daily reconciliation, attribution categories and the operator LIVE gate
  • cloud: widen whole-job settlement to every contract dimension
  • gateway: add optional governed generation routing seam
  • model-ops: add ASR discovery family
  • model-ops: add fail-closed discovery family registry
  • model-ops: add multimodal discovery families
  • model-ops: add object detection discovery
  • model-ops: add OCR discovery families
  • model-ops: add report-only generation discovery
  • model-ops: add VQA and KIE discovery
  • model-ops: discover guardrail candidates
  • model-ops: discover sparse retrievers
  • model-ops: discover structured extraction candidates
  • model-ops: discover text classification candidates
  • model-ops: discover text rerankers
  • model-ops: read the executor’s trust clearance from the onboarding artifact
  • model-ops: register guardrail discovery family
  • model-ops: register structured extraction discovery
  • model-ops: register text classification discovery
  • model-ops: scan guardrail model candidates
  • model-ops: scan structured extraction candidates
  • model-ops: scan text classification candidates
  • model-ops: support report-only family scans
  • model-ops: wire the trust_remote_code scan into the AUTO gate
  • models: onboard Qwen/Qwen3-Embedding-8B (scaffolded)
  • sdk: estimate() on both SDKs for the /v1/estimate dry run
  • server: bill sampled video frames as images
  • telemetry: export the sealed cold-start metric remotely and chart it
  • benchmark: validate sorted deployment evidence
  • catalog: bound GLiNER sequence lengths
  • ci: reseal staging catalog evidence
  • cloud: accept client-scoped WorkOS sessions
  • cloud: accept Terraform canonical false in bootstrap guard
  • cloud: accept WorkOS environment issuers
  • cloud: address billing bootstrap review
  • cloud: address CodeRabbit review — SES client bounds, send pacing, re-arm audit
  • cloud: address CodeRabbit review on the /v1/estimate dry run
  • cloud: address CodeRabbit review on the key-limit + recharge-cap PR
  • cloud: address CodeRabbit review on the settled-charge surfaces
  • cloud: address CodeRabbit review on the usage export
  • cloud: address prerequisite gate review
  • cloud: address the CodeRabbit review on the batch/jobs generation PR
  • cloud: address the CodeRabbit round on the metering GA integration
  • cloud: address the second CodeRabbit round on /v1/estimate
  • cloud: admit an image-witnessed authoritative zero at settlement
  • cloud: admit exact task-role tag bootstrap
  • cloud: admit observed Modal US locations
  • cloud: align benchmark smoke with rated profile
  • cloud: align runner authorization with smoke protocol
  • cloud: allow nested ECR repositories
  • cloud: async-surface reservation parity — video guard, byte cap, tokenizer allowance
  • cloud: attest benchmark runtime identity
  • cloud: bind a rate_grade key to the price and hardware its envelope measured
  • cloud: bind nightly health checks to staging
  • cloud: bind paused forwarder candidate before prod cutover
  • cloud: bind the charge-slot backstop to the stream, not the disconnect grace
  • cloud: block the generation in/out split on a named measurement
  • cloud: bound gateway cold-start recovery
  • cloud: bound nightly evidence execution
  • cloud: bound the sealed resolver cache and cap proxy response bodies
  • cloud: cap identity convergence retry seam
  • cloud: claim the hold before releasing it, and fold case like the registry
  • cloud: classify benchmark evidence failures
  • cloud: close cold recovery review gaps
  • cloud: close cold-start evidence failure modes
  • cloud: close final metering review faults
  • cloud: close meter-only review gaps
  • cloud: close metering GA review findings
  • cloud: close rate evidence review gaps
  • cloud: close reporter lease invariants
  • cloud: close six gate findings on the /v1/estimate dry run
  • cloud: close the batch-2a gate findings on editable limits + recharge cap
  • cloud: close the CodeRabbit round on the candidate-promotion branch
  • cloud: close the CodeRabbit round on the integration head
  • cloud: close the CodeRabbit round on the reconciliation branch
  • cloud: close the CodeRabbit round on the waves-1-3 integration head
  • cloud: close the final CodeRabbit round on /v1/estimate
  • cloud: close the gate findings on the resolution-seam verdict
  • cloud: close the promotion gate’s operator-assertion escapes
  • cloud: close the second CodeRabbit round on the credit surfaces
  • cloud: CodeRabbit review — extract params must be explicit per line
  • cloud: CodeRabbit review fixes — bounded sealed-class cache, doc reconciliation
  • cloud: CodeRabbit review fixes — quote authority degradation, single-model projection
  • cloud: CodeRabbit review fixes — sealed posture floor, generator provenance
  • cloud: CodeRabbit round 2 — evidence doc, strict decode, single query walk
  • cloud: CodeRabbit round 3 — exhaustive ceiling match, honest parse row
  • cloud: CodeRabbit round on the control-plane prerequisite gate
  • cloud: CodeRabbit round-2 — redact resolver errors, tighten class guard
  • cloud: CodeRabbit round-6 — quote survives a dead pool, pin e2e clock
  • cloud: CodeRabbit round-7 — atomic cache admission, sweep resilience
  • cloud: compile exact generation customer routes
  • cloud: compile exact generation routes
  • cloud: correct console wallet path to /console/usage — verified against sie-web
  • cloud: correct the race() return annotation in the snapshot test
  • cloud: cover regular fleet cold starts
  • cloud: cover regular Modal fleet cold starts
  • cloud: decide license exclusion at the registry resolution seam
  • cloud: declare generation input_tokens in the launch catalog
  • cloud: default customer charging to disabled
  • cloud: disable recharge schedule without teardown
  • cloud: distinguish authoritative zero from missing count at settlement
  • cloud: enforce the image witness on both sides of the rating boundary
  • cloud: expose collector test configs to non-root
  • cloud: fail a tier book whose price point has no exact Metronome decimal
  • cloud: fail closed on incomplete preflight reads
  • cloud: fail closed when a declared skip cannot be recorded
  • cloud: fail loudly when test prerequisites vanish
  • cloud: gate batch usage on settlement evidence, not a running counter
  • cloud: gate estimates on resolved model identity
  • cloud: gate fixes for the usage export — day-sliced scan, pool isolation, uniform masking
  • cloud: harden cold-start acceptance proof
  • cloud: harden dynamic cold-start membership
  • cloud: harden staging release verification
  • cloud: harden the multipart licensing gate and lock the route wiring
  • cloud: index count-gated SES resources via one()/splat
  • cloud: isolate benchmark evidence runtime
  • cloud: log truncated usage-export streams; pin legacy-row null passthrough
  • cloud: make extract canonical for Docling billing
  • cloud: make generation rate evidence prefix-cache safe
  • cloud: make test prerequisites and tracing deterministic
  • cloud: make the jobs loopback estimate allowance-aware and prove the guard end to end
  • cloud: make the jobs planner uphold the invariant settlement now enforces
  • cloud: make the reconciliation bars statements the cost basis supports
  • cloud: make the response cap operative, unbilled, and honest
  • cloud: make the witnessed zero reach the wire and cover both rails
  • cloud: merge candidate promotion and close #2628 review
  • cloud: mirror the widened model canonicalization in the control plane
  • cloud: observe the sealed cold-start metric; scope it to the prometheus path
  • cloud: parse the multipart model under the route’s real body limit
  • cloud: pin benchmark credential versions
  • cloud: pin the sealed tariff evidence; harden the split and the posture tests
  • cloud: pin the storage eligibility cutoff to UTC
  • cloud: preflight input token rate bounds
  • cloud: preserve cold-start evidence gaps
  • cloud: preserve inactive prod rollback tasks
  • cloud: price sealed-lane memory at the Sandbox rate, not standard
  • cloud: propagate connector request identity
  • cloud: publish a committed charge even when settlement then faults
  • cloud: pull the authoritative invoice inside the single-flight lock
  • cloud: recipient chain, re-arm edge, staleness bound, per-row commit
  • cloud: reconcile what the batchgen merge hid from git
  • cloud: record an admission outcome on every licensing rejection
  • cloud: redact benchmark operator preflight
  • cloud: reject a negative storage-accrual lookback
  • cloud: reject an explicit storage-accrual day that is not closed
  • cloud: reject expired recovery evidence
  • cloud: reject null WorkOS client claims
  • cloud: release the hold when the gateway cannot deliver the result
  • cloud: renumber monthly-cap migration to 0040 — 0039 taken by sibling W3 branch
  • cloud: renumber rotation-successor index migration to 0043 — 0042 taken by #2580
  • cloud: report diagnostic cleanup failure
  • cloud: report every live-preflight failure in one pass
  • cloud: report malformed release pointers
  • cloud: reseal release catalog dependents
  • cloud: restore the spend-limit error contract on the W3 base
  • cloud: retain exact benchmark deployment evidence
  • cloud: retain exact benchmark deployment evidence\n\nRefs #2573
  • cloud: retry benchmark identity convergence
  • cloud: round-8 minors — env var name, authority restore, day-sweep guard
  • cloud: route expected no-charge stream releases off the fault alerts
  • cloud: route readiness failures through the gate without widening the opt-out
  • cloud: route the async reservation through the shared video and token seams
  • cloud: scope authoritative zero evidence
  • cloud: scope rate preflight to exact authority
  • cloud: scope the generation input bound to books that price input
  • cloud: seal generation campaign artifact paths
  • cloud: tell the truth about the video byte rail, and bound the tests that prove it
  • cloud: tighten console origin validation, pace off the previous row
  • cloud: tighten evidence diagnostic contract
  • cloud: treat an explicit null usage as an empty block, not a refusal
  • cloud: treat an explicitly null video as absent, not as a video
  • cloud: type-gate the chat_template_kwargs allowlist
  • cloud: unblock billing bootstrap migrations
  • cloud: unseal console-free release subsets and offload-smoke
  • cloud: unseal sealed release subsets, report every preflight failure, and gate degraded acceptance
  • cloud: valid IAM tag charset + keep bootstrap secret-gate count at 2
  • cloud: validate tier scheme inputs canonically
  • cloud: wire the meter-only operator gate
  • gateway: cap embeddings requests at 256 inputs
  • gateway: harden governed generation dispatch
  • gateway: preserve self-host grammar routing
  • loadtest: use job token for GHCR pulls
  • model-ops: accept GLiNER discovery models
  • model-ops: address authorization review feedback
  • model-ops: address evidence review feedback
  • model-ops: authenticate the evidence producer and bind the identity to execution
  • model-ops: bind CodeRabbit review completion
  • model-ops: bind crash fallback repository
  • model-ops: bind executor authorization
  • model-ops: bind trusted completion status
  • model-ops: close final trust gate review gaps
  • model-ops: close malformed trust evidence gap
  • model-ops: close three bypasses in the trust_remote_code gate
  • model-ops: dispatch the Stage 4-5 gate completion explicitly
  • model-ops: distinguish disabled discovery families
  • model-ops: exclude OCR-only KIE models
  • model-ops: explicitly dispatch discovery handoffs
  • model-ops: gate readiness on bot reviews
  • model-ops: harden bot review handoff
  • model-ops: harden cold Modal image builds
  • model-ops: harden completion status fallbacks
  • model-ops: harden report-only scanning
  • model-ops: harden trust gate provenance
  • model-ops: keep a failed gate origin recoverable
  • model-ops: narrow document OCR compatibility
  • model-ops: never log gh’s stderr, which carries GH_TOKEN by construction
  • model-ops: preserve completed review mutations
  • model-ops: re-derive the clearance in matrix-block, bind evidence to the revision
  • model-ops: re-grant contents:read to the dispatch job
  • model-ops: recognize Copilot bot aliases
  • model-ops: reconcile gate recovery reviews
  • model-ops: refuse an unset RUNNER_TEMP instead of falling back into the workspace
  • model-ops: refuse instead of crashing when the evidence fetch faults
  • model-ops: refuse unusable evidence instead of falling back to blockers
  • model-ops: reject ambiguous family matches
  • model-ops: reject failed Copilot reviews
  • model-ops: repair gate recovery replay
  • model-ops: retry transient Modal build downloads
  • model-ops: scan all loader-visible code
  • model-ops: scan zero-shot classification models
  • model-ops: suppress dynamic artifact failure logging
  • model-ops: tolerate pending status retries
  • model-ops: validate gate head before status writes
  • model-ops: validate review completion evidence
  • model-ops: wake the adapter executor from an agentless dispatch job
  • model-ops: wake the adapter-exec executor by explicit dispatch
  • quality-eval: retain models pending vanilla floors
  • quality-eval: surface unfloored matrix rows
  • release: bootstrap a fresh production database
  • release: compare fresh proof types exactly
  • release: retry malformed Modal attestations
  • release: retry transient Modal attestation output
  • release: verify empty replay target digest
  • server: address CodeRabbit review on the video meter
  • server: bill SigLIP padded text work
  • server: bound isolation fan-out and close the NaN decode escapes
  • cloud: index the rotation-successor probe on api_keys
  • New capabilities: add agent shell GPU fallbacks; add agent shell placement controls; add launch catalog contract canary; add Modal-native placement primitives; add native placement primitives; allow snapshotless native lanes
  • Reliability and operations: align rate-book recovery checks; harden dispatch preflight gate; harden final generation gates; harden retained release recovery; preserve capacity diagnostics on timeout races
  • Performance: cache MUVERA projection state
  • cloud: add agent shell GPU fallbacks
  • cloud: add agent shell placement controls
  • cloud: add launch catalog contract canary
  • cloud: add Modal-native placement primitives
  • cloud: add native placement primitives
  • cloud: allow snapshotless native lanes
  • cloud: attest realized placement identity
  • cloud: establish broad-native release baseline
  • cloud: govern prod legacy billing cutover
  • cloud: preflight offline tokenizer dependencies
  • server: pinnable LoRA adapter revisions in lora_paths
  • cloud: accept scheduled orphan-hold sweeps
  • cloud: address native placement review
  • cloud: address release review feedback
  • cloud: align rate-book recovery checks
  • cloud: align tokenizer snapshot dependencies
  • cloud: allow absent generation in final smoke
  • cloud: allow bounded prod platform bootstrap
  • cloud: attest final gateway replacement
  • cloud: bound snapshot lifecycle boots
  • cloud: bound worker identity collection
  • cloud: canonicalize generation GPU profiles
  • cloud: close bounded prod bootstrap plan
  • cloud: close launch review findings
  • cloud: declare bootstrap alert container
  • cloud: default orphan hold sweep ttl
  • cloud: derive exact rate topology from broad release
  • cloud: derive gateway gpu admission from plan
  • cloud: diagnose bounded worker capacity waits
  • cloud: drain legacy generation revisions
  • cloud: enforce global canary request identity
  • cloud: fail closed on tokenizer pin drift
  • cloud: finalize gateway after generation deploy
  • cloud: frame reusable candidate outputs
  • cloud: gate billable release commit safely
  • cloud: gate every gateway replica on dispatch readiness
  • cloud: guard broad placement pending realized identity
  • cloud: harden dispatch preflight gate
  • cloud: harden final generation gates
  • cloud: harden retained release recovery
  • cloud: integrate launch remediation release gates
  • cloud: keep slow batch leaders warm
  • cloud: make Vercel deploy creation transport-safe
  • cloud: package excluded gateway target
  • cloud: package gateway from crate target
  • cloud: preflight gateway generation dispatch
  • cloud: preflight route and worker hash parity
  • cloud: preserve capacity diagnostics on timeout races
  • cloud: preserve launch route errors
  • cloud: preserve snapshot cancellation cleanup
  • cloud: reconcile launch remediation release gates
  • cloud: recover ambiguous Vercel deploy creates
  • cloud: recover bounded prod bootstrap outputs
  • cloud: recover rate-book releases drain-first
  • cloud: redact readiness diagnostics
  • cloud: refresh staging catalog evidence digests
  • cloud: rehearse immediate launch accounts
  • cloud: reject partial readiness placement
  • cloud: reserve token inputs from governed bounds
  • cloud: respect Better Stack token scope
  • cloud: restore authenticated gateway rollback health
  • cloud: restrict rehearsal recovery fields
  • cloud: serialize bootstrap state reads
  • eval: make run-eval-group.sh and its tests portable to macOS
  • eval: reject symlink escapes from the /tmp guard; reap the requirements temp dir
  • loadtest: use valid AWS IAM purpose tag
  • models: declare tokenizer dependencies
  • quality: address remaining review findings
  • quality: close nightly model gaps and add exact quarantine
  • quality: gate bounded NVIDIA MUVERA checks
  • quality: isolate parameterized FiNER evals
  • quality: preserve floor metric contracts
  • quality: record bounded concurrency-one targets
  • quality: record faithful CMedQA floors
  • quality: record faithful quality floors
  • quality: reject non-finite qrel scores
  • quality: restore bounded FiNER coverage
  • quality: restore Mixedbread query expansion
  • quality: stabilize scoped PR quality runs
  • release: close prod bootstrap output contract
  • release: redact bootstrap state read failures
  • smoke: restore blocking-probe assertions and cold-touch retry coverage
  • worker: copy sie_telemetry crate into worker image builds
  • server: cache MUVERA projection state
  • New capabilities: accept native multimodal generation; support routed media measurement campaigns; add production bootstrap rate book; add pre-public task measurement candidates; bind launch catalog to commercial rate book; bind launch rates to commercial book
  • Reliability and operations: harden generation quality sharding; harden multimodal batch validation; harden queue readiness for generation defaults; harden shard failure capture; restore governed retrieval floors
  • Performance: bind bounded extraction evidence; bind current nonvision launch evidence; recover granite guardian queue evidence; reuse siglip preprocessing
  • api: accept native multimodal generation
  • bench: support routed media measurement campaigns
  • billing: add production bootstrap rate book
  • catalog: add pre-public task measurement candidates
  • catalog: bind launch catalog to commercial rate book
  • catalog: bind launch rates to commercial book
  • catalog: prepare launch serving evidence
  • catalog: prepare non-public launch catalog checkpoint
  • catalog: record current nonvision evidence
  • catalog: record Snowflake and R3 evidence
  • catalog: validate native multimodal recipes
  • catalog: verify document vision recipes
  • catalog: verify existing launch recipes
  • catalog: verify GLiNER2 large launch evidence
  • catalog: verify R3 pipeline on L4
  • catalog: verify redaction task evidence
  • catalog: verify Snowflake on preferred L4
  • ci: complete req14 gates after lane runs
  • ci: remove runner-held waits from official-recipe gates
  • cloud: activate accounts and route launch events
  • cloud: add bounded prod platform bootstrap
  • cloud: add exact launch rate campaign suite
  • cloud: add immutable managed release pipeline
  • cloud: add launch account credit flow
  • cloud: add launch rate campaign suite
  • cloud: add Qwen task evidence campaign
  • cloud: add secure staging catalog evidence dispatch
  • cloud: add secure staging evidence dispatch
  • cloud: archive Modal measurement releases
  • cloud: assemble commercial rates from exact evidence
  • cloud: authenticate Modal rate evidence
  • cloud: bracket generation rate window
  • cloud: compile isolated measurement topology
  • cloud: derive commercial rates from exact resource seconds
  • cloud: derive commercial rates from Modal billing
  • cloud: execute document vision recipes
  • cloud: execute remaining reviewed recipes
  • cloud: route signup notifications directly to Slack
  • cloud: verify BGE and GTE launch evidence
  • deploy: add scalr-bootstrap root for internal terraform state
  • deploy: run scalr-bootstrap under opentofu with live scalr backend
  • deploy: scalr-held terraform state for internal roots
  • deploy: split scalr workspaces into cli- and vcs-driven
  • generation: type structured output contracts
  • model-ops: split feasibility PROCEED-WITH-ADAPTER-AUTO vs human
  • model-ops: trust_remote_code security pre-gate
  • quality-eval: in-pipeline wall-clock stop with partial report
  • sdk: add typed Responses client
  • 1841: stamp sealed metric units on the Modal collector, and make the CP filters cover what the suite reads
  • adapters: serialize output_hidden_states forwards — close the #2144 recorder-race class
  • adapters: serialize output_hidden_states forwards to close the #2144 recorder-race class
  • api: address multimodal contract review
  • api: close multimodal parity gaps
  • bench: add faithful PyLate official recipe
  • bench: address launch evidence review
  • bench: align Any2Any dispatch semantics
  • bench: align ColBERT reference validation
  • bench: align pinned dataset metadata
  • bench: bind generation evidence to runtime protocol
  • bench: bind Modal measurement provenance
  • bench: bind Qwen3.6 launch evidence to runtime identity
  • bench: canonicalize campaign route identities
  • bench: clean up failed server startups
  • bench: close operation evidence review gaps
  • bench: discard partial server results
  • bench: emit strict quality JSON
  • bench: fail closed on benchmark operation identity
  • bench: forward only encode/score to EvalRunner in quality path
  • bench: harden generation quality sharding
  • bench: harden multimodal batch validation
  • bench: harden queue readiness for generation defaults
  • bench: harden shard failure capture
  • bench: keep image types out of runtime imports
  • bench: pin ragged PyLate reranking fix
  • bench: pin trusted dataset loaders
  • bench: preserve canonical FewRel dataset identity
  • bench: preserve generation evidence imports
  • bench: preserve invalid campaign archives
  • bench: preserve profile benchmark identity
  • bench: preserve safe Modal shard failures
  • bench: reject extraction item errors
  • bench: reject incompatible reference operations early
  • bench: reject incomplete shard assembly
  • bench: reject partial performance measurements
  • bench: remove stale generation import
  • bench: report empty perf evidence explicitly, not as ambiguous
  • bench: restore governed retrieval floors
  • bench: route profiled performance requests
  • bench: satisfy review static analysis
  • bench: start named profile routes
  • bench: stop plural options poisoning choice extraction
  • bench: type bounded image dispatch
  • bench: validate multimodal encode batches
  • bench: version routed rate candidates
  • billing: bind Modal tags to pretraffic contract
  • billing: bind zero-page job evidence
  • billing: preserve zero-page job evidence
  • billing: project observed units through reservation plans
  • billing: restrict zero terminals to pages
  • billing: verify multimodal fixture units
  • catalog: bind named output evidence
  • catalog: disable unsafe qwen speculative default
  • catalog: harden native acceptance validation
  • catalog: raise GTE Muvera repetitions
  • catalog: route mxbai edge through native attention
  • catalog: use pinned GTE query length
  • catalog: validate named profile evidence
  • ci: authenticate quality dependency sync
  • ci: diff quality evals from current PR base
  • ci: fail closed on mixed Modal output
  • ci: isolate privileged req14 completion
  • ci: isolate quality eval evidence
  • ci: isolate quality runner credentials
  • ci: let scalr-bootstrap track the opentofu pin
  • ci: repair Qwen quality shard workflow
  • ci: verify trusted req14 callback checkout
  • cloud: accept canonical Modal volume mount
  • cloud: accept Modal GCP provider identity
  • cloud: address release workflow review follow-ups
  • cloud: align launch evidence contracts
  • cloud: align release lock namespace
  • cloud: align release probes with scoped listing
  • cloud: allow release image manifest readback
  • cloud: allow runtime secret namespace
  • cloud: allow scoped release key discovery
  • cloud: allowlist release child environments
  • cloud: approve permissive launch licenses
  • cloud: attest LightOnOCR image compatibility
  • cloud: attest Modal snapshot memory limits
  • cloud: bind billing to executing Modal app
  • cloud: bind billing to Modal environment
  • cloud: bind commercial evidence to routed candidates
  • cloud: bind ECR login to OIDC identity
  • cloud: bind rate probes to archived model pins
  • cloud: bind score probes to campaign batches
  • cloud: bind topology plan semantics
  • cloud: bootstrap purpose-bound Slack sinks
  • cloud: bound and isolate release preflight reads
  • cloud: classify ECS start failures first
  • cloud: complete managed release operability follow-ups
  • cloud: complete release operability audit
  • cloud: dispatch exact catalog profiles
  • cloud: dispatch quality shards in Modal batch mode
  • cloud: fail closed in candidate image lookup
  • cloud: fail closed on ECR lookup errors
  • cloud: finish release child isolation
  • cloud: finish release preflight hardening
  • cloud: guard acceptance projection freshness
  • cloud: guard catalog projection freshness
  • cloud: harden commercial evidence joins
  • cloud: harden launch identity and approval invariants
  • cloud: harden launch rate campaign probes
  • cloud: harden managed release promotion
  • cloud: harden notification launch paths
  • cloud: harden screenshot acceptance
  • cloud: harden vision acceptance bounds
  • cloud: include smart chat task evidence
  • cloud: isolate launch credit approvals
  • cloud: keep LightOn preflight CPU-safe
  • cloud: make topology imports typecheck
  • cloud: map static lanes to catalog pools
  • cloud: measure both SigLIP towers
  • cloud: pin batch worker execution identity
  • cloud: pin dev worker identity
  • cloud: pin launch measurement resources
  • cloud: pin native worker execution identity
  • cloud: pin staging batch memory limit
  • cloud: preflight live release dependencies
  • cloud: preflight Modal SGLang runtime
  • cloud: preflight release task arguments
  • cloud: preserve concurrent topology output
  • cloud: preserve governed release tool paths
  • cloud: preserve sidecar image contract
  • cloud: prime catalog before gateway rollout
  • cloud: probe exact campaign request batches
  • cloud: reconcile commercial rates with launch profiles
  • cloud: reconcile measurement topology with launch catalog
  • cloud: refresh staging evidence pins
  • cloud: reject duplicate generation shards
  • cloud: reject malformed release metadata
  • cloud: reject mismatched catalog evidence
  • cloud: reject undecodable evidence
  • cloud: reject unversioned S3 evidence
  • cloud: require complete pii redaction proof
  • cloud: reseal licensed launch release locks
  • cloud: reseal production rate release locks
  • cloud: restore CUDA toolkit to SGLang sidecars
  • cloud: retain published measurement evidence
  • cloud: retire legacy environment ECR state
  • cloud: scope Better Stack credentials by product
  • cloud: separate guardrail policy fixtures
  • cloud: use lightweight generation evidence contract
  • cloud: validate launch evidence integrity
  • cloud: verify control-plane runtime dependencies
  • cloud: verify Modal markers from files
  • cloud: verify redact pii acceptance
  • cloud: verify redact PII postprocessing
  • deploy: pin internal scalr workspaces to opentofu
  • gateway: validate schema keywords by context
  • generation: close structured grammar review gaps
  • grammar: bound pointer index parsing
  • grammar: harden schema normalization
  • loadtest: manage github-token in terraform + sync via ESO, pin sie-perf-lab ref
  • loadtest: persist github-token in terraform, sync via ESO, pin perf-lab ref
  • loadtest: remove illegal variable interpolation from output description
  • model-ops: close review gaps in the AUTO verdict contract
  • model-ops: require an explicit matrix_task key in auto_authoring
  • profiles: honor routed outputs and retrieval recipes
  • quality-eval: address Copilot review on the wall-clock stop
  • quality-eval: remove shadow tree on exception paths, not just returns
  • quality-eval: select governed official metric
  • quality-eval: serve non-default smoke bundles from a pip —target shadow so the overlay reaches the bundle floor
  • quality-eval: validate wall_budget_minutes inside run_pipeline
  • release: authenticate gateway rollback probe
  • release: keep sie-audio-prep internal to builds
  • sdk: accept nullable grammar metadata
  • sdk: keep grammar types on supported module path
  • security: include aws edge in standing review
  • security: pin hf_revision on the 10 unpinned trust_remote_code models
  • server: apply trained Rotary ColBERT projection
  • server: defer LightOn OCR compatibility patch
  • server: hide unsupported visual muvera profiles
  • server: mask ColBERT expansion keys in flash attention
  • server: meter SigLIP tokens from attention masks
  • server: pin launch guardrail policies
  • server: preserve stream helper compatibility
  • server: preserve trusted code revision on fallback
  • server: restore ColBERTv2 retrieval recipe
  • server: restore Jina ColBERT retrieval recipe
  • server: serialize remaining tokenizer access
  • sie_server: register remote-code snapshot dir on sys.path for ST-Router models
  • sie-bench: isolate PyLate reference evaluation
  • sie-bench: preserve expansion token attention
  • smoke: fail loud when a bundle overlay undershoots its declared pins
  • terraform: manage github-token container only, drop placeholder version
  • tooling: derive overlay cache version from lock
  • tooling: gate Modal quality shard fanout
  • tooling: harden launch failure handling
  • tooling: isolate Modal shell worktree environment
  • tooling: make empty Modal status succeed
  • tooling: mount tracked playbook fixtures on Modal
  • tooling: namespace Modal JIT caches by ABI
  • tooling: preserve Modal status errors
  • tooling: preserve multiline Modal commands
  • tooling: read Modal lock from mounted workspace
  • tooling: resolve modal from task path
  • tooling: scope Modal inventory conflicts
  • tooling: validate Modal description types
  • tooling: version Modal datasets cache
  • vision: harden launch evidence
  • vision: pad fused detection batches
  • catalog: bind bounded extraction evidence
  • catalog: bind current nonvision launch evidence
  • generation: recover granite guardian queue evidence
  • vision: reuse siglip preprocessing
  • New capabilities: bind exact rate measurement windows; shard OCR quality by page; support package revision evidence; attest worker execution identity; fan out quality shards on Modal; pin immutable Docling artifacts
  • Reliability and operations: harden OCR shard provenance; harden package revision evidence; fail fast on managed gateway lock drift; attest OCR shard execution identity; clarify complete rate evidence
  • Performance: concurrent OCR page dispatch so server batching actually engages
  • bench: bind exact rate measurement windows
  • bench: shard OCR quality by page
  • bench: support package revision evidence
  • cloud: attest worker execution identity
  • cloud: fan out quality shards on Modal
  • server: pin immutable Docling artifacts
  • server: support immutable Docling model artifacts
  • tooling: attach named Modal shell secrets
  • bench: attest OCR shard execution identity
  • bench: clarify complete rate evidence
  • bench: clarify rate-window mismatch errors
  • bench: harden OCR shard provenance
  • bench: harden package revision evidence
  • bench: mark serving basis non-commercial
  • bench: require exact rate resource evidence
  • ci: install sdk deps for catalog tests
  • ci: queue release rerun when main moves
  • ci: regenerate launch catalog on release PR
  • ci: regenerate release PR when main moves
  • cloud: bind benchmark evidence to deployed workers
  • cloud: bind benchmark model identity
  • cloud: bind quiesced ECS candidates before scaling
  • cloud: bind reviewed catalog recipe executors
  • cloud: bind worker resource attestation
  • cloud: fail fast on managed gateway lock drift
  • cloud: include compiler version dependency
  • cloud: isolate catalog compiler dependencies
  • cloud: meter detection catalog smoke images
  • cloud: preserve deployment validation ordering
  • cloud: preserve identity across billing handoff
  • cloud: reject noninteger GPU identity counts
  • cloud: remove member key spend ceiling
  • cloud: require clean governed release sources
  • cloud: validate integral worker identity fields
  • cloud: validate managed gateway lock before release
  • eval: constrain temporary artifact root
  • eval: serve canonical profile routes
  • tooling: bind release locks from clean head
  • tooling: close Modal fanout lifecycle races
  • tooling: fail closed on Modal volume sync
  • tooling: isolate Modal uv cache
  • tooling: isolate quality shards on Modal v2
  • tooling: keep active modal shells alive
  • tooling: preserve Modal fanout failures
  • tooling: reject symlink fanout outputs
  • tooling: resolve modal in detached shells
  • sie-bench: concurrent OCR page dispatch so server batching actually engages
  • New capabilities: add rate-grade evidence contract; add reusable local queue stack; add upstream GLiNER NER reference; govern vision launch evidence; seed load-test request selection; SIESparseImageWrapper for served visual sparse retrieval
  • Reliability and operations: capture generation timeout evidence; harden local queue stack lifecycle; drop stale metrics registry dependency; finalize billing money hardening; harden account lifecycle rehearsal
  • Performance: record SPLADE 512-token run; record extraction L4 measurements; compact sparse output in native precision; fuse packed SPLADE pooling
  • bench: add rate-grade evidence contract
  • bench: add reusable local queue stack
  • bench: add upstream GLiNER NER reference
  • bench: govern vision launch evidence
  • bench: seed load-test request selection
  • bench: SIESparseImageWrapper for served visual sparse retrieval
  • billing: meter multimodal catalog units
  • candle: add and optimize native SPLADE sparse encoding
  • candle: add native SPLADE sparse encoding
  • cloud: add exact qwen35 perf probe
  • cloud: add fail-closed catalog acceptance gate #1943
  • cloud: add fail-closed catalog acceptance runner
  • cloud: add fast chat catalog acceptance
  • cloud: add managed-edge abuse and tenant gates
  • cloud: add OCR and speech acceptance probes
  • cloud: add operator key administration
  • cloud: add operator key administration and keyed usage
  • cloud: add OWLv2 acceptance executor
  • cloud: add OWLv2 catalog acceptance
  • cloud: add staging account lifecycle rehearsal
  • cloud: add staging rehearsal rate authority
  • cloud: audit measured launch rate-book coverage
  • cloud: audit measured rate book coverage
  • cloud: bind verified DINO and Whisper evidence
  • cloud: declare exact lane placement resources
  • cloud: declare exact lane resources
  • cloud: enforce conserved usage rating
  • cloud: enforce exact managed catalog authority
  • cloud: execute text acceptance recipes
  • cloud: exercise DINO catalog acceptance
  • cloud: expose operator billing health
  • cloud: gate catalog tasks on evidence
  • cloud: govern launch catalog topology
  • cloud: govern vision launch catalog
  • cloud: integrate control-plane launch readiness
  • cloud: integrate exact request billing authority
  • cloud: notify operators of new signups
  • cloud: pack Whisper into shared launch pool
  • cloud: record generation launch evidence
  • cloud: record Qwen3.5 launch evidence
  • cloud: rotate staging rehearsal rate book
  • cloud: serve release-locked model catalog
  • cloud: serve release-locked model catalog #1944
  • cloud: settle exact pair and generation usage
  • cloud: ship proprietary vision launch catalog
  • cloud: validate extraction catalog recipes
  • cloud: validate structured output acceptance
  • cloud: verify qwen35 A100 model evidence
  • cloud: wire privacy-safe console RUM
  • cluster: opt-in bake cache write via SIE_BAKE_CACHE_WRITE
  • core: expose authoritative request usage
  • core: report authoritative score pair usage
  • evidence: propagate worker execution identity
  • extract: add pinned PII quality gate
  • gateway: add strict Cohere rerank compatibility
  • model-ops: feasibility accepts transformers-5-line bundle choice
  • model-ops: scaffold re-renders cloud release locks with the model YAML
  • models: onboard Snowflake/snowflake-arctic-embed-s (scaffolded)
  • models: raise qwen36 launch context to 8k
  • playbooks: upstream SparseEncoder reference arm for sparse-visual models
  • quality-eval: wire the signature-DB loop — adjudicated-skip + verdict write-back
  • reranking: aggregate launch runtime and release locks
  • reranking: harden Qwen score runtime
  • sdk: expose terminal billing metadata
  • sdk: preserve terminal request metadata on errors
  • serve visual-SPLADE via transformers5 (sparse-vision adapter + measurement path)
  • server: bind package artifact identity
  • server: bind package-backed model artifacts
  • server: transformers514 bundle + SparseEncoder vision sparse adapter
  • sie-bench: add generation quality sharding
  • sie-bench: add generic generation quality sharding
  • telemetry: centralize observability on OTLP
  • vision: harden OSS runtime contracts
  • vision: ship OSS launch runtime and evidence
  • audio: sandbox image builds the audio wheel opportunistically
  • bench: align local generation lane routing
  • bench: attest generation execution revision
  • bench: bind broker artifact ARN
  • bench: capture generation timeout evidence
  • bench: clarify direct revision semantics
  • bench: close execution identity provenance gaps
  • bench: exercise generation queue sidecar
  • bench: harden local queue stack lifecycle
  • bench: isolate generation smoke stack lifecycle
  • bench: isolate queue smoke ports
  • bench: make fake stack no-op explicit
  • bench: preserve ChemProt entity offsets
  • bench: preserve exact reranker task identity
  • bench: preserve NER task parameters
  • bench: reject coerced execution identity
  • bench: remove failure log content
  • bench: render generation smoke help
  • bench: resolve canonical AG News dataset
  • bench: route canonical model profiles
  • bench: select encode output families
  • bench: separate model and machine profiles
  • bench: serialize image rerank inputs
  • bench: use SPIA field labels for PII eval
  • bench: verify broker artifact bucket before reads
  • bench: warm generation at measured concurrency
  • billing: preserve legacy outbox drain
  • billing: reject unexpected metronome rates
  • billing: verify vendor-bound raw usage
  • candle: bound CUDA compiler parallelism
  • candle: include CUDA math constants for SPLADE
  • candle: preserve tiny SPLADE log1p weights
  • ci: bind cloud gateway cache to OSS source
  • ci: install catalog compiler dependency
  • ci: load sidecar image for release smoke tests
  • ci: make cloud gateway checks deterministic
  • ci: supersede obsolete pull request runs
  • ci: validate release guard wait settings
  • ci: wait for release check reruns
  • cloud: accept numeric whisper fixture transcript
  • cloud: address billing review findings
  • cloud: address rate authority review
  • cloud: align acceptance with native labels
  • cloud: align rehearsal rates with measured audit
  • cloud: benchmark qwen35 through native generate
  • cloud: bind acceptance hash to source artifact
  • cloud: bind exact lane cpu limits
  • cloud: bind generation grammar pool
  • cloud: bind structured schema semantics
  • cloud: bound outbox poison isolation
  • cloud: clarify idempotent deletion tests
  • cloud: clean up metered rehearsal credits
  • cloud: close console release preflight gaps
  • cloud: close launch integration review gaps
  • cloud: close launch review gaps
  • cloud: close operator key launch gaps
  • cloud: converge managed catalog exactly
  • cloud: cover multimodal rehearsal units
  • cloud: declare redact PII labels
  • cloud: derive migration log stream
  • cloud: drop stale metrics registry dependency
  • cloud: enforce billing activation invariants
  • cloud: finalize billing money hardening
  • cloud: govern legacy billing cutover
  • cloud: harden account lifecycle rehearsal
  • cloud: harden billing money flows
  • cloud: harden generation acceptance evidence
  • cloud: harden launch rate evidence
  • cloud: harden lifecycle rehearsal cleanup
  • cloud: harden signup alert routing
  • cloud: harden Stripe clawbacks and usage delivery
  • cloud: include registry compiler dependency
  • cloud: keep cleanup accounts suspended
  • cloud: keep live secret values out of state
  • cloud: keep proxy dry-run provenance current
  • cloud: keep qwen35 measurements pending
  • cloud: keep qwen35 performance pending
  • cloud: match Better Stack canonical alert fields
  • cloud: pin exact console release commits
  • cloud: preserve charged structured failures
  • cloud: preserve committed key recovery
  • cloud: preserve custom catalog evidence
  • cloud: preserve proxy token shape in dry-run
  • cloud: rebase catalog acceptance gate
  • cloud: reconcile exact metronome rating
  • cloud: reconcile native vision label bindings
  • cloud: refresh release lock artifact pins
  • cloud: reject duplicate billing bindings
  • cloud: reject empty key scopes
  • cloud: reject malformed catalog exports
  • cloud: reject unknown extraction recipes
  • cloud: reject unsupported lane providers
  • cloud: report gross Stripe funding
  • cloud: require console release domain
  • cloud: require ordered whisper transcript
  • cloud: restore Metronome billing authority
  • cloud: restore telemetry branch CI
  • cloud: retain charged acceptance failures
  • cloud: tighten billing terminal invariants
  • cloud: tolerate inaccessible gateway candidates
  • cloud: validate acceptance readiness reasons
  • cloud: validate edge origin before secret sync
  • cloud: validate lazy snapshot lanes
  • cloud: validate modal cloud placement
  • cloud: validate native label binding kinds
  • cloud: validate rate tariff shape
  • colpali: serialize model forwards — transformers output-recorder race
  • colpali: stop GPU memory accumulation on Vidore evals
  • control-plane: make wallet authority explicit
  • core: preserve unit count wire order
  • deploy: bound generation readiness process
  • deploy: harden generation readiness gate
  • deploy: honor local weight source precedence
  • deploy: prove exact generation revision readiness
  • deploy: resolve model staging weight sources
  • eval: add faithful Qwen3-VL reference
  • eval: bound multimodal reference scoring
  • eval: bound reranker candidates by corpus
  • eval: isolate VL reference dependencies
  • eval: keep gate plan aligned with dispatch
  • eval: preserve VL text coverage
  • eval: reject empty reranker qrels
  • evidence: bind immutable model identity
  • evidence: close worker identity contract gaps
  • evidence: harden worker identity boundaries
  • extract: close review validation gaps
  • extract: filter unconstrained relation endpoints
  • keda: keep gateway coordination under autoscaling
  • keda: simplify forward OTLP upgrade
  • keda: support forward-only OTLP upgrade
  • keda: verify lean forward upgrade
  • keep audio prep release pin in sync
  • loadtest: skip doomed bench-only scenarios when no release deployed
  • metering: reject invalid score counts
  • model-ops: align exact vision gate evidence
  • model-ops: gate parses the trigger’s ‘Triggered by #N’ issue reference
  • model-ops: void gate attempts that provably began no spend
  • model: restore SPLADE quality sequence cap
  • models: pin hf_revision for gte-Qwen2-7B remote-code load
  • models: restore SPLADE 512-token capacity
  • models: serve gte-Qwen2-7B via its own bidirectional forward
  • playbooks: v2-native task extraction + timeout contract for the sparse arm
  • quality-eval: identity-based signature dedup + branch hash (#2096 review)
  • quality-eval: make runner wake parallel, fail-loud, and demand-accurate
  • quality-eval: reset stale WakeRequestedAt against launch time
  • quality-eval: tolerate YAML date-typed fields in signature identity (#2096 review)
  • quality: tolerate missing gpu identity tooling
  • refresh managed release locks
  • release: refresh deployment locks during release regeneration
  • release: run deployment lock refresh in project env
  • reranking: address exact-head review
  • reranking: emit authoritative fake usage
  • reranking: emit fake score usage
  • reranking: reject blank compatibility inputs
  • server: empty bundle model_filter must advertise zero models, not all
  • server: remove unused fake adapter import
  • sie-bench: harden generation sharding evidence
  • sie-tools: keep session failures auth-scoped
  • sie-tools: preserve account lifecycle errors
  • smoke: align OTel collector contract checks
  • smoke: align OTLP metric assertions
  • smoke: restore kind smoke on main
  • splade: align telemetry and model docs
  • splade: harden profile and sparse defaults
  • tooling: fail fast before Modal dataplane runs
  • tooling: stop modal shells by app id
  • tools: preserve baked mise in modal shells
  • vision: address review edge cases
  • worker: propagate effective sequence caps
  • worker: release OOM tracebacks in the no-recovery branch too (CodeRabbit)
  • benchmarks: record SPLADE 512-token run
  • bench: record extraction L4 measurements
  • candle: compact sparse output in native precision
  • candle: fuse packed SPLADE pooling
  • candle: pack SPLADE BERT inference
  • vision: bulk-convert detection outputs
  • New capabilities: add GLiGuard and fix extraction catalog behavior; add GLiGuard and fix LightOnOCR and GLiREL extraction; add cloud benchmark campaigns; add native R3-Skill evaluation; add staging cloud smoke path; add two-client staging rehearsal
  • Reliability and operations: fence broker rollback recovery; harden cloud campaign evidence; harden cloud campaign readiness; harden live campaign preflight; harden managed campaign control plane
  • Performance: cache exact bf16 gelu values; fuse ModernBERT exact GeGLU; optimize ModernBERT CUDA hot path; vectorize bf16 gelu lookup gate
  • add GLiGuard and fix extraction catalog behavior
  • add GLiGuard and fix LightOnOCR and GLiREL extraction
  • bench: add cloud benchmark campaigns
  • bench: add native R3-Skill evaluation
  • bench: add staging cloud smoke path
  • bench: add two-client staging rehearsal
  • bench: complete managed cloud campaigns
  • bench: cover cloud task modalities
  • bench: inspect diagnostic campaign archives
  • bench: operationalize cloud measurement campaigns
  • bench: operationalize managed SIE Cloud measurement campaigns
  • ci: fake-stack-test job — Fake Engine regression suite + Cloud dispatcher loopback
  • ci: fake-stack-test job + five Fake Engine regression tests + Cloud loopback
  • cloud-gateway: serve job chunk results via signed capability URLs
  • cloud-gateway: serve offload job results via signed capability URLs
  • cloud: apex-domain cutover + console deploy model
  • cloud: control-plane audit-read + org-settings + residency-read endpoints (console Tier-2)
  • cloud: export worker-lane OTLP traces to the collector (lane-only, #1740)
  • cloud: gateway OTLP trace+metric export — clean end-to-end telemetry spine
  • cloud: govern Modal deployments with typed specs
  • cloud: per-org alert-threshold toggles (0006 alert_prefs) gating _fire_alerts
  • cloud: per-org console service keys via /internal/console-keys
  • cloud: per-org console service keys via POST /internal/console-keys
  • cloud: relocate the schedule runner in-VPC (EventBridge -> /internal/schedules/run-due)
  • cloud: resolve fresh keys on-demand to kill the ~5-10s gateway 401 window
  • cloud: surface the saved card’s brand/last4/exp (Stripe payment-method retrieve)
  • cloud: WP4 billing-depth control-plane endpoints (invoices, usage, alert toggles, payment method)
  • cloud: WP4 billing-depth endpoints — invoices, usage series, alert-config toggles
  • control-plane: make auto-recharge fire end-to-end
  • mac-smoke: additive fake phase 0 — weightless surface contract
  • model-ops: add stage-5 quality+perf onboarding gate
  • model-ops: add-model trigger — labeled issue starts the onboarding agent
  • model-ops: stage-5 quality+perf onboarding gate — measure vs target, bounded fix loop, evidence packet
  • models: add Tencent R3-Skill support
  • models: onboard intfloat/multilingual-e5-small (scaffolded)
  • quality-eval: #1810 experiment runner — parallel Modal sessions, GPU selection, budget backstop
  • quality-eval: autonomous experiment runner — parallel Modal sessions, GPU selection, budget backstop (Req14 v2 3/4)
  • quality-eval: end-to-end pipeline wiring + dedup/budget semantics (Req14 v2 4/4)
  • quality-eval: wire triage→planner→runner→autofix into a bounded pipeline
  • server: add deterministic weightless fake adapter family
  • server: failure injection for the fake adapter family
  • server: fake adapter family — deterministic weightless sie-fake models
  • server: synthetic memory-tracker mode — deterministic pressure/eviction
  • server: synthetic memory-tracker mode for deterministic eviction
  • support Iso ModernColBERT on Candle
  • tests: lost-result 120s characterization on the full queue topology
  • tests: lost-result characterization on the full queue topology
  • bench: accept canonical UTF-8 broker artifacts
  • bench: accept omitted optional request refs
  • bench: align broker contract boundary nesting
  • bench: align broker log retention
  • bench: align campaign runtime dependencies
  • bench: avoid imagebuilder template collision
  • bench: bind ami runtime kernel evidence
  • bench: bound R3 quality regression gate
  • bench: bound staging smoke below service knee
  • bench: close campaign review and ci gaps
  • bench: close campaign review edge cases
  • bench: expose safe broker failure diagnostics
  • bench: externalize broker deployment settings
  • bench: fence broker rollback recovery
  • bench: harden cloud campaign evidence
  • bench: harden cloud campaign readiness
  • bench: harden live campaign preflight
  • bench: harden managed campaign control plane
  • bench: make SSM dispatch at most once
  • bench: pass authorized request to orchestrator
  • bench: pin runnable worker manifests
  • bench: prepare campaign clients before start
  • bench: preserve broker terminal state
  • bench: preserve cloud campaign timing evidence
  • bench: preserve Fargate scratch permissions
  • bench: probe exact start control key
  • bench: publish worker image to owned ecr path
  • bench: reject malformed canonical artifacts
  • bench: relocate runtime caches to scratch
  • bench: remove stale quality import
  • bench: retain ami assertion diagnostics
  • bench: retry transient ECS role propagation
  • bench: retry transient R3 score timeouts
  • bench: roll task revisions without downtime
  • bench: secure ephemeral identity creation
  • bench: secure staging cloud validation
  • bench: skip absent profile deletion
  • bench: support repeated maintenance pauses
  • bench: trust scheduler group for worker expiry
  • bench: validate R3 evaluation inputs
  • candle: close ModernBERT audit gaps
  • ci: disable implicit cloud tool installs
  • ci: pin nested mise runtime
  • ci: verify cloud IaC toolchain
  • cloud-gateway: complete loopback class-fix via one canonical helper; align Python mirror + docs (#1906 final review)
  • cloud-gateway: IP-literal loopback check + recover stuck payload-volume reload (CodeRabbit #1906 r2)
  • cloud-gateway: reject non-ASCII signature hex without panicking
  • cloud-gateway: validate capability-URL origins, bound TTL, and fix object-store minting (CodeRabbit #1906)
  • cloud: address CodeRabbit — same-item batch id pairing + metrics protocol env + /metrics deny log
  • cloud: address CodeRabbit — serialize env-mutating metrics-protocol test with a module-local ENV_LOCK
  • cloud: address follow-up review
  • cloud: address Modal deployment review
  • cloud: alert on zero-metering invoice drift; reject ambiguous usage_events apportionment
  • cloud: align dev backend benchmark
  • cloud: align vendor billing with ledger debits
  • cloud: batch/offload jobs retry retryable per-item MODEL_LOADING instead of failing the chunk
  • cloud: bound on-miss key resolver’s single-flight map + harden lookup
  • cloud: CodeRabbit CP review — alert-configs PUT 404s, invoice-payload 502 guard, ty-ignore format
  • cloud: console SIE_GATEWAY_URL is the public edge, not the raw Modal URL
  • cloud: customer-scope the saved-card lookup — close cross-tenant card-PII leak
  • cloud: deploy otel-collector + merge OTLP endpoint before lane snapshot so lane tracing activates
  • cloud: derive outbox credits from the ledger debit and alert on reconciliation drift
  • cloud: diagnose stale-checkout CP destroy + bridge VERCEL_API_TOKEN
  • cloud: diagnose stale-checkout CP destroy + bridge VERCEL_API_TOKEN in cloud-provision
  • cloud: don’t persist raw billing_email in audit_log detail (PII)
  • cloud: final WP4 review — strict invoice currency, atomic alert toggle, UTC-naive parse, card guard
  • cloud: gateway OTLP export — blocking reqwest client for span batch processor + HTTP/TLS metric transport
  • cloud: gateway OTLP export — Tokio runtime for span batch processor + HTTP/TLS metric exporter
  • cloud: harden console-key lifecycle
  • cloud: harden whole-estate finalization
  • cloud: hide console keys from /me/usage + /me revoke (security review)
  • cloud: preserve billing alert and meter attribution
  • cloud: refresh recovery handle on batch chunk re-spawn (avoid orphaned container + 409 spam)
  • cloud: release the scheduler lock_conn on any pre-hand-off failure
  • cloud: stage the gen-lane model in model-stage and env-qualify the unstaged-weights hint
  • cloud: stop deploy from dirtying satellite uv.locks
  • cloud: stop deploy from dirtying satellite uv.locks (clean-tree guard)
  • cloud: WP4 review follow-ups — usage upper bound, invoice-fallback diagnosability, card exp contract
  • control-plane: hand auto-recharge sweep off the EventBridge request path
  • control-plane: harden auto-recharge collection
  • control-plane: make recharge retries lease-safe
  • fake-engine: address two-axis review findings across the stack
  • fake-engine: CI loopback interop + CodeRabbit review findings
  • gateway: attest deployed execution revision
  • gateway: include unanimous error_code in all_items_failed 500 logs
  • gateway: reject malformed queue-request bodies with a typed 400
  • harden classification quality verification
  • model-ops: apply CodeRabbit review findings on the stage-5 gate
  • model-ops: apply review findings on the friction PR
  • model-ops: deterministic escalation when the agent itself crashes
  • model-ops: filter vanilla floor material to fleet metric families
  • model-ops: harden add-model trigger per review
  • model-ops: raise directly in _read_json per code-quality finding
  • model-ops: read vanilla numbers from the verdict’s evidence field
  • model-ops: rehearsal follow-ups — agent guardrail, record valves, auto-review, log-tail noise
  • model-ops: rehearsal follow-ups — guardrails, record valves, log-tail noise
  • model-ops: scaffolder crash on present-but-null sibling fields
  • model-ops: seed-floor emits fleet metric families, not the raw mteb vector
  • model-ops: stage-5 gate read vanilla numbers from verdict evidence, not probe measurements
  • model-ops: tighten the agent-crash fallback per review
  • ocr: address review lifecycle findings
  • preserve parameterized quality gate targets
  • quality-eval: accurate backstop messages for CPU playbooks (review S1)
  • quality-eval: guard non-dict cluster signals in the pipeline (review)
  • quality-eval: harden experiment runner per PR #1883 review (traversal, stale-verdict, rollback, build-clock)
  • quality-eval: harden pipeline per PR #1891 review (report/verdict/workflow)
  • quality-eval: reap process groups in rollback to avoid zombies (re-review)
  • quality-eval: reject non-object verdict JSON in escalate_verdict (re-review)
  • quality: address R3 review feedback
  • quality: complete R3 faithful gate workflow
  • resolve file-backed adapter impacts
  • resolve onboarding bundle metadata
  • resolve query-qualified quality targets
  • reuse resolved uv in vanilla probe
  • server: allow dependency-free model selections
  • server: recommend bundle for full selection
  • server: support cloud catalogs in bundle checks
  • server: validate LightOnOCR bundle runtime
  • candle: cache exact bf16 gelu values
  • candle: fuse ModernBERT exact GeGLU
  • candle: optimize ModernBERT CUDA hot path
  • candle: vectorize bf16 gelu lookup gate
  • cloud: index usage_events_outbox (account_id, occurred_at) for the usage read (0007)
  • ocr: serve generative models with SGLang
  • ocr: serve generative OCR with SGLang continuous batching
  • New capabilities: add catalog-smoke step — per-model serve_i6pn load verification; add env-specific Modal deploy provisioning; align prod-us manifest to latest i6pn fast-path topology; commit generated price book so the provisioner locates it; deploy-both IaC hardening — env-derived Modal env + env-targeted B6 smoke; event-driven fast-lane promotion + credit ramp
  • Reliability and operations: drop the speculative mise install retry chain; harden deploy tooling — modal —json ANSI, async snapshot-commit race, Vercel sensitive-var upsert; harden lane_deploy driver — modal —json ANSI + async snapshot-commit race; retry server launch on startup death; scope the port sweep to our serve chain and harden the start guard
  • cloud: add catalog-smoke step — per-model serve_i6pn load verification
  • cloud: add env-specific Modal deploy provisioning
  • cloud: align prod-us manifest to latest i6pn fast-path topology
  • cloud: commit generated price book so the provisioner locates it
  • cloud: deploy-both IaC hardening — env-derived Modal env + env-targeted B6 smoke
  • cloud: event-driven fast-lane promotion + credit ramp
  • cloud: normalize the OSS terminal failed readiness state to MODEL_LOAD_FAILED on the Modal lane
  • cloud: productionize the i6pn realtime dispatch fast path end-to-end
  • cloud: readiness-gated i6pn dispatch — thread loaded_models to the wire, gate cold workers
  • model-ops: deterministic add-a-model scaffolder — verdict → YAML + matrix + target
  • model-ops: deterministic add-a-model scaffolder + onboarding runbook
  • model-ops: structured feasibility verdict — eval-model stage-2 gate
  • model-ops: structured feasibility verdict for add-a-model stage 2
  • quality-eval: add fix-proposal planner (tiered Fable-5) + experiment-plan schema
  • quality-eval: bounded auto-fix — agent PRs for floor re-baselines, human merge gate (Req 14 stage 4)
  • quality-eval: bounded auto-fix CLI for floor re-baseline
  • quality-eval: fix-proposal planner (tiered Fable-5) + experiment-plan schema (Req14 v2 2/4)
  • quality-eval: make the planner reason from signals + consume the LLM’s predicted evidence
  • quality-eval: onboarding smoke playbook — served-route bring-up on Modal
  • quality-eval: provenance-gated preflight + vanilla-lane floor template
  • quality-eval: provenance-gated preflight + vanilla-lane floor template (Req14 v2 1/4)
  • ci: drop the speculative mise install retry chain
  • ci: give the scheduled quality run its own concurrency group
  • ci: stop rolling the un-retried rustup ensure in mac-smoke-nightly
  • cloud: address catalog-smoke review — gate internal token, harden revoke + warm-up bound
  • cloud: address CodeRabbit round-2 review on #1782
  • cloud: address env deploy review comments
  • cloud: address PR review comments
  • cloud: bound wedged model loads + typed dead-letter on the lazy lane
  • cloud: bound/adapt slab credit holds so a funded org stops self-wedging at 402
  • cloud: catalog-smoke presents SIE_API_KEY when the gateway enforces managed keys
  • cloud: don’t send type on Vercel env update (sensitive-var 400)
  • cloud: drop candle-only bge-large-en from the default L4 lane’s models_served
  • cloud: gateway must not 401 valid keys during cold-replica key-snapshot warmup
  • cloud: gen-lane warm floor + gateway-config secret merge in deploy-edge
  • cloud: harden deploy tooling — modal —json ANSI, async snapshot-commit race, Vercel sensitive-var upsert
  • cloud: harden lane_deploy driver — modal —json ANSI + async snapshot-commit race
  • cloud: heartbeat TTL for i6pn discovery keys (corpse expiry on hard stop)
  • cloud: narrow B6 local gateway bundle to the deployed models_served (#1771 hash parity)
  • cloud: pg-connector smoke builds ModalLaneExecutor from SIE_LANE_APP_MAP
  • cloud: pg-connector smoke builds ModalLaneExecutor from SIE_LANE_APP_MAP (#1782 missed caller)
  • cloud: scope _modal_cli TERM=dumb to captured calls only (CodeRabbit #1807)
  • cloud: skip B6 offload-batch in external-CP mode (unshareable payload store)
  • cloud: stage refs/main alias so offline lanes resolve pinned-SHA models
  • cloud: stop provisioning unbacked Stripe metered prices
  • cloud: stop provisioning unbacked Stripe metered prices (prepaid-credits + Metronome rating is the model)
  • cloud: tighten edge deploy conflict resolution
  • cloud: Vercel env update sends value only (sensitive-var target 400)
  • cloud: warm managed key before catalog-smoke sweep to beat the key-propagation race
  • cloud: wire the default gen lane into the gateway census + app map
  • gateway: bump yanked spin 0.10.0 to 0.10.1 in Cargo.lock
  • mac-smoke nightly score-path contract + unknown-model 500s
  • mac-smoke: call /v1/score with the slash-form model id
  • mac-smoke: kill stale server by port between phases
  • mac-smoke: retry server launch on startup death
  • mac-smoke: scope the port sweep to our serve chain and harden the start guard
  • model-ops: harden feasibility validator per review
  • model-ops: harden scaffolder per review
  • poc: pass explicit github_token to claude-code-action
  • quality-eval: charset-gate vanilla_vector metric keys (review C1)
  • quality-eval: classify malformed encode responses as REFUTED, not a crash
  • quality-eval: controlled exit on planner output-write failure (CodeRabbit)
  • quality-eval: harden autofix per CodeRabbit review round 2
  • quality-eval: harden autofix preflight against output injection (review round 1)
  • quality-eval: require string cell fields before charset validation
  • server: colbert-small multivector encode returns one vector-set per input, worker-item order
  • server: Florence-2 processor compat with transformers 4.57
  • server: guard cross_encoder predict tokenization against the concurrent-metering race (class-fix for #1800)
  • server: load tokenizer by pinned revision so offline chat-template rendering doesn’t need refs/main
  • server: return 404 instead of 500 for unknown models on inference endpoints
  • server: terminal Failed readiness + Florence-2 processor compat + pin launch-set revisions
  • server: terminal Failed readiness state so the sidecar dead-letters permanently-failed model loads
  • New capabilities: add flat→per-account Files-store migration; co-located AWS us-east-1 reverse-proxy edge for api.superlinked.com; consolidate data-plane deploy + manifest-drive proxy-auth/i6pn; namespace Files-store objects by account; opt-in Modal proxy-auth layer for config/lane/OTLP ingress; parameterize Modal environment + prod-us tf root by SIE env (§B2)
  • Reliability and operations: fail fast on terraform plan errors before the destroy-check; bake SIE_MODAL_PROXY_AUTH into image env so the proxy-auth knob doesn’t diverge local↔remote; make proxy_auth_env authoritative in both directions; pin npm for TypeScript release publishing; aws-edge S3 backend uses native state locking (use_lockfile)
  • cloud: add flat→per-account Files-store migration
  • cloud: co-located AWS us-east-1 reverse-proxy edge for api.superlinked.com
  • cloud: consolidate data-plane deploy + manifest-drive proxy-auth/i6pn
  • cloud: namespace Files-store objects by account
  • cloud: opt-in Modal proxy-auth layer for config/lane/OTLP ingress
  • cloud: parameterize Modal environment + prod-us tf root by SIE env (§B2)
  • cloud: per-account Files-store namespacing + migration
  • cloud: prod bring-up terraform prerequisites
  • cloud: prod bring-up terraform prerequisites — prod-us auth-token outputs, aws-edge prod state
  • cloud: prod enablement — aws-edge provisioner step, prod-us root, Modal-env parameterization
  • cloud: prod-us terraform root — env-isolated control plane, separate state bucket
  • cloud: reconcile OTLP transport + proxy-auth sender headers
  • cloud: stand up the AWS edge via cloud-provision
  • dispatcher: full IP-pin for object-store + postgres verify-full egress
  • dispatcher: org-ownership assertion on upload:// ends — defense in depth behind the gateway check
  • dispatcher: resolve-then-pin connector egress — SSRF TOCTOU backstop
  • ci: fail fast on terraform plan errors before the destroy-check
  • ci: pin npm for TypeScript release publishing
  • cloud: address CodeRabbit review on the prod-enablement PR
  • cloud: aws-edge S3 backend uses native state locking (use_lockfile)
  • cloud: bake SIE_MODAL_PROXY_AUTH into image env so the proxy-auth knob doesn’t diverge local↔remote
  • cloud: collision-aware Files migration move
  • cloud: correct console callback route + guard the CP terraform apply against silent destroys
  • cloud: correct scheduler/slo-tuner DSN-secret comments + enforce non-ingress secret refs in deploy-posture audit
  • cloud: default Metronome credit type to fiat USD-cents so low-balance alerts and grants provision
  • cloud: gateway payload store must be a local path, not the s3 payload_store_url
  • cloud: gateway payload store must be a local path, not the s3 payload_store_url (#1743 follow-through)
  • cloud: gateway ships the cloud-storage backend; provisioner wires SIE_CONFIG_SERVICE_URL
  • cloud: guard config-URL resolver against a malformed manifest
  • cloud: harden Phase-3 CI + sibling-sweep per adversarial review (Refs #1757 #1758)
  • cloud: make proxy_auth_env authoritative in both directions
  • cloud: parameterize the deploy-env slug in managed Modal app names
  • cloud: parameterize the OTLP collector app name with the deploy-env slug
  • cloud: provision §7.6 prepaid credit-pack price + STRIPE_PRICE_ID
  • cloud: redact all secret-bearing Config fields in Debug
  • cloud: redact ModalProxyToken Debug + cover /gen handshake forwarding
  • cloud: resolve gateway app name from gateway_app.APP_NAME, not ctx.env
  • dispatcher: address CodeRabbit review on the #1742 egress pins
  • dispatcher: harden connector egress pin against CodeRabbit findings
  • stabilize tilt e2e regressions
  • use pinned cargo-zigbuild in tilt builds
  • New capabilities: add GTE RoPE embedding kernel; CLI WorkOS device-flow auth — retire the shared-token shortcut for customer commands; SIE Cloud managed service — AWS control plane, Modal data plane, provisioning + billing kernel; OSS amendments from the SIE Cloud build — engine, SDKs, gateway/sidecar seams; add deterministic regression-ledger detector for scheduled quality runs; add floor-archaeology playbook (CPU-only)
  • Reliability and operations: retry model-dependent calls on MODEL_LOADING; harden regression-ledger live mode and audit input validation; harden multi-gpu worker routing; compile fused gated activation op; import CUDA pointer trait
  • candle: add GTE RoPE embedding kernel
  • cloud: CLI WorkOS device-flow auth — retire the shared-token shortcut for customer commands
  • cloud: SIE Cloud managed service — AWS control plane, Modal data plane, provisioning + billing kernel
  • oss: OSS amendments from the SIE Cloud build — engine, SDKs, gateway/sidecar seams
  • quality-eval: add deterministic regression-ledger detector for scheduled quality runs
  • quality-eval: add floor-archaeology playbook (CPU-only)
  • quality-eval: add Modal parity-probe playbook (diagnostic-only)
  • quality-eval: add official-recipe, fp32-ab, config-sweep, determinism playbooks
  • quality-eval: add report-only triage classifier + workflow for scheduled quality runs
  • quality-eval: add validation-playbook harness + hypothesis/verdict contract
  • quality-eval: report-only triage classifier + gated workflow for scheduled quality runs
  • quality-eval: seed triage signature DB and short-circuit known clusters
  • quality-eval: spread runner fleet across capacity pools with always-on floors
  • quality-eval: validation playbooks — adversarial repro harness (Req 14 Proj 1, stage 3)
  • server: support multi-gpu worker placement
  • worker: route by child queue pressure
  • worker: route multi-gpu sidecar children
  • worker: support multi-gpu sidecar children
  • candle: compile fused gated activation op
  • candle: import CUDA pointer trait
  • mac-smoke: retry model-dependent calls on MODEL_LOADING
  • quality-eval: address CodeRabbit + code-quality review comments
  • quality-eval: address CodeRabbit review findings
  • quality-eval: address triage review findings
  • quality-eval: detect EBS cache volume via lsblk model string
  • quality-eval: enforce budget cap on all eval playbooks + budget-stop artifacts
  • quality-eval: fp32 ST arm via throwaway -st model, not a profile
  • quality-eval: harden playbook verdict logic (review round 1)
  • quality-eval: harden regression-ledger live mode and audit input validation
  • quality-eval: harden triage input handling per CodeRabbit review
  • quality-eval: keep register-runners.sh bash-3.2 compatible
  • quality-eval: tolerate stop/start race in wake instance-stopped wait
  • quality-triage: share data-bearing baseline discovery with triage live mode
  • quality-triage: skip data-less scheduled runs in ledger baseline auto-discovery
  • req6: align candle chart after rebase
  • req6: harden multi-gpu worker review fixes
  • req6: harden multi-gpu worker routing
  • req6: resolve main rebase fallout
  • req6: wire pressure-aware multi-gpu routing
  • server: keep ipc ping alive on gpu health errors
  • sidecar: bound queue admission through scheduler completion
  • sidecar: refresh ready slots on cancel fanout failure
  • sidecar: release cancelled scheduler pressure
  • sidecar: wire local ingest through adapter pool
  • New capabilities: opt-in template_as_prompt for ST prompt-length-aware templates; use_model_encode — delegate to checkpoint-native encode(); add Candle OOM pressure recovery; add Candle residency eviction; per-pool gpu_driver option + confidential H100 tracing example
  • Reliability and operations: harden splade flash gate + add flash-vs-native parity test; splade-v3 floors were cosine-epoch artifacts — re-baseline to dot self-measures + harden flash gate; restore executable bit on mac-smoke mise task; faithful ModernBERT forward + hardened 1_Dense loading; restore e5 query/passage templates on the multilingual-e5-large sentence_transformer profile
  • adapters: opt-in template_as_prompt for ST prompt-length-aware templates
  • adapters: use_model_encode — delegate to checkpoint-native encode()
  • add Candle OOM pressure recovery
  • add Candle residency eviction
  • terraform/azure: per-pool gpu_driver option + confidential H100 tracing example
  • adapters: harden splade flash gate + add flash-vs-native parity test
  • adapters: honor caller instruction in use_model_encode without a template
  • address candle catalog review comments
  • address Candle residency review comments
  • align Candle idle eviction lifecycle
  • align candle workers with catalog lane config
  • bench: cap docling pubtabnet-html at 2000 pages to fit the job budget
  • bench: splade-v3 floors were cosine-epoch artifacts — re-baseline to dot self-measures + harden flash gate
  • bundles: pin flash-attn-4 prerelease so sglang 0.5.10.post1 resolves
  • ci: restore executable bit on mac-smoke mise task
  • colbert_modernbert_flash: apply trained ColBERT Dense projection (+ guard, tokenizer, re-baseline)
  • colbert_modernbert_flash: faithful ModernBERT forward + hardened 1_Dense loading
  • colbert: GTE muvera also at simhash 6 - CMedQA/MMarco evals OOM at dim 81920
  • colbert: mxbai muvera at simhash 6 (FDE dim 20480) - runner dies at 81920
  • colbert: retune GTE/mxbai muvera for the Dense-projected representation
  • colbert: serve pylate Dense chains and align fallback with flash
  • deploy: cap tester l4 lanes
  • deploy: zero and cap tester L4 lanes
  • deploy: zero AWS worker warm floor
  • deploy: zero tester worker warm lanes
  • gateway: fan out gpu-agnostic demand across pool profiles for scale-from-zero
  • gateway: normalize explicit GPU demand labels
  • helm: wake a deterministic owner lane on gpu-agnostic demand for multi-profile bundles
  • models: NV-Embed-v2 — serve via checkpoint-native encode() (latent-attention recipe), NFCorpus 0.3715→0.4479
  • models: restore e5 query/passage templates on the multilingual-e5-large sentence_transformer profile
  • models: restore e5 templates on multilingual-e5-large ST profile
  • models: serve NV-Embed-v2 via its native latent-attention ST path
  • preserve Candle idle evictor failure logs
  • quality-eval: fail fast when sie-server dies during readiness wait
  • quality-eval: unbreak 7B sglang lanes (flash-attn-4 resolution) + fail-fast readiness + docling job budget
  • restore kind smoke candle routing
  • New capabilities: add Candle profile routing and catalog models; add candle runtime diagnostics; add candle xlm-roberta flash attention; add dev AMI bake script; add dev ec2 cleanup workflow; add dev EC2 launch script
  • Reliability and operations: harden native worker routing; harden native worker runtime; restore sidecar model readiness; delete Python NATS pull loop from worker hot path; harden SIE dev AMI workflow
  • Performance: add GEMM diagnostics switch; fuse xlm-r qkv projection; optimize bge-m3 xlm-r runtime path; enable candle reduced precision gemms
  • add Candle profile routing and catalog models
  • add candle runtime diagnostics
  • add candle xlm-roberta flash attention
  • add dev AMI bake script
  • add dev ec2 cleanup workflow
  • add dev EC2 launch script
  • add SIE API key rotation tooling
  • candle: add native adapter runtime
  • dev: add personal SIE CUDA/Rust AMI workflow
  • document ssm ssh dev launcher
  • gateway: count dropped stale NAKs from abandoned attempts
  • helm: add opt-in OTLP tracing wiring + bundled collector to sie-cluster
  • helm: add optional bundled Tempo tracing backend
  • helm: add SIE tracing dashboard and tempo datasource uid
  • helm: opt-in OTLP tracing wiring + bundled collector in sie-cluster
  • helm: optional bundled Tempo tracing backend + Grafana datasource for sie-cluster
  • improve candle replica concurrency
  • improve dev AMI cache scalability
  • observability: bound tracing shutdown + trim env inputs across all four runtimes
  • observability: bound tracing shutdown and trim env inputs across all four runtimes
  • observability: gate Rust tracing on SIE_TRACING_ENABLED
  • observability: unify tracing enable switch across gateway, sidecar, and worker
  • quality-eval: gate score-on-retrieval candidate cells once floor committed
  • reattach dev cache volume
  • server: land Apple Silicon as a primary device onto main
  • server: land Apple Silicon as a primary device onto main (re-land of #1455)
  • sie_gateway: emit gateway.publish span on non-streaming queue-publish path
  • sie_server_rust: add OTLP distributed tracing
  • sie_server_rust: add OTLP distributed tracing to the Rust native worker
  • terraform: dedicated single-AZ node group for stateful observability
  • tilt: enable tracing with bundled Tempo
  • tracing: emit sidecar.dispatch span from sie_server_sidecar
  • tracing: inject queue trace context for non-generation endpoints
  • tracing: instrument sie_server_sidecar with OTLP (sidecar.dispatch span)
  • tracing: propagate trace context through the non-streaming worker loop
  • tune candle forward concurrency
  • tune dev EC2 launch defaults
  • wire dev AMI cache disk helpers
  • #1340: stella per-task instruction + re-baseline stale dense/colbert floors (blast radius was #1432)
  • adapters: guard empty-document MaxSim batches; strengthen single-doc test
  • address API key review feedback
  • address API key rollback review
  • address rust worker smoke review comments
  • address sidecar PR review feedback
  • align rust worker kind smoke wiring
  • align rust worker warm smoke behavior
  • align worker pool config hashes
  • back up API key secret before apply
  • candle: align cublaslt with candle cuda stream
  • candle: align readiness with loaded model state
  • candle: copy vendored cuda dependency in image
  • candle: harden native worker routing
  • candle: harden native worker runtime
  • candle: import cublaslt cuda extension
  • candle: keep catalog additions candle-only
  • candle: publish loaded models in health
  • candle: repair layer norm cuda lifetimes
  • candle: restore sidecar model readiness
  • candle: run bge-m3 profile in bf16
  • candle: vendor cublaslt for cuda build
  • candle: vendor layer norm cuda dependency
  • clean stale Python queue ownership drift
  • complete empty preformed worker requests
  • compress dev ami user data
  • config: defer cancellation after config commit starts
  • config: handle idempotency in-flight races
  • config: shield idempotency failure cleanup
  • copy rust worker vendor deps before docker cook
  • dashboard: harden run_status backfill per review feedback
  • dashboard: keep cancelled nightlies off the Quality on main widget
  • delete Python NATS pull loop from worker hot path
  • dev: address dev AMI review hardening
  • dev: harden SIE dev AMI workflow
  • dev: persist Codex state on dev cache disk
  • docs: drop otel.name from span-attribute table; tighten tracing dashboard test
  • export profile bundle targets
  • gateway: append forwarded embeddings headers to preserve multi-value tracestate
  • gateway: cancel abandoned direct batch fallback work
  • gateway: don’t report a succeeded SSE generation as inter_chunk_timeout on consumer lag
  • gateway: drop stale NAKs from abandoned attempts in handle_nak
  • gateway: forward traceparent/tracestate through /v1/embeddings rewrite
  • gateway: harden batch direct fallback cancellation
  • gateway: ignore pre-latch stale NAKs after republish
  • gateway: ignore quick-xml <0.41 RustSec advisories in cargo-deny
  • gateway: keep signalling SSE incompleteness on consumer lag; skip redundant cancel
  • gateway: keep signalling SSE incompleteness on lag; only skip cancel when done
  • gateway: record demand on backpressure so KEDA scales the lane
  • gateway: route engine-pin errors through endpoint_error_response
  • gateway: separate queue timeouts from model loading
  • gateway: use real RUSTSEC id for second quick-xml advisory
  • harden candle bge-m3 runtime bring-up
  • harden dev EC2 launcher dry-run
  • harden dev profile home fallback
  • harden SIE API key tooling
  • helm: make bundled collector downstream OTLP TLS configurable
  • helm: set OTEL_EXPORTER_OTLP_PROTOCOL=grpc on traced pods
  • helm: suppress bundled collector when explicit tracing endpoint set
  • improve candle embedding runtime throughput
  • install gateway pool config hashes
  • install gpg agent for dev AMI bake
  • keep dev AMI user data under limit
  • log candle batch failure context
  • models: load stock tokenizer for gte-Qwen2-7B-instruct
  • models: restore per-task MTEB instruction for stella_en_v5
  • models: serve embeddinggemma-300m via SentenceTransformerDenseAdapter
  • models: serve gte-Qwen2-1.5B via faithful sentence_transformer adapter
  • models: serve gte-Qwen2-1.5B via sentence_transformer adapter
  • muvera: center tokens before SimHash to fix answerai @muvera Blackwell collapse
  • muvera: drop same-dim count-sketch + retune colbert muvera configs
  • muvera: drop same-dim count-sketch, retune colbert configs + floors
  • persist dev cache volume
  • persist dev kubeconfig
  • playground: keyboard-accessible tabs + associated form labels
  • polish rollback backup examples
  • preserve candle compute cap diagnostics
  • preserve control-plane bundle hash in sidecar
  • quiet dev AMI tool installs
  • remove candle replica fallback path
  • remove Python worker pull-loop path
  • repair candle ci routing checks
  • repair candle rebase fallout
  • repeat dev AMI bake markers
  • require explicit API key rollback path
  • resolve latest dev AMI base
  • respect hf cache env in rust worker
  • restore default candle cuda stream
  • route kind rust worker build through docker task
  • sdk: raise actionable error on encode result-count desync
  • select rustls provider at binary startup
  • select rustls provider for rust services
  • server: /v1/embeddings maps OOM to 503 RESOURCE_EXHAUSTED, not 500
  • server: address CI + review comments on the Apple Silicon re-land
  • server: apply profile runtime options + append EOS on cluster encode path
  • server: apply profile runtime options and append EOS on cluster encode path
  • server: bound continuous-batch drain so a busy LoRA can’t starve siblings
  • server: bound the continuous-batching drain so a busy LoRA can’t starve siblings
  • server: cap LoRA drain to snapshot backlog so post-snapshot arrivals don’t overshoot
  • server: make SGLang load eviction headroom-aware
  • server: move load headroom behind adapter contract
  • server: serialize hot reload registry mutations
  • server: skip EOS on empty input; cover non-default profile on queue path
  • server: stop mislabelling uint8 embeddings as “binary” on the queue path
  • server: stop mislabelling uint8 embeddings as binary on the queue path
  • server: unregister from MemoryManager even if adapter.unload() raises
  • server: unregister model from MemoryManager even if adapter.unload() raises
  • set home during dev AMI bake
  • sidecar: bound scheduler bulk enqueue waves
  • sidecar: keep only config hash wiring
  • sidecar: start batch cancel before direct pulls
  • sie_server_rust: only log OTLP init success when exporter built
  • stabilize rust worker kind smoke
  • terraform: address CodeRabbit review on obs node group
  • terraform: auto-derive observability node group AMI from instance arch
  • terraform: make observability node group ami_type/disk_type configurable
  • tighten rust worker smoke diagnostics
  • tilt: use single string for candle-helm-smoke cmd (Starlark has no implicit concat)
  • tilt: wait for complete trace before asserting span chain
  • tolerate missing zstd devel package
  • tracing: assert items/metadata alignment before zip in scheduler batch
  • tracing: attach inbound trace context for non-generation publishes
  • tracing: parent sidecar.dispatch on first valid inbound context
  • verify dev AMI bake via ssm
  • candle: add GEMM diagnostics switch
  • candle: fuse xlm-r qkv projection
  • candle: optimize bge-m3 xlm-r runtime path
  • enable candle reduced precision gemms
  • fuse candle xlm-roberta qkv projection
  • gateway: build the HRW ring from the by_bundle index, not a whole-cluster scan
  • improve candle worker concurrency and build cache
  • restore candle xlm-roberta qkv layout
  • server: vectorize bge_m3_flash CLS gather to drop N per-encode GPU syncs
  • sidecar: bulk enqueue scheduler batches
  • restore gateway sidecar rustls defaults
  • Reliability and operations: unblock ColBERT score-on-retrieval — score_pairs 500 + muvera dense-encode + gate co-serve; include sie-bench in image context
  • #1430: unblock ColBERT score-on-retrieval — score_pairs 500 + muvera dense-encode + gate co-serve
  • docker: include sie-bench in image context
  • New capabilities: publish sie-bench image to GHCR on release; per-tool GPU routing and env-overridable GLiNER models
  • Reliability and operations: extend tester structured generation timeouts; retry failed payload cleanup deletes; align sie-test qwen27 lane routing; build sie-bench image CPU-only with headless opencv; tolerate cold-start provisioning errors in scale-from-zero loadtests
  • ci: publish sie-bench image to GHCR on release
  • sie-mcp: per-tool GPU routing and env-overridable GLiNER models
  • align sie-test qwen27 lane routing
  • bench: address CodeRabbit review (non-root image, clearer comment, checkout creds)
  • bench: build sie-bench image CPU-only with headless opencv
  • bench: tolerate cold-start provisioning errors in scale-from-zero loadtests
  • extend tester structured generation timeouts
  • gateway: dedupe offloaded payload cleanup keys
  • gateway: retry failed payload cleanup deletes
  • gateway: track exact offloaded payload keys
  • loadtest: tolerate cold-start provisioning errors + publish sie-bench to GHCR
  • persist sie-test qwen warm lanes
  • scale sie-test spot lanes to zero
  • server: drain removed profile variants during config apply
  • server: preflight profile variant config updates
  • server: preserve profile variants during config reconcile
  • server: serialize async config registry updates
  • server: tighten profile variant reconciliation
  • sie-mcp: chunk GLiNER extract/redact for large docs; flag the other 4B
  • New capabilities: enable gateway-API model pinning on sie-test; add per-profile quality floors for bge-m3 sparse/multivector
  • Reliability and operations: exclude per-profile target.<profile>.json floors from result loaders; run docker task in project env; skip payload-store cleanup DELETEs for inline requests; guard label-window overflow + defer LEDGAR for v1.0 uni-encoders; gate subset-parameterized tasks; keep skipping candidate-gen dirs
  • Performance: release MaxSim per-batch intermediates to cap peak memory
  • deploy: enable gateway-API model pinning on sie-test
  • quality-eval: add per-profile quality floors for bge-m3 sparse/multivector
  • bench: exclude per-profile target.<profile>.json floors from result loaders
  • ci: run docker task in project env
  • deploy: address CodeRabbit review on sie-test pinning
  • gateway: record offload before the store write (PR review)
  • gateway: skip payload-store cleanup DELETEs for inline requests
  • gliclass: guard label-window overflow + defer LEDGAR for v1.0 uni-encoders
  • quality-eval: gate subset-parameterized tasks; keep skipping candidate-gen dirs
  • sie_bench: batch-local query padding in MaxSim to fix #1461 OOM
  • sie_bench: release MaxSim per-batch intermediates to cap peak memory
  • New capabilities: floor Florence-2-base on COCO detection (AP 0.2785); floor Florence-2-base on olmOCR-bench (accuracy 0.0824); fail fast when a bundle is placed on the wrong platform; add Gemma 4 generation models (E2B, E4B, 26B-A4B); first-class cuda13 platform for the gemma bundle; first-class cuda13 platform for the gemma bundle (build/CI/release/deploy)
  • Reliability and operations: raise uv HTTP timeout for the gemma cu130 wheel downloads; harden logical pool queue routing; harden worker-direct batch fallback; return retryable model loading for generation; gate mteb task instruction to instruction-following models
  • bench: floor Florence-2-base on COCO detection (AP 0.2785)
  • bench: floor Florence-2-base on olmOCR-bench (accuracy 0.0824)
  • deploy: fail fast when a bundle is placed on the wrong platform
  • server: add Gemma 4 generation models (E2B, E4B, 26B-A4B)
  • server: first-class cuda13 platform for the gemma bundle
  • server: first-class cuda13 platform for the gemma bundle (build/CI/release/deploy)
  • server: keep pinned models loaded and exclude them from LRU eviction
  • support logical pools backed by queue pools
  • terraform: add marci-dev Azure A10 single-node smoke-test example
  • terraform: Azure marci-dev A10 smoke-test example + node-subnet NSG LB fix
  • align retryable generation review feedback
  • bench: gate mteb task instruction to instruction-following models
  • bench: gate MTEB task instructions to instruction-following models
  • bench: resolve profile output_types from adapter_options.runtime for bge-m3 sparse/multivector evals
  • ci: raise uv HTTP timeout for the gemma cu130 wheel downloads
  • ci: run the gemma cuda13 build smoke + cover it in release verify/stamp/warm
  • disable payload store in kind smoke
  • handle sidecar progress ack failures
  • harden logical pool queue routing
  • harden worker-direct batch fallback
  • loadtest: disable payload store in nightly deploy
  • merge runtime pool capacity changes
  • preserve queue delivery budget
  • return retryable model loading for generation
  • server: correct stale cuda13/gemma platform references
  • server: make docker bake/verify platform-aware (cuda13/gemma in mixed builds)
  • server: make transformers version bundle-controlled
  • terraform: allow public LoadBalancer/ingress inbound on AKS node subnet NSG
  • terraform: clear internal mise references from public Azure module
  • terraform: clear internal mise/tooling references from public AWS and GCP modules
  • terraform: drop internal references from public Azure module README and variables
  • New capabilities: floor measure-first dark models, retire donut-rvlcdip; provision and enable the payload store by default; fix <OD> bbox conversion, re-task base-ft/large to COCO detection
  • Reliability and operations: align gliner label inputs; align gliner v2.5 conll protocol; improve generation failure observability; coalesce whitespace-only instruction to default; pool post-RMSNorm last_hidden_state
  • bench: floor measure-first dark models, retire donut-rvlcdip
  • deploy: provision and enable the payload store by default
  • florence2: fix <OD> bbox conversion, re-task base-ft/large to COCO detection
  • bench: align gliner label inputs
  • bench: align gliner v2.5 conll protocol
  • deploy: address payload-store review feedback
  • improve generation failure observability
  • qwen3_vl_embedding: coalesce whitespace-only instruction to default
  • qwen3_vl_embedding: pool post-RMSNorm last_hidden_state
  • sie_bench: coerce ClassLabel int names to str in classification loader
  • sie_bench: repoint extract-kie FUNSD to live nielsr/funsd dataset
  • sie-cluster: make the self-host deploy skill robust to its own runtime
  • sie-server: make transformers5 bundle reachable via pip
  • sync generation visibility api contract
  • New capabilities: add per-pool pinned-model set to the pool config API; per-pool pinned-model set in the pool config API; support profile-qualified ids in per-pool pinned-model set; enable S3 payload store for >1MB work items
  • Reliability and operations: bound pinned-model metric labels and add handler validation tests; scope worker preload models by lane; alias PEFT LoRA adapter names; expose model_cache_bucket_url as a root output
  • gateway: add per-pool pinned-model set to the pool config API
  • gateway: per-pool pinned-model set in the pool config API
  • gateway: support profile-qualified ids in per-pool pinned-model set
  • tester-cluster: enable S3 payload store for >1MB work items
  • gateway: bound pinned-model metric labels and add handler validation tests
  • scope worker preload models by lane
  • server: alias PEFT LoRA adapter names
  • tester-cluster: expose model_cache_bucket_url as a root output
  • Reliability and operations: harden release artifact workflows; keep score traffic from idle-evicting rerankers; serialize worker use with unload state; share score media batch cost
  • ci: harden release artifact workflows
  • server: keep score traffic from idle-evicting rerankers
  • server: serialize worker use with unload state
  • server: share score media batch cost
  • New capabilities: answer_questions transient-QA tool (Req 12 #1309)
  • Reliability and operations: guard model ready timeout against liveness budget; wire model ready timeout through helm; add Qwen3.6 27B 32K RTX serving profile; route grammar requests to a non-speculative profile (NEXTN bypasses Outlines FSM); make worker config reconciliation no-op safe
  • Performance: extract shared vision patch-embed rebind helper; batch generate() across pages to close #601 L4 throughput gap
  • mcp: answer_questions transient-QA tool (Req 12 #1309)
  • add Qwen3.6 27B 32K RTX serving profile
  • address qwen 27b review feedback
  • generate: route grammar requests to a non-speculative profile (NEXTN bypasses Outlines FSM)
  • guard model ready timeout against liveness budget
  • make worker config reconciliation no-op safe
  • route qwen 27b fp8 cold starts
  • wire model ready timeout through helm
  • adapters: extract shared vision patch-embed rebind helper
  • lighton_ocr: batch generate() across pages to close #601 L4 throughput gap
  • Reliability and operations: align pool-scoped bundle hashes; avoid sticky missing bundle hashes; clarify missing profile inheritance; fail closed on missing bundle metadata; stabilize keda all-marker e2e
  • config: align pool-scoped bundle hashes
  • config: avoid sticky missing bundle hashes
  • config: clarify missing profile inheritance
  • config: fail closed on missing bundle metadata
  • tilt: stabilize keda all-marker e2e
  • New capabilities: demand-side token-reduction benchmark (Req 12, #1311); add describe_image tool (caption + zero-shot tags); add describe_image tool (caption + zero-shot tags) — Req 12 #1310; cap describe_image payload size before cluster calls; claude.ai connector surface — OAuth bridge + skill ZIP (Req 12 #1312); sie_mcp edge with docs_to_markdown tool (Req 12 #1306)
  • Reliability and operations: harden replace snapshot IPC; return retryable OpenAI provisioning errors; bind OAuth authorization codes to client_id; doctor classifies probe read-timeouts as cold, not unreachable; align GPU memory pressure defaults
  • Performance: skip tag embedding when top_k <= 0
  • bench: demand-side token-reduction benchmark (Req 12, #1311)
  • mcp: add describe_image tool (caption + zero-shot tags)
  • mcp: add describe_image tool (caption + zero-shot tags) — Req 12 #1310
  • mcp: cap describe_image payload size before cluster calls
  • mcp: claude.ai connector surface — OAuth bridge + skill ZIP (Req 12 #1312)
  • mcp: sie_mcp edge with docs_to_markdown tool (Req 12 #1306)
  • mcp: structured extraction + structured generation tools (Req 12 #1308)
  • mcp: wire measured token-reduction figures into savings metadata
  • tools: add sie doctor — per-capability cluster diagnostics
  • tools: Florence-2 fallback for image OCR
  • tools: sie_tools — Claude Code context-offload client for managed clusters
  • align GPU memory pressure defaults
  • config: detect bundle config hash drift
  • config: fingerprint model pool ownership
  • config: harden replace snapshot IPC
  • config: replace drifted export snapshots
  • deps: bump sidecar prometheus for protobuf advisory
  • gateway: address provisioning review feedback
  • gateway: align provisioning contract docs
  • gateway: decode native media JSON bytes
  • gateway: dereference structured output schema refs
  • gateway: make provisioning non-2xx universally
  • gateway: preserve ref sibling schema semantics
  • gateway: return retryable OpenAI provisioning errors
  • mcp: address review feedback on structured tools
  • mcp: bind OAuth authorization codes to client_id
  • mcp: blank-env fallback for model ids; honor SIE_MCP_IMAGE_TOP_K=0
  • mcp: deep-copy committed token-reduction figures in build_metadata
  • mcp: validate embedding shapes in _top_k_tags
  • sdk: normalize score image payloads for wire transport
  • sdk: normalize score images for wire transport
  • server: guard readiness for removed configs
  • server: honor pool-aware model configs
  • server: render qwen3 vl reranker document images in user prompt
  • server: render Qwen3-VL reranker document images in user prompt
  • sie-cluster: add spot toleration to AKS worker pool
  • tools: address doctor review feedback
  • tools: doctor classifies probe read-timeouts as cold, not unreachable
  • worker: keep SGLang loads off event loop
  • mcp: skip tag embedding when top_k <= 0
  • New capabilities: add grant_admin_to_creator (opt-in AAD-RBAC for caller); lock model-cache storage account to cluster VNet by default; install kubelogin and convert kubeconfig after AKS get-credentials; wire Azure provider tooling; add azure (AKS) terraform module; ship values-aks.yaml AKS overlay with the Azure module
  • Reliability and operations: harden storage_allowed_ip_ranges CIDR validation; harden release guarded merge checks; resubscribe stale NATS health stream; emit az aks get-credentials --overwrite-existing; drop unreachable final_registry guard so ACR path can fire
  • azure-terraform: add grant_admin_to_creator (opt-in AAD-RBAC for caller)
  • azure-terraform: lock model-cache storage account to cluster VNet by default
  • cluster: install kubelogin and convert kubeconfig after AKS get-credentials
  • cluster: wire Azure provider tooling
  • deploy: add azure (AKS) terraform module
  • helm: ship values-aks.yaml AKS overlay with the Azure module
  • azure-terraform: emit az aks get-credentials --overwrite-existing
  • azure-terraform: harden storage_allowed_ip_ranges CIDR validation
  • ci: harden release guarded merge checks
  • cluster: address review feedback on Azure provider wiring
  • cluster: drop unreachable final_registry guard so ACR path can fire
  • cluster: set TF_VAR_* on Azure destroy path (same as create)
  • deploy: revert system pool default to Standard_D4s_v3 (zoned everywhere)
  • gateway: resubscribe stale NATS health stream
  • sidecar: preserve msgpack work item payloads
  • New capabilities: add azure blob payload store support; add server-side copy fast path for cloud weight sync; informational generation eval CI gate over committed floors; add vision (image) input to generate(); preserve text/image content-part ordering; vision (image) input for generate()
  • Reliability and operations: harden cloud cache sync paths; clear HIGH Dependabot alerts (docling, rustls-webpki); ensure cloud weight sync creates local parents; evict stale gateway workers on shutdown; fall back to relay on S3/GCS server-side copy failure
  • Performance: engage conformant image preprocessing for v1; engage conformant image preprocessing for v1 (1.8x)
  • add azure blob payload store support
  • add server-side copy fast path for cloud weight sync
  • bench: informational generation eval CI gate over committed floors
  • generate: add vision (image) input to generate()
  • generate: preserve text/image content-part ordering
  • generate: vision (image) input for generate()
  • support azure blob cluster cache
  • tester-cluster: rtx6000 g7e.4xlarge + sglang preload + hf-token wiring
  • address azure cache review feedback
  • address final cloud storage review issues
  • bench: harden generation eval gate per review
  • deps: clear HIGH Dependabot alerts (docling, rustls-webpki)
  • ensure cloud weight sync creates local parents
  • evict stale gateway workers on shutdown
  • fall back to relay on S3/GCS server-side copy failure
  • generate: address CodeRabbit review on vision input
  • generate: address huronat review on vision input (F2-F8)
  • generate: image-free content_parts field must not shadow layout
  • generate: reject both-present image-bearing content layouts
  • harden cloud cache sync paths
  • loadtest-ci: self-heal orphaned cluster + stale lock in preflight
  • nemo_colembed: trim left-padding rows from v1 conformant doc embeddings
  • normalize local weight sync destination
  • quality-adapter: gate v1 Vidore3 on English; finalize ?lang= plumbing
  • skill: add bash language tag to hfCache —set fenced block (MD040)
  • skill: move inline comments off shell continuation lines so the helm snippet pastes cleanly
  • support cloud source weight sync
  • tester-cluster: update rtx6000-spot machineType doc to g7e.4xlarge to match terraform
  • nemo_colembed: engage conformant image preprocessing for v1
  • nemo_colembed: engage conformant image preprocessing for v1 (1.8x)
  • New capabilities: defer sie-config NATS startup and honor log levels; refresh KEDA Tilt local dev branch; M4 dense encoders — mxbai-embed-large-v1, arctic-embed-l-v2.0, modernbert-embed-base; add daily guarded stable releases
  • Reliability and operations: accept dense dim in qwen3 vl embedding adapter; preserve model query templates in mteb eval; scale single-profile bundles on gpu-agnostic demand; consolidate runtime ninja install; install ninja in cuda runtime
  • defer sie-config NATS startup and honor log levels
  • dev: refresh KEDA Tilt local dev branch
  • models: M4 dense encoders — mxbai-embed-large-v1, arctic-embed-l-v2.0, modernbert-embed-base
  • release: add daily guarded stable releases
  • accept dense dim in qwen3 vl embedding adapter
  • bench: preserve model query templates in mteb eval
  • dev: address KEDA Tilt PR review
  • helm: scale single-profile bundles on gpu-agnostic demand
  • server: consolidate runtime ninja install
  • server: install ninja in cuda runtime
  • server: install ninja in CUDA SGLang runtime
  • terraform: deny non-HTTPS access on state and quality-eval S3 buckets
  • New capabilities: configure GPU disk sizing and generate smoke; support static queue pools
  • Reliability and operations: fail fast on invalid static pool config; pin kind smoke workers to default queue pool; canonicalize static queue pool names; stabilize GPU disk Terraform test
  • configure GPU disk sizing and generate smoke
  • gateway: support static queue pools
  • address GPU disk review comments
  • ci: pin kind smoke workers to default queue pool
  • gateway: canonicalize static queue pool names
  • gateway: fail fast on invalid static pool config
  • stabilize GPU disk Terraform test
  • Breaking change: Queue work subjects and pool streams use the new sie.work.{pool}.{machine_profile}.{bundle}.{model} shape only; legacy subject filters are intentionally not preserved.; workers will subscribe to sie.work.*.<poolName> instead of sie.work.*.default. Deployed alone (without the matching gateway/sidecar update that publishes/filters on the new subject) this will break routing on every cluster. To preserve the old shared-queue behavior, set workers.common.queuePool: "default" explicitly.
  • New capabilities: route work by queue pool lanes; default SIE_POOL to pool name (not “default”)
  • Reliability and operations: harden queue lane admission; bump vitest 2.1.9 -> 4.1.0 (CVE-2026-47429); align lane defaults and tilt e2e; preserve worker-group queue defaults
  • gateway: Queue work subjects and pool streams use the new sie.work.{pool}.{machine_profile}.{bundle}.{model} shape only; legacy subject filters are intentionally not preserved.
  • helm: workers will subscribe to sie.work.*.<poolName> instead of sie.work.*.default. Deployed alone (without the matching gateway/sidecar update that publishes/filters on the new subject) this will break routing on every cluster. To preserve the old shared-queue behavior, set workers.common.queuePool: "default" explicitly.
  • gateway: route work by queue pool lanes
  • helm: default SIE_POOL to pool name (not “default”)
  • deps: bump vitest 2.1.9 -> 4.1.0 (CVE-2026-47429)
  • gateway: harden queue lane admission
  • helm: align lane defaults and tilt e2e
  • helm: preserve worker-group queue defaults
  • Breaking change: workers.pools.<name>.bundle (string), workers.pools.<name>.minReplicas, workers.pools.<name>.maxReplicas, workers.pools.<name>.extraEnv, and workers.pools.<name>.imageBundle are replaced by workers.pools.<name>.bundles.<bundle>.{minReplicas, maxReplicas, extraEnv, imageBundle, enabled}. workers.common.bundle is removed (no longer consumed). StatefulSet, ScaledObject, PDB, and image-prepull DaemonSet names change from worker-{pool} to worker-{pool}-{bundle}, so in-place upgrades require deleting the old resources first.
  • New capabilities: agent-jobs text-gen readiness — code/SQL/tools/guard evals + Qwen3.6-27B + precision routing; transfer sie-cluster claude skill; P(unsafe) logprob threshold for CHECK POLICY precision; split worker pools into pool × bundles schema; surface code/sql/guard capabilities; resolve job aliases in configs/resolve; add sglang worker pool for generative models
  • Reliability and operations: expose unauthenticated metrics scrape port; expose unauthenticated metrics scrape port safely for prom; preserve gateway metrics scrape labels; drop unsupported ebnf advertisement + restore guardian a100 guard threshold; fail-fast on missing Spider DBs + order-sensitive SQL exec accuracy
  • helm: workers.pools.<name>.bundle (string), workers.pools.<name>.minReplicas, workers.pools.<name>.maxReplicas, workers.pools.<name>.extraEnv, and workers.pools.<name>.imageBundle are replaced by workers.pools.<name>.bundles.<bundle>.{minReplicas, maxReplicas, extraEnv, imageBundle, enabled}. workers.common.bundle is removed (no longer consumed). StatefulSet, ScaledObject, PDB, and image-prepull DaemonSet names change from worker-{pool} to worker-{pool}-{bundle}, so in-place upgrades require deleting the old resources first.
  • agent-jobs text-gen readiness — code/SQL/tools/guard evals + Qwen3.6-27B + precision routing
  • agents: transfer sie-cluster claude skill
  • guard: P(unsafe) logprob threshold for CHECK POLICY precision
  • helm: split worker pools into pool × bundles schema
  • models: surface code/sql/guard capabilities; resolve job aliases in configs/resolve
  • tester-cluster: add sglang worker pool for generative models
  • agents: address sie cluster review comments
  • bench: fail-fast on missing Spider DBs + order-sensitive SQL exec accuracy
  • gateway: expose unauthenticated metrics scrape port
  • gateway: expose unauthenticated metrics scrape port safely for prom
  • guard: reject multi-candidate sampling + keep logprobs consistent on rewrite
  • guard: robust verdict thresholding, logprob hygiene, decoded-token logprobs
  • helm: fail-fast on missing/invalid bundle replica bounds
  • helm: preserve gateway metrics scrape labels
  • helm: use sidecar binary for image pre-pull
  • models: drop unsupported ebnf advertisement + restore guardian a100 guard threshold
  • sie_server: honor params.instruction in Florence-2 extract
  • tester-cluster: cap rtx6000 default bundle to avoid over-subscription
  • tools: via-SIE EBNF response_format shape + request/preload model split
  • New capabilities: 5-domain generation bench + via-sie quality matrix + gateway schema gaps; add e0-02 all-minilm time-share experiment; land coalesce_ms=5 + max_batch_requests=12 as Rust defaults; add —via-sie smoke path (route through sie_server); add min_tokens + system_prompt + temperature for G4 retry; close Qwen3.6-27B gap — min_tokens=10 + max=768 + ctx=4096
  • Reliability and operations: set verbose=True on SIEServer so launch errors surface; document worker-sidecar metrics wiring; gate sidecar nats reconnect refresh; harden sidecar config recovery; budget loadtest barrier timeouts
  • Performance: anchor min_batch_cost floor at max_batch_tokens // 4; tighten adaptive wait ceiling + revert gte-multilingual 32k; rebind vision Conv3d patch-embed to F.linear; raise max_batch_tokens 16k → 32k to stop IPC-batch shred
  • 5-domain generation bench + via-sie quality matrix + gateway schema gaps
  • add e0-02 all-minilm time-share experiment
  • batch_config: land coalesce_ms=5 + max_batch_requests=12 as Rust defaults
  • bench-27b: add —via-sie smoke path (route through sie_server)
  • bench-27b: add min_tokens + system_prompt + temperature for G4 retry
  • bench-27b: close Qwen3.6-27B gap — min_tokens=10 + max=768 + ctx=4096
  • bench-27b: launch full SIE stack (NATS+worker+gateway) for —via-sie
  • bench+model: via-sie 4-task n=300 sweep + NEXTN smaller-draft on 27B
  • bench: 0.6B via-sie validated; harness + 27B config gains
  • bench: 5-shot CoT for CaseHOLD (item 5 — close 27B target gap)
  • bench: fix Qwen3-0.6B GPQA (parrot bug) + 27B diagnostics; final matrix
  • bench: improve perf eval output handling
  • docling: accept image input + run on OCR-bench quality path
  • gateway+worker: chat surface accepts min_tokens + chat_template_kwargs
  • gateway: strengthen generation isolation guardrails
  • latency: tighten FetchExpiryController defaults to 2/15/50
  • model+bench: RTX-PRO-6000 FP8 profile for Qwen3.6-27B + 6000 validation
  • model: bump Qwen3-0.6B serving context 1024→4096 for prod simple-task use
  • models: add Marqo/marqo-fashionSigLIP (SigLIP open_clip, fashion image-text)
  • ocr: docling accepts images + quality eval prefers documents
  • reconcile live worker config in sidecar
  • RTX PRO 6000 FP8 profile for Qwen3.6-27B + SIE-on-6000 generative benchmark matrix
  • scheduler: load-aware pipeline_depth autotune (S14 follow-up)
  • scheduler: production-parity defaults + serial pipeline (carveout p99 fix)
  • scheduler: restore SIE_RUST_PIPELINE_DEPTH=2 default (deep-saturation fix)
  • scheduler: SIE_PULL_QUANTUM_INCLUDE_QUEUE_MS for Py-main parity
  • scheduler: SIE_RUST_WAVE_CADENCE env toggle (default on)
  • scheduler: step adaptive controller once per wave (Python parity)
  • sidecar: add worker config and pool admission reconciliation
  • sidecar: wire generation direct dispatch
  • sie_server: add MinerU2.5-Pro-2604-1.2B doc OCR adapter
  • sie_server: carve out QueueExecutor + IPC types for Rust worker POC
  • sie_server: integrate MinerU2.5-Pro-2604-1.2B doc OCR adapter
  • sie_server: UDS msgpack IPC server for Rust worker sidecar
  • sie_worker_rust: close parity gaps with Python pull loop + smoke test
  • sie_worker_rust: scaffold Rust worker sidecar crate (Phase 1c)
  • sie_worker_rust: wire end-to-end NATS -> IPC -> publish loop (Phase 1d)
  • sie-bench: synchronize loadtest measurement start
  • worker/rust: IPC connection pool — lift the sidecar’s last serialization bottleneck
  • worker/rust: narrate the hot path — structured INFO, slow-RPC + heartbeat-streak WARNs, full error chains
  • worker: introduce InferenceBackend trait + BackendRouter
  • worker: native Candle BERT backend behind candle feature
  • accept dense_dim in dense adapters
  • adapters: replace Qwen3-VL vision Conv3d patch-embed with matmul
  • adapters: route Qwen3-VL VLMs through flash attention (Vidore3 throughput)
  • address pr review quality issues
  • bench-27b: drop bundle from SIEServer (sie-server rejects bundle+models combo)
  • bench-27b: set verbose=True on SIEServer so launch errors surface
  • bench-27b: skip chat_template_kwargs on via-sie (gateway rejects unsupported field)
  • bench-27b: wait for sie-server /healthz (not /health)
  • bench: bump casehold/gpqa max_tokens to 2048 (CoT truncation)
  • bench: let via-SIE smoke serve a profile-variant model end-to-end
  • bench: resolve CPU deps for quality server
  • catalog: include eval-matrix tasks so dispatch filter accepts them
  • ci: address analyzer findings and stale queue test
  • ci: avoid nested mise in integration fixture
  • ci: keep sidecar out of warm cache
  • ci: refresh gateway openapi contract
  • correct e0 vm runbook paths
  • deploy: add sidecar registry resources
  • deploy: address server sidecar review feedback
  • deploy: align server sidecar naming
  • deploy: align server sidecar naming and kind preload smoke
  • deploy: align tilt sidecar image naming
  • deploy: document worker-sidecar metrics wiring
  • deploy: keep sidecar on GHCR by default
  • deploy: normalize server sidecar naming
  • deploy: publish server sidecar image
  • deploy: rename sidecar container to worker-sidecar
  • deploy: wire SIE server sidecar for kind smoke
  • deploy: wire worker sidecar image across kind and cloud
  • gate sidecar nats reconnect refresh
  • gateway+server: queue is the only mode — kill direct-mode cruft
  • gateway: suppress H9 first-chunk-fallback on single-worker pools
  • harden sidecar config recovery
  • impact-map: keep profiles distinct when adapter_options differ
  • keep generation machinery off default queue path
  • loader: wire profile runtime.default_sampling into the adapter
  • modal: report actual GPU on remote, not stale env-default
  • model: bump Qwen3.6-27B default/h100 mem_fraction_static 0.85 → 0.92
  • orchestrator: thread CLI -p profile through to client.extract
  • preserve worker batch identity and publish image
  • product: update design audit for topical docs
  • quality_eval: take results-bearing JSON envelope in load_eval_json
  • quality: batch3 of CodeQL findings + bench KIE bug
  • quality: batch3 of CodeQL findings + bench KIE root-cause
  • quality: batches 1+2 of CodeQL quality findings
  • quality: close CodeQL quality-tab findings
  • quality: drop redundant inline imports in donut + registry
  • quality: repair adapter eval harness regressions
  • remove e0 preflight httpx dependency
  • require rust sidecar for queue workers
  • review: 0.6B ctx test 1024->4096, loader except logs, README gaps resolved
  • review: recompute 27B target delta_vs_baseline for the 2048 scores
  • run directory creation
  • run e0 vm scripts via uv
  • scheduler: autotune signal — observed_p50/target_p50 ratio
  • scope bundle config hash cache per registry
  • security: bump astro to ^6.4.2 for website
  • security: bump gateway deps to patched versions
  • security: bump product/gtm Python lockfiles
  • security: bump product/gtm/content/slides npm transitives
  • security: bump root pnpm deps + add overrides for transitives
  • security: bump root Python deps to patched versions
  • security: bump sie_dashboard npm deps to patched versions
  • security: bump sie_ts_sdk standalone pnpm transitives
  • security: bump sst to ^4 to drop vulnerable aws-sdk v2
  • security: cap vite at ^6 + add Node engines to website
  • security: close ~190 Dependabot alerts across 9 manifests
  • security: sanitize one-pager template with DOMPurify
  • security: use Reflect.construct for WebSocket headers shim
  • sie_bench: send SIE profile via X-SIE-MACHINE-PROFILE header
  • sie_bench: use rapidfuzz for OmniDocBench edit distance
  • sie_server: clear CUDA cache on uncovered VLM paths + drop private sem _value access
  • sie_server: VLM cache clears on uncovered paths + drop private sem _value access
  • sie-bench: budget loadtest barrier timeouts
  • slow sidecar nats consumer reconcile
  • smoke: launch sie_server worker with -b sglang, not -m <model>
  • smoke: preload the target model in via-sie worker
  • test: restore donut helper call contract
  • worker-sidecar: harden queue carveout contracts
  • worker/rust: one long-lived pull stream — kill 30s ack_wait stall
  • worker/rust: re-copy src after cargo chef cook so real build isn’t a stub
  • worker/rust: set CUDA_COMPUTE_CAP at build time (default 89, L4)
  • worker/rust: stop shipping the cargo-chef stub binary as the real build
  • worker: harden Candle backend + align dispatcher error contract
  • worker: harden payload store + error paths; surface silent success bugs
  • worker: SGLang adapter accepts min_new_tokens kwarg + 27B via-sie validated
  • adaptive: anchor min_batch_cost floor at max_batch_tokens // 4
  • batching: tighten adaptive wait ceiling + revert gte-multilingual 32k
  • glm_ocr: rebind vision Conv3d patch-embed to F.linear
  • gte-multilingual-base: raise max_batch_tokens 16k → 32k to stop IPC-batch shred
  • mineru_vl: O(L) incremental no-repeat-ngram for greedy decode
  • ocr: swap pure-Python Levenshtein DP for rapidfuzz
  • rope_flash: vectorize CLS/mean pooling, eliminate per-item .item() sync
  • server: FP16 on GPU, coalesce sized for IPC bursts, starvation self-heal
  • restore adaptive batching defaults to 15/50ms
  • scheduler: drop depth autotune (signal didn’t pan out in S17)
  • New capabilities: add Qwen3.6-27B model + migrate to CUDA 12.9
  • Reliability and operations: isolate generation direct dispatch from shared queues; resolve 18 open CodeQL alerts; use SHA256 (not SHA1) for actor_id log tag; colocate tests under infra/, update sync contract
  • server: add Qwen3.6-27B model + migrate to CUDA 12.9
  • isolate generation direct dispatch from shared queues
  • security: resolve 18 open CodeQL alerts
  • security: use SHA256 (not SHA1) for actor_id log tag
  • terraform-sync: colocate tests under infra/, update sync contract
  • security: drop advanced CodeQL setup
  • Breaking change: fail-closed authentication (default-deny)
  • New capabilities: generation quality-gate scoring core (roadmap §5, trust-critical); generation-quality regression gate over the existing scorers; add cohere measurements for us-east-1; add openai measurements for us-east-1; add voyage measurements for us-east-1; regex/EBNF response_format + developer role (roadmap 1.7)
  • Reliability and operations: forward provision_timeout_s in SIEImageTextWrapper.encode; raise image-task eval timeouts to fix Flickr30k nightly; exclude favicon + OG image from auth middleware; keep public surfaces vague about what’s behind auth; revert NextAuth function-form, use try/catch on Resource
  • Performance: warm one Lambda, bump timeout, narrow S3 verdict fetch
  • gateway: fail-closed authentication (default-deny)
  • bench: generation quality-gate scoring core (roadmap §5, trust-critical)
  • bench: generation-quality regression gate over the existing scorers
  • benchmarks: add cohere measurements for us-east-1
  • benchmarks: add openai measurements for us-east-1
  • benchmarks: add voyage measurements for us-east-1
  • chat: regex/EBNF response_format + developer role (roadmap 1.7)
  • dashboard: add executive quality summary widget on landing
  • dashboard: add hover-tooltip on ‘to verify’ explaining WARN
  • dashboard: brand alignment foundation - palette, fonts, header
  • dashboard: brand favicon, opengraph image, light-mode hover fix
  • dashboard: brand foundation - palette, fonts, header logo
  • dashboard: brand surfaces - retokenise landing page + widget
  • dashboard: public preview shell on / so Slack unfurls work
  • dashboard: retokenise badges - brand-elevated neutrals, DM Mono labels
  • dashboard: retokenise landing surfaces onto brand palette
  • dashboard: retokenise loadtest pages onto brand surfaces
  • dashboard: retokenise quality pages onto brand surfaces
  • dashboard: retokenise shared components and sign-in onto brand palette
  • dashboard: treat WARN as passing-with-verify, shade gauge amber
  • gateway,sdk,server: add native generate endpoint with improved admission control and validation
  • gateway,sdk,server: add OpenAI-compatible chat completions with streaming and sampling extensions
  • gateway,server: add multi-turn tool-use support with OpenAI-compatible message format
  • gateway: /v1/completions (legacy OpenAI Completions, raw-prompt)
  • gateway: /v1/completions streaming (text_completion SSE)
  • gateway: /v1/generate accepts seed/logprobs/logit_bias/n/best_of/lora_adapter (M8)
  • gateway: /v1/responses (OpenAI Responses API, MVP)
  • gateway: /v1/responses structured array input (conversation history)
  • gateway+worker: per-choice OpenAI streaming for n>1 (H4, H5, M4)
  • gateway: accept OpenAI multimodal content-parts; reject images (no VL model)
  • gateway: add routing salt + byte-preserving key mode (M11)
  • gateway: advertise lora_adapters on /v1/models + pre-validate unknown names
  • gateway: fail-closed authentication (default-deny)
  • gateway: meaningful system_fingerprint on chat responses (roadmap 1.3/§5)
  • gateway: refactor streaming and routing with improved error handling and metrics
  • gateway: register /v1/moderations as explicit 501 (roadmap 1.8, phase 3)
  • gateway: serve a rendered API reference at /docs (Redoc)
  • gateway: unify /v1/embeddings on the OpenAI error envelope (roadmap 1.4)
  • generation: best_of — over-generate + rank by logprob, return top n
  • generation: complete M4 req2 generation primitive with streaming, structured outputs, and routing
  • generation: multi-candidate n>1 (non-streaming) end-to-end (roadmap 1.5)
  • generation: multi-LoRA serving (one base, N adapters, per-request) (roadmap 6.2)
  • generation: ship generate() primitive — Qwen3.5-4B + NEXTN/MTP + xgrammar, adapter perf at parity with raw SGLang
  • generation: streaming n>1 — per-candidate SSE interleave
  • helm/sie-cluster: bundle cert-manager + trust-manager (opt-in) with self-signed TLS mode
  • openapi: add tool_calls support to chat completion schema
  • python-sdk: expose typed params for chat n/logprobs/lora_adapter/etc (M7)
  • quality-eval: add heartbeat logging and improve long-running process observability
  • routing: cache-aware (prefix-hash) routing (roadmap §6.3)
  • sie_bench: add Cohere as a first-class eval source
  • sie_bench: add Cohere multimodal embeddings
  • sie_bench: add Cohere rerank backend for native MTEB rerank tasks
  • sie_bench: add OpenAI Embeddings as a first-class eval source
  • sie_bench: add Voyage provider source plumbing
  • sie_bench: implement Voyage text embedding runner
  • sie_dashboard: add /quality/compare to diff two quality runs
  • terraform: add default_tags Project=sie/Cluster on all AWS providers
  • terraform: add on-demand RTX 6000 baseline pool to tester-cluster
  • terraform: add uniform project=sie label across all GCP clusters
  • terraform: idle-stop and on-demand wake for quality-eval runner fleet
  • terraform: scale quality-eval fleet to 5+5, smart-wake, 4h timeout
  • terraform: wake-runners retries + watchdog queued-jobs backstop
  • tester-cluster: add on-demand L4 worker pool + capacityType node pins
  • ts-sdk: handle 202 provisioning in chatCompletions + expose missing fields (H1+M6)
  • worker: SGLang owns grammar; worker preflight opt-in only (H8, ADR-0002)
  • worker: wire mixed-pool fairness scheduler into the pull-loop (opt-in)
  • worker: WorkClassScheduler core for mixed-pool fairness (roadmap §6.1)
  • bench: declare olmocr[bench] dep for OCR-bench quality eval
  • bench: drop —with-deps from playwright install (no sudo on g7e)
  • bench: forward provision_timeout_s in SIEImageTextWrapper.encode
  • bench: install playwright chromium for OCR-bench KaTeX rendering
  • bench: playwright install —with-deps for OCR-bench chromium
  • bench: raise image-task eval timeouts to fix Flickr30k nightly
  • bench: respect similarity() inputs in ColBERT/ColPali wrappers
  • bench: run olmocr.bench.tests off orchestrator’s asyncio loop
  • centralize worker_id subject normalization (M5)
  • chart: pin image-prepull DaemonSets to GPU nodes
  • ci: copy assets/ into the gateway Docker build (redoc bundle)
  • dashboard: close remaining ‘SIE Dashboard’ leaks in <title> and og:image:alt
  • dashboard: drop edge runtime on opengraph-image for OpenNext
  • dashboard: exclude favicon + OG image from auth middleware
  • dashboard: keep public surfaces vague about what’s behind auth
  • dashboard: pick healthy daily via coverage + health gates
  • dashboard: require >=50 pairs on main-run fallback in nightly picker
  • dashboard: revert NextAuth function-form, use try/catch on Resource
  • dashboard: short tab title for authed users, neutral for unauth
  • dev: set explicit auth opt-in for local gateway launchers (post fail-closed)
  • gateway,sdk,server: add wire-level validation, improve resource cleanup, and enhance observability across request lifecycle
  • gateway,sdk,server: prevent metric cardinality DoS and fix non-idempotent retry logic
  • gateway,sdk,server: strengthen validation and eliminate silent failures across request lifecycle
  • gateway,sdk,server: validate numeric fields and improve error handling
  • gateway: add NATS config trusted producers helm override
  • gateway: document ModelCapabilities in OpenAPI + refresh on profile delta-update
  • gateway: generation timeouts bypass legacy request-timeout ceiling (H7)
  • gateway: scope LoRA adapter capabilities per profile (M10)
  • gateway: strict allow-list + 400 contract on /v1/completions (H3)
  • gateway: strict allow-list + 400 contract on /v1/responses (H2)
  • gateway: tighten chat sampler/token-cap + tool-history validation (M1, M13)
  • gateway: trust chart-rendered sie-config pod name for NATS deltas
  • generation: cancel tombstone prevents first-chunk fallback double-execution (H9)
  • generation: LoRA lora_path is a top-level /generate field, not a sampling param
  • generation: tighten lossy tool-control flags (M14)
  • grammar: resolve tokenizer adapter for Outlines processor factories and remove anchors from regex patterns
  • helm/sie-cluster: guard validateTls probe against deployments with nil labels
  • helm/sie-cluster: guard validateTls probe against nil deployment labels
  • helm/sie-cluster: include cert-manager mode in presence-check gate
  • helm/sie-cluster: label-based cert-manager detection + bidirectional runtime check
  • helm/sie-cluster: make self-signed root-CA namespace configurable
  • helm/sie-cluster: one-step bundled cert-manager install with self-signed TLS
  • helm/sie-cluster: probe cert-manager controllers cluster-wide
  • helm/sie-cluster: regenerate Chart.lock with synced digest
  • helm: trim and drop empty entries in ingress.hosts
  • modal: exclude cargo target/ from sandbox image mount
  • pytorch_embedding: accept and forward revision kwarg
  • quality_eval: tolerate stdout noise around eval JSON envelope
  • quality-eval: handle paginated jobs API and align last two filters
  • release-docker,warm-cache: address review findings
  • server,gateway: add GPU-aware health probes to detect and recover from wedged CUDA contexts
  • server,gateway: GPU-aware health probes to detect & recover from wedged CUDA contexts
  • sie_bench: register OPENAI_SOURCE so —save-targets openai actually saves
  • sie_dashboard: address compare-page review nits
  • sie_dashboard: label compare log links with run IDs
  • sie_server: base64-decode JSON image inputs
  • sie_server: enforce media bytes contract at every consumer
  • sie_server: install cv2 system libs for docling extract
  • streaming: no-silent-drop on chunk-queue backpressure (H6)
  • terraform: detect in-flight workflow runs via explicit status query
  • terraform: drop redundant Project overrides in quality-eval-l4
  • terraform: drop watchdog_idle_minutes default to 5, ignore PIP drift
  • terraform: grant ec2:DescribeInstanceStatus to wake role
  • terraform: per-runner idle-stop via GitHub Actions runners API
  • terraform: require positive activity observation before idle-stop
  • terraform: scope quality-eval IAM on Role tag instead of Project
  • terraform: seed GPU node-group desired_size from min_size
  • terraform: watchdog backstop covers queued-status runs
  • terraform: watchdog grants HANG_MINUTES grace from LaunchTime
  • terraform: watchdog ignores LastBusyAt older than LaunchTime
  • tester-cluster: pin worker pool nodeSelectors to gpu-type as well
  • dashboard: warm one Lambda, bump timeout, narrow S3 verdict fetch
  • release-docker: drop matrix consolidation, keep deps push retry
  • New capabilities: default payload store to model-cache bucket /payloads; typed InputTooLongError for extract 400 INPUT_TOO_LONG
  • Reliability and operations: bump dev-g6-spot to g6.2xlarge; bump dev-g6-spot to g6.2xlarge so default worker pool fits; default workers to shared queue pool; pin opencv-python-headless to drop X11 runtime deps; tolerate config conflicts in bootstrap, gate on sie-config
  • infra: default payload store to model-cache bucket /payloads
  • sdks: typed InputTooLongError for extract 400 INPUT_TOO_LONG
  • aws-example: bump dev-g6-spot to g6.2xlarge
  • aws-example: bump dev-g6-spot to g6.2xlarge so default worker pool fits
  • chart: default workers to shared queue pool
  • cluster.py,aws.py: address review suggestions 1 & 2
  • deps: pin opencv-python-headless to drop X11 runtime deps
  • gateway: tolerate config conflicts in bootstrap, gate on sie-config
  • gateway: tolerate config conflicts in bootstrap, gate on sie-config ready
  • sdk: widen sie-sdk requires-python to >=3.12
  • terraform-aws: default ECR creation off, prefix repo names with project_name
  • terraform-aws: trim slashes from ecr_repository_prefix
  • terraform-google-sie: wait for identity pool before binding WI
  • terraform: relax required_version from ~> 1.14.3 to >= 1.14
  • New capabilities: add ColQwen3 + Nemotron ColEmbed v2 visual doc retrieval; add text classification task support; add post-download load timeout with stall-based download bounds; add scope-able workflow_dispatch with model/profile/task filters; add new INPUT_TOO_LONG ErrorCode; enforce overflow_policy in gliclass adapter
  • Reliability and operations: surface empty matrix and add measurement-mode for unbaselined adapters; emit task_class in quality-adapter JSON output; score detection eval predictions from result[“objects”]; annotate empty-diff path that bypasses impact_map; capture real exit code from impact_map in resolve-impact.sh
  • adapters: add ColQwen3 + Nemotron ColEmbed v2 visual doc retrieval
  • extraction: add text classification task support
  • model-loader: add post-download load timeout with stall-based download bounds
  • quality-adapter: add scope-able workflow_dispatch with model/profile/task filters
  • server: add new INPUT_TOO_LONG ErrorCode
  • server: enforce overflow_policy in gliclass adapter
  • server: route INPUT_TOO_LONG to HTTP 400 in extract API
  • server: validate overflow_policy in resolve_runtime_options
  • bench: emit task_class in quality-adapter JSON output
  • bench: score detection eval predictions from result[“objects”]
  • ci: annotate empty-diff path that bypasses impact_map
  • ci: capture real exit code from impact_map in resolve-impact.sh
  • ci: collapse adapter-equivalent profiles in quality-adapter matrix
  • ci: pin mise to 2026.5.5 in loadtest workflows
  • ci: surface empty matrix and add measurement-mode for unbaselined adapters
  • docker: install libspatialindex-c6 in worker images
  • gliclass: catch IndexError empty-tensor crash as InputTooLongError
  • gliclass: raise InputTooLongError from argmax-empty backstop
  • probe-chart: make 3-OLD/2-NEW sample asymmetry explicit in title
  • quality-adapter: namespace quality_eval tests to avoid conftest collision
  • quality-adapter: split adapter_paths on commas before —changed-dirs
  • quality-adapter: split Pair column so pair_key | stops breaking the table
  • quality-adapter: split Pair column to stop pair_key | breaking the table
  • New capabilities: default score_pairs() in BaseAdapter + baseline reranking targets; bump cold-start schema to v6 with deserialize/warmup split; per-model perf concurrency defaults for OCR adapters; adapter-triggered quality eval on persistent L4 runner; make destroy conditional on workflow_dispatch input; nightly loadtest pipeline + baseline recorder
  • Reliability and operations: raise loadtest job timeout to GH Actions ceiling (360 min / 6h); clarify experimental NATS health mode; lang tag on fenced block; fail fast on invalid scenarios; surface error/no_results rows in MD; harden parse_label against unexpected filenames; adapt OCR perf Item shape to model.inputs and fail loudly on errors
  • adapters: default score_pairs() in BaseAdapter + baseline reranking targets
  • bench,charts: bump cold-start schema to v6 with deserialize/warmup split
  • bench: per-model perf concurrency defaults for OCR adapters
  • ci: adapter-triggered quality eval on persistent L4 runner
  • ci: make destroy conditional on workflow_dispatch input
  • ci: nightly loadtest pipeline + baseline recorder
  • colbert: add score_pairs support and expand model coverage
  • dashboard: add status and kind filters to runs list
  • dashboard: introduce run-group concept (run = 3 scenarios)
  • dashboard: loadtest results dashboard (Next.js + SST + DynamoDB)
  • dashboard: render every metric in the perf-lab archive
  • dashboard: scaffold loadtest dashboard (Next.js + SST)
  • dashboard: status and kind filters on runs list
  • dashboard: track run_status; gh-API one-time backfill
  • docling: add ocr profile defaulting do_ocr=true
  • gateway: expose OpenAPI contract
  • gateway: unify API errors and align probe contracts
  • helm: expose probes value trees for worker/gateway/config
  • helm: tighten startup/readiness probes for faster pod-ready
  • helm: TLS termination via cert-manager + BYO matrix docs
  • helm: wire probe templates to values trees
  • infra: opt-in S3 cluster model cache
  • ltfr: cache-vs-no-cache compare chart, with-cache run data, and 8 single-mode chart refresh
  • matrix: add task_class stamping to eval measurements
  • server,bench: split deserialize/warmup in cold-start instrumentation (v6)
  • server: cap torch CPU threads at worker startup
  • sie_server: per-stage timing markers in lifespan for engine_boot attribution
  • sie_server: split adapter.warmup() out of load() with cold-start log markers
  • tools: bump cold-start bench to v5 with scenario flag
  • tools: LTFR per-scenario bench tooling + results (issue #652)
  • tools: ltfr-bench orchestrator (issue #652)
  • bench: adapt OCR perf Item shape to model.inputs and fail loudly on errors
  • bench: address CodeRabbit review on PR #779
  • bench: correctly detect v6 split presence in flattened runs[]
  • bench: derive emitted gpu_load_s from v6 deserialize+warmup split when available
  • chart: pass —cluster-cache to sie-server and correct populate command in docs
  • charts: vertical legend so ‘image pull + container init’ and ‘node prov’ aren’t clipped
  • ci+terraform: three deterministic root causes for loadtest pipeline
  • ci: address CodeRabbit findings on quality-adapter PR
  • ci: address CodeRabbit’s second-pass review on quality-adapter
  • ci: address CodeRabbit’s third-pass review on quality-adapter
  • ci: auto-clear stale terraform state lock from prior runner crashes
  • ci: drop double cuda12 suffix + force codebuild for missing images
  • ci: forensic dump on argo failure + LB-ENI release before destroy
  • ci: gate stale-lock clear behind force_unlock input + pass —aws-region to destroy
  • ci: override registry/gpu-selector/tolerations + Python heredoc
  • ci: parse markdown bench output → result.json synthesis
  • ci: pass WORKFLOW env to run_scenarios.sh in loadtest.yml
  • ci: preflight env-var check in run_scenarios + finalize scripts
  • ci: provision GH PAT secret + in-cluster github-token before bootstrap
  • ci: raise loadtest job timeout to GH Actions ceiling (360 min / 6h)
  • ci: read bench-config from local clone, not raw.githubusercontent.com
  • ci: right-size bench pod + worker pod resources for cluster shape
  • cluster: move orphan-LB sweep into cmd_destroy, drop parallel script
  • cluster: use project_name (not example name) for orphan-LB VPC tag
  • cluster: use project_name for orphan-LB VPC tag lookup
  • dashboard,ci: keep run_status consistent between S3 and DynamoDB
  • dashboard,ci: wire real Prometheus matrix shape + extend headlines
  • dashboard: drop time-based legacy run grouping (was unsafe)
  • dashboard: GPU util shown as 0-100 (was being multiplied by 100 again)
  • dashboard: include duration_seconds in run-meta.json (was DynamoDB-only)
  • dashboard: normalize array-shaped searchParams before .trim()
  • deps: bump plotly to >=6.1.1 for kaleido compat
  • docker: stub bundles/ and models/ in deps stage
  • docling: cache DocumentConverter per (device, ocr_enabled)
  • docling: mark adapter unloaded in unload()
  • docling: thread device through PdfPipelineOptions accelerator_options
  • gateway: address latest coderabbit contract notes
  • gateway: address PR review for NATS health mode
  • gateway: address probe and SDK review findings
  • gateway: align CreatePoolRequest OpenAPI with runtime validation
  • gateway: clarify experimental NATS health mode
  • gateway: close remaining review contract gaps
  • gateway: preserve embeddings timing headers
  • gateway: preserve scale-from-zero request path
  • gateway: reject unsupported embeddings token arrays
  • helm: validate ACME server and privateKeySecretRef in validateTls
  • infra: grant kms:Decrypt to workers when model cache uses SSE-KMS
  • infra: normalize whitespace-only model cache string inputs
  • infra: treat empty model_cache_kms_key_id as unset
  • infra: use flat lifecycle key for s3-bucket module v5
  • loadtest-ci: force-delete orphan elbv2 LBs before terraform destroy
  • loadtest-ci: force-delete orphan LBs and stop swallowing destroy failures
  • loadtest-ci: poll workflow phase instead of argo submit —wait
  • loadtest-ci: poll workflow phase instead of relying on argo —wait
  • ltfr-bench,notes: lang tag on fenced block; fail fast on invalid scenarios; surface error/no_results rows in MD
  • ltfr-bench: hoist imports to top; guard payload.results shape
  • ltfr-bench: mark request_failed rows in scenario-a/b MD tables
  • ltfr-bench: preserve failure context in aggregated rows; add request_failed status
  • ltfr-bench: treat no_results cells as failures in exit code
  • ltfr-charts: strip legend clip-path so labels render full width
  • ltfr: tighten UID/timestamp guards in capture_image_pull_events
  • multi_pod_cold_start: raise on ASG terminate fail; UID-filter pull events; isolate scenario-c pod
  • paddleocr_vl: pass use_cache=True to generate
  • paddleocr_vl: pass use_cache=True to generate to enable KV-cache
  • review: tighten score_pairs options handling and query text validation
  • sie-server: include model and bundle directories in wheel distribution
  • terraform: detect HF-cache EBS by NVMe model + size, not by Linux name
  • terraform: set resolve_conflicts_on_* = OVERWRITE on EKS addons
  • tools: drop module-level docstrings (AGENTS.md rule)
  • tools: guard fig_per_cell_table aggregation against empty results
  • tools: guard mean() against empty engine_boot_s in aggregate()
  • tools: harden parse_label against unexpected filenames
  • tools: mark cold-start-bench.py executable (EXE001)
  • tools: remove module docstring from cold_start_charts.py (repo rule)
  • New capabilities: add OmniDocBench OCR quality loader; support /v1/score with dense/sparse/colbert/hybrid modes; add Marqo/marqo-ecommerce-embeddings-B via open_clip backend
  • Reliability and operations: add terminal failed state to model registry (sie-test#85)
  • bench: add OmniDocBench OCR quality loader
  • bge-m3: support /v1/score with dense/sparse/colbert/hybrid modes
  • siglip: add Marqo/marqo-ecommerce-embeddings-B via open_clip backend
  • server: add terminal failed state to model registry (sie-test#85)
  • Breaking change: openapi.json is now a committed artifact that must be regenerated and committed when API changes are made
  • New capabilities: add GLM-OCR adapter; add Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B multimodal adapters; add GLiNER2 and GLiNER-bi adapters; add Qwen3-Reranker-0.6B and 4B causal LM reranker support; add SigLIP 2 base-patch16-224 vision-language encoder; add minimal cache weights snapshot command for offline deployments
  • Reliability and operations: retry only transient connection errors under wait_for_capacity; surface unrouteable models loudly and helm-repo-add on pristine hosts; emit identical NATS payload to bundle and _all subjects; surface mixed-profile unrouteable models and keep snapshot consistent on writes; add retry logic for deadsnakes PPA to handle Launchpad outages
  • Performance: cache JPEG-encoded corpus images across queries; lazily JPEG-encode corpus images on first use; cache SDK version parse, integer audit latency, UUIDv7; cut hot-path allocations, fuse numpy decode, tighten backpressure
  • openapi: openapi.json is now a committed artifact that must be regenerated and committed when API changes are made
  • adapters: add GLM-OCR adapter
  • adapters: add Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B multimodal adapters
  • add GLiNER2 and GLiNER-bi adapters
  • add Qwen3-Reranker-0.6B and 4B causal LM reranker support
  • add SigLIP 2 base-patch16-224 vision-language encoder
  • admin: add minimal cache weights snapshot command for offline deployments
  • bench: honor SIE_BENCH_SERVER_READY_TIMEOUT in eval orchestrator
  • ci: nightly loadtest gate against dedicated EKS cluster
  • ci: nightly loadtest gate, ephemeral cluster per run
  • extract: add Docling adapter for PDF/DOCX/HTML extraction
  • extract: add Docling adapter for PDF/DOCX/HTML parsing
  • extract: plumb document items and structured data results
  • observability: add Prometheus metrics to sie-config and expand sie-gateway coverage
  • observability: Prometheus metrics for sie-config and sie-gateway
  • oom: implement defensive exception fan-out and improve recovery metrics
  • oom: improve error semantics and budget exhaustion detection
  • openapi: add static spec export and validation
  • router: import Rust gateway source tree
  • server: add reactive OOM recovery and proactive idle eviction
  • social: daily social content pipeline with 5-source drafts + engagement
  • types: add document input modality across SDKs, server, and metadata
  • adapters: add input validation guards for empty/failed visual inputs
  • adapters: address review findings for Qwen3-VL adapters
  • adapters: clarify video placeholder, validate token IDs, fix torch_dtype key
  • add client-side hour filter to search_x_posts (was date-level only)
  • address CodeRabbit review feedback
  • address follow-up PR review nits
  • address remaining CodeRabbit feedback (round 2)
  • address review findings — negative truncation guard, score() options, constant dedup
  • bench: show correct unit labels for MP/s throughput in —print-gap
  • bundles: declare Qwen3-VL adapters in default bundle
  • ci: use Blacksmith runner in CI
  • client: retry only transient connection errors under wait_for_capacity
  • cluster: address PR #701 review comments
  • cluster: correct kubectl flag combo and reorder LB sweep before helm uninstall
  • cluster: helm uninstall before terraform destroy to clean up AWS LB leftovers
  • cluster: unblock end-to-end mise run cluster create --build
  • config,cluster: surface unrouteable models loudly and helm-repo-add on pristine hosts
  • config: emit identical NATS payload to bundle and _all subjects
  • config: surface mixed-profile unrouteable models and keep snapshot consistent on writes
  • docker: add retry logic for deadsnakes PPA to handle Launchpad outages
  • docker: propagate failure when all add-apt-repository retries exhausted
  • docling: per-task converter, hf_revision guard, callable typing (CodeRabbit)
  • docs: Update packages/sie_server/Dockerfile.cuda11
  • fail closed on missing/unparsable timestamps in lookback filter
  • gateway,config,sdk: resiliency, concurrency, and cross-service hash parity
  • gateway,config: address PR review — 404 for unknown models, 202 on default routing, full YAML propagation
  • gateway,config: harden auth, trusted NATS producers, and recovery path; drop gateway HA default
  • gateway,sdk: map upstream timeouts to 503+MODEL_LOADING for SDK retry
  • gateway: add GET /v1/models/{model} detail route
  • gateway: address PR #716 review feedback
  • gateway: align /v1/models error and list shapes
  • gateway: drop double-counted REQUEST_COUNT / REQUEST_LATENCY emit
  • gateway: emit X-SIE-Error-Code header on model-loading 503
  • gateway: keep record_request async to match main’s call shape
  • gateway: make sie-config single source of truth for bundles with live resync
  • gateway: normalize model ids in NATS work subjects + docs/tooling/ha cleanup
  • gateway: pre-instantiate request/demand metric families on startup
  • gateway: prioritize epoch-rewind branch; harden no-thrash test; correct arch-guide on ephemeral restart
  • guard score() and score_pairs() against empty input lists
  • helm: default clusterRouting to “queue” on import-sie-router-rust
  • helm: enable NATS + JetStream by default to match queue clusterRouting
  • helm: fail fast when gateway has no bundle source
  • kind-smoke: add —no-pool-isolation for static clusters + contract-drift fixes
  • kind-smoke: address bot review feedback
  • kind-smoke: enable configStore and harden config/gateway tests
  • kind-smoke: enable JetStream on test NATS and drop duplicate subchart
  • kind-smoke: wire sie-config image and helm overrides into kind cluster fixture
  • kind-smoke: wire sie-config image into kind cluster fixture
  • observability: address PR review blockers on metrics PR
  • sdk: cluster cache prefix probe uses list, not head
  • sdk: cluster cache prefix probe uses list, not head (Refs #732, #654)
  • sdk: has_children filters folder-marker objects (Refs #732, #654)
  • sdk: preserve caller-supplied document format over inferred (CodeRabbit)
  • sdk: retry mid-flight transport disconnects, not just timeouts
  • sdk: retry on connection errors and generic 503s
  • sie_config: address PR review feedback
  • terraform/aws: set 100GB root volume on cpu node group to avoid DiskPressure
  • terraform/gcp: undo router→gateway rename on GCP Cloud Router + NAT
  • tests: include sie-config in expected missing-image list
  • tests: restore docker gateway smoke test after router rename
  • tmux-scripts: improve robustness of session parsing and argument handling
  • types: adapt to ty 0.0.32 stricter ignore handling
  • use searchTerms for X tweet-scraper actor (was searchQueries)
  • bench: cache JPEG-encoded corpus images across queries
  • bench: lazily JPEG-encode corpus images on first use
  • docker: add —link + move ARG BUNDLE to eliminate cross-bundle layer noise
  • docker: normalize mtimes so shared venv layer is dedupable
  • docker: reorder stages for maximum BuildKit cache reuse
  • docker: split worker venv into shared + bundle-specific layers
  • gateway: cache SDK version parse, integer audit latency, UUIDv7
  • gateway: cut hot-path allocations, fuse numpy decode, tighten backpressure
  • gateway: fuse msgpack_numpy decode into the response path
  • gateway: move score-endpoint unwrap instead of cloning
  • gateway: pass msgpack items through as rmpv::Value
  • gateway: publish work items concurrently + borrow shared fields
  • gateway: tighten cold-pool backpressure + cheaper QPS counter
  • gateway: trim per-request work on the inference hot path
  • Breaking change: Removed --model CLI args from worker startup; use SIE_PRELOAD_MODELS env var or --preload flag instead
  • New capabilities: add ModernBERT flash dense embedding support with fallback mechanism; add OCR quality benchmarks (olmOCR-bench); add OCR quality benchmarks with olmOCR-bench; add pages/sec throughput metric for OCR perf eval; add perf metrics to OCR eval pipeline; also report query throughput in mpix/s for image queries
  • Reliability and operations: add missing NATS Helm repo to release workflow; don’t set NODE_AUTH_TOKEN for OIDC npm publishes; harden affinity spill with bounds check, clamp, and debug log; make rejected requests visible to KEDA scaling metrics; remove redundant tokenizer validation and unused template parameter
  • workers: Removed --model CLI args from worker startup; use SIE_PRELOAD_MODELS env var or --preload flag instead
  • adapters: add ModernBERT flash dense embedding support with fallback mechanism
  • bench: add OCR quality benchmarks (olmOCR-bench)
  • bench: add OCR quality benchmarks with olmOCR-bench
  • bench: add pages/sec throughput metric for OCR perf eval
  • bench: add perf metrics to OCR eval pipeline
  • bench: also report query throughput in mpix/s for image queries
  • benchmarks: add MTEB NFCorpus evaluation results for ModernBERT-based embedders
  • bench: report vision corpus throughput in mpix/s
  • bench: report vision corpus throughput in mpix/s instead of items/s
  • deps: migrate from pynvml to nvidia-ml-py package
  • haystack: add haystack_integrations namespace aliases
  • haystack: add namespace-convention aliases
  • observability: add anonymous usage telemetry
  • sdk: add max_concurrency param to SIEAsyncClient to prevent connection pool exhaustion
  • server: add lightonai/LightOnOCR-2-1B OCR adapter with next bundle
  • workers: implement model preloading at startup to reduce first-request latency
  • adapters: remove redundant tokenizer validation and unused template parameter
  • address PR review — panel title, namespace variable
  • bench: handle unloaded images in pixel count computation
  • bench: use concurrent async requests for OCR perf eval
  • bench: validate image entries before computing pixel counts
  • bench: validate pixel counts before using them for image corpus throughput
  • build: downgrade dockerfile syntax version to 1 for broader compatibility
  • ci: add missing NATS Helm repo to release workflow
  • ci: don’t set NODE_AUTH_TOKEN for OIDC npm publishes
  • dashboard: queue routing dashboard accuracy and usability
  • docs: add update date to portfolio header
  • docs: correct PR reference in reranker reclassification note
  • docs: populate reranker data and simplify table header
  • docs: update stale model counts after reranker reclassification
  • haystack: rename namespace alias to sie
  • install uv via curl instead of COPY —from ghcr.io
  • preload smoke test checks model.loaded instead of nonexistent workers field
  • readme: heading format
  • release: add LanceDB integrations to release-please config
  • router: add overflow spill to break model affinity deadlock
  • router: harden affinity spill with bounds check, clamp, and debug log
  • router: make rejected requests visible to KEDA scaling metrics
  • sie-bench: account for in-flight drain in throughput calculation
  • sie-bench: use union wall-clock for multiprocess throughput merge
  • tester-cluster: patient KEDA scale-down for worker pools
  • New capabilities: add async, chunking, and streaming to Weaviate document enricher; improve DLQ routing and score response handling; implement Config Management API with NATS-based distribution and review fixes; add LanceDB integration (Python + TypeScript); queue routing dashboard + NATS exporter + router image tag; queue routing dashboard, NATS prom exporter, router image tag
  • Reliability and operations: correct cluster routing condition, stream max_age units, and reconnect state ordering; add recreate strategy for router deployment when nats config restore is enabled; restore Chart.yaml deps from main, keep appVersion v-prefix; queue routing dashboard PromQL for NATS wait; configurable NATS fetch budget, Helm-wired queue params
  • Performance: decouple scanner and SIE batch sizes in enrich_table; stream enrich_table batch-by-batch instead of full materialization; use Lance scanner for column projection in enrich_table; bypass FastAPI for hot proxy paths via raw ASGI middleware
  • add async, chunking, and streaming to Weaviate document enricher
  • dlq,pull-loop: improve DLQ routing and score response handling
  • implement Config Management API with NATS-based distribution and review fixes
  • integrations: add LanceDB integration (Python + TypeScript)
  • observability: queue routing dashboard + NATS exporter + router image tag
  • observability: queue routing dashboard, NATS prom exporter, router image tag
  • sdk: add get_model() and configure LanceDB release workflows
  • terraform: add AWS eval-eu EKS cluster with multi-GPU support
  • terraform: add evaluation cluster setup for AWS with multi-GPU support and updated configurations
  • terraform: add node labels, adjust pool sizes for tester cluster
  • config,queue,nats: correct cluster routing condition, stream max_age units, and reconnect state ordering
  • handle BytesIO images in LlamaIndex and validate Weaviate classify config
  • helm: add recreate strategy for router deployment when nats config restore is enabled
  • helm: restore Chart.yaml deps from main, keep appVersion v-prefix
  • helm: use generic release-please updater for appVersion
  • helm: use generic updater for both Chart.yaml version fields
  • helm: use l4-spot/rtx6000-spot naming convention for spot profiles
  • integrations: address CodeRabbit review findings for LanceDB PR
  • observability: queue routing dashboard PromQL for NATS wait
  • queue-routing: configurable NATS fetch budget, Helm-wired queue params
  • queue-routing: resolve bugs, add configurable NATS params, fix score wire format
  • queue-routing: score response format and DLQ fallback routing key
  • release: use NPM_TOKEN for initial sie-lancedb publish
  • router: use “scores” key in queue-mode score responses
  • terraform: add GPU subnet coverage validation
  • terraform: relax AZ validation and clarify defaults
  • terraform: review fixes for tester cluster infra
  • terraform: switch tester-cluster to us-east-2
  • terraform: Switch tester-cluster to us-east-2 and update deployment docs
  • terraform: validate gpu_node_groups for duplicate and reserved names
  • test: add buildx builder pause recovery and improve build error diagnostics
  • update adapter tests and address code review feedback
  • use OCI registry URI for helm chart in README
  • lancedb: decouple scanner and SIE batch sizes in enrich_table
  • lancedb: stream enrich_table batch-by-batch instead of full materialization
  • lancedb: use Lance scanner for column projection in enrich_table
  • router: bypass FastAPI for hot proxy paths via raw ASGI middleware
  • router: reduce thread pool pressure by inlining small deserialization
  • router: remove msgpack_numpy global patch and BaseHTTPMiddleware
  • router: replace stdlib json with orjson for 3-10x faster serialization
  • sdk+router: lazy msgpack_numpy.patch and pure ASGI middleware
  • Reliability and operations: increase docker smoke test timeouts and add retry; include $platform in worker image tag format; revert pool names to machine profile names; remove —provenance flag (requires public repo)
  • helm: include $platform in worker image tag format
  • helm: revert pool names to machine profile names
  • increase docker smoke test timeouts and add retry
  • remove —provenance flag (requires public repo)
  • Reliability and operations: add sie-qdrant and sie-weaviate to release-please config; point sync-terraform default repos to production; remove —provenance from npm publish for private repo; correct image.tag comment to reflect actual format; remove duplicate platform suffix from worker image tag
  • add sie-qdrant and sie-weaviate to release-please config
  • ci: point sync-terraform default repos to production
  • ci: remove —provenance from npm publish for private repo
  • helm: correct image.tag comment to reflect actual format
  • helm: remove duplicate platform suffix from worker image tag
  • remove internal-only references from COMPATIBILITY.md
  • New capabilities: add profiling script for sparse encoding hot path; add GitHub Actions workflow to sync Terraform modules to registry repos; apply QoL improvements from PR #484 review comments; switch default GPU from g5 (A10G) to g6 (L4); add rerank/score support to TEI runner; implement configurable document length limits and custom prefix token registration
  • Reliability and operations: restore triggering ref for source checkout; restore quality by enabling causal attention and QK-normalization; restore dev-l4-spot zones to us-central1 for GPU availability; check /metrics endpoint in test_prometheus_metrics_exist; add per-attempt timeout to lease renewal fetch
  • Performance: optimize MoE expert dispatch with sorted-expert routing; batch MaxSim scoring across documents on GPU; batch sparse aggregation with segment_reduce and fuse relu; batch split_embeddings + validate ColBERT performance
  • adapters: add profiling script for sparse encoding hot path
  • add GitHub Actions workflow to sync Terraform modules to registry repos
  • apply QoL improvements from PR #484 review comments
  • aws: switch default GPU from g5 (A10G) to g6 (L4)
  • bench: add rerank/score support to TEI runner
  • colbert: implement configurable document length limits and custom prefix token registration
  • deploy: move namespace, SA, and HF token secret management to Helm chart
  • deploy: prepare Terraform modules for public registry publishing
  • deploy: rewrite example module sources to registry references
  • deploy: rewrite Helm and internal references for public release
  • deploy: two-artifact model — GCP Terraform infra-only, batteries-included Helm chart
  • docker: add —docker-platform flag to docker build task
  • extend create_pool API/SDK with minimum_worker_count and bundle
  • helm: add batteries-included sub-chart dependencies to sie-cluster
  • helm: add image pre-pull DaemonSet for GPU worker pools
  • helm: add step to build Helm chart dependencies in Kind smoke tests
  • helm: default router to image-embedded model configs
  • helm: enable image pre-pull DaemonSet by default
  • helm: port health gates from Terraform to Helm post-install hooks
  • helm: remove prometheus alias, bump to v0.2.0, standardize chart
  • infra: add Modal GPU sandbox for remote benchmark execution
  • infra: add rollout warning and explicit image_type for GCFS
  • infra: enable GCFS image streaming on GPU node pools
  • infra: set min_node_count=1 on L4 spot GPU node pools
  • integrations: add Qdrant integration
  • integrations: add Qdrant integration with native sparse vector support
  • integrations: add Weaviate v4 integration with Go module spec
  • multiprocess loadtest + SDK aiohttp migration
  • sdk: add version negotiation headers between SDK and server
  • sdk: default wait_for_capacity=True and timeout=900s
  • sdk: version negotiation header (SDK ↔ server)
  • sie-bench: add dataset/input_type fields for mTEB corpus inputs
  • sie-bench: built-in multiprocess loadtest mode
  • skills: add eval-model skill for HF model assessment
  • skills: add eval-model skill for HF model integration assessment
  • sync Terraform modules to registry repos
  • tei-runner: add /embed_sparse support for sparse models
  • tei: add /embed_sparse support and auto-detect pooling mode
  • terraform/aws: restore cluster autoscaler helm release to infra module
  • terraform/aws: strip k8s resources, restructure as infra module with examples
  • terraform: add cluster name and artifact registry variables; update node pool configuration
  • terraform: add EBS CSI driver, NVIDIA device plugin, default StorageClass
  • terraform: strip gcp k8s/ layer; examples use infra-only module
  • tools: add ColBERT query vs document profiling script
  • tools: add dense P50 latency profiling script
  • adapters: sort IDF unique_ids to satisfy SparseVector contract
  • add missing production example to tf validate; fix tempfile leak; remove module docstring
  • address PR #478 review feedback
  • address PR review — GPU alert formula, kubectl parsing, CI path filter
  • address review feedback for npm publish
  • address review findings - race prevention, cleanup, lighter checkout
  • alloy: add stage.cri{} before stage.json to unwrap CRI log envelopes
  • alloy: explicitly set configMap name and key for sub-chart wiring
  • alloy: scope pod discovery to current node via field selector
  • bench: complete g5 to g6 migration in AWS eval configs and GPU mapping
  • benchmarks: use TEI /embed_all for ColBERT multi-vector models
  • bench: skip loading candidates_model for single-model servers
  • chart: update home URL and Helm install command in README
  • CI compatibility and consistent env var usage
  • CI compatibility for sync-terraform workflow
  • ci: add contents: read permission to publish-pypi-oidc job
  • ci: add helm repo add + dep build to kind-smoke workflow
  • ci: g5 refactored to g6 already
  • cluster: build concrete helm command in status from infra_outputs
  • cluster: guard helm/kubectl post-create log when outputs are empty
  • deploy: clean terraform init artifacts before push
  • deploy: correct smoke test TypedDict access and helm dry-run args
  • deploy: correct StatefulSet rollout semantics, PDB scope, and KEDA pause
  • deploy: remove dangling kubernetes_namespace_v1.sie references from health_gates.tf
  • deploy: restore triggering ref for source checkout
  • deploy: update default destination repos for GCP and AWS modules
  • deploy: use triggering ref for source checkout in sync-terraform
  • disable LoRA adapter layers after loading to prevent quality corruption
  • docs: clarify optional image push in AWS and GCP README files
  • docs: update Helm chart path in AWS and GCP README files
  • fix integration test
  • helm,hook: deploy/helm/sie-cluster/templates/hooks/prometheus-ready-test.yaml
  • helm: add before-hook-creation to Job delete policies; document count==0 expectation
  • helm: address coderabbit findings on health gate hooks
  • helm: address non-blocking review findings from PR #336
  • helm: address review findings in batteries-included sub-chart PR
  • helm: address reviewer suggestions for health gate hooks
  • helm: address second-pass review findings
  • helm: aggregate buckets by le in p95 latency alert
  • helm: bump chart version to 0.1.1 (patch, not minor)
  • helm: clarify prometheusAddress comment — ignored when sub-chart is installed
  • helm: correct kube-prometheus-stack semver constraint and remove hardcoded grafana password
  • helm: correct misleading validation comment in router-deployment.yaml
  • helm: don’t emit ScaledObject CRDs unless KEDA is confirmed present
  • helm: downscope KEDA RBAC to Role/RoleBinding; remove runtime apk installs
  • helm: fix loki service URL and extract alloy config to file
  • helm: fix three blocking review issues in sub-chart dependencies
  • helm: improve temporary values file handling in helm_template function
  • helm: improve, simplify, and modularize sie-cluster chart
  • helm: move ‘app.kubernetes.io/part-of’ label to selector labels for consistency
  • helm: remove autoscaling.enabled from values-aws.yaml
  • helm: render KEDA ScaledObjects via post-install hook to avoid CRD chicken-and-egg
  • helm: replace hardcoded namespace in provisioning alert rules
  • helm: require non-empty hfToken.value when hfToken.create is true
  • helm: sub-chart naming, Loki compactor, event exporter ECR, Grafana folders
  • helm: use autoscaling.prometheusAddress in prometheus hook; remove stub health_gates.tf
  • helm: use full FQDN for Prometheus service in KEDA and health gates
  • helm: use router.service.port in NOTES.txt instead of hardcoded 8080
  • infra: update min_node_count default in top-level GCP module
  • normalize SDK version warned-set key to major.minor
  • pool error types, add pool/progress test coverage
  • profiling: add flash variant registry, device validation, top-level import
  • profiling: sync GPU before tensor timing, move script to tools/
  • profiling: use in-place relu_ to match production code path
  • qwen3: restore quality by enabling causal attention and QK-normalization
  • readme: correct helm chart path
  • readme: correct helm install command
  • release: track all package versions via release-please extra-files
  • release: track TS SDK version.ts via release-please
  • replace corrupted bge-m3 NanoFiQA2018 target + set bfloat16 precision
  • review items
  • router: increase pool lease TTL to survive rolling upgrades
  • router: resolve default pool GPU for scale-up when gpu/pool omitted
  • router: use effective_pool instead of pool_name for default pool GPU extraction
  • sdk: defer aiohttp session creation to fix “no running event loop” in SIEAsyncClient
  • sie_bench: improve —print-gap report accuracy and readability
  • tei-runner: validate /embed_all returns per-token embeddings
  • tei-runner: validate output_type in TEIRunner init
  • terraform/aws: add full -backend-config flags to production init command
  • terraform/aws: add precondition asserting >=2 GPU-capable AZs exist
  • terraform/aws: address review findings post-restructure
  • terraform/aws: correct helm chart path in dev-g5-spot example comment
  • terraform/aws: filter VPC AZs to only zones offering the GPU instance type
  • terraform/aws: fix invalid splat on instance type offerings locations
  • terraform/aws: remove provider aws block from child module
  • terraform/aws: use var.project_name in VPC subnet cluster tags
  • terraform: add validation for GPU node pool zones to ensure they match the configured region
  • terraform: restore dev-l4-spot zones to us-central1 for GPU availability
  • terraform: update GPU instance type description for clarity and add dev-g6-spot example
  • terraform: update stale k8s module references in comments
  • terraform: upgrade AWS modules and fix deprecations
  • test: check /metrics endpoint in test_prometheus_metrics_exist
  • test: update EKS tests from g5 to g6 after GPU instance type change
  • ts-sdk: add per-attempt timeout to lease renewal fetch
  • use prepack instead of prepublishOnly
  • validate minimum_worker_count input and soften docstrings
  • adapter: optimize MoE expert dispatch with sorted-expert routing
  • adapters: batch MaxSim scoring across documents on GPU
  • adapters: batch sparse aggregation with segment_reduce and fuse relu
  • adapters: batch split_embeddings + validate ColBERT performance
  • adapters: batch split_embeddings in ColBERT adapters
  • adapters: eliminate GPU overhead from IDF query encode path
  • florence2: greedy decoding for OCR (-23% P50)
  • florence2: switch OCR configs from beam search to greedy decoding
  • server,bench: add batch coalescing, query warmup, and benchmark stability improvements
  • server: dispatch immediately when worker is idle
  • server: optimize BertFlashAdapter inference path (+35% corpus throughput)
  • server: reduce batch wait timeout 10ms \u2192 2ms for lower Doc P50
  • Breaking change: remove florence2 and gliner standalone bundles — extraction adapters (gliner, glirel, gliclass) are now included in the default bundle
  • New capabilities: add native MTEB reranking task support with MRR metric; encode-dense matrix eval — 3 models × 8 tasks; add date-prefixed versioning for chronological filename ordering; reorganize and expand model size lookup table with alphabetical ordering; add detailed perf metrics, metric filter, and threshold selector; add marimo benchmark dashboard notebook
  • Reliability and operations: restore /var/cache/apt mounts, keep /var/lib/apt removed; add HF_TOKEN auth and config kwargs and fix stella models; add dense projection support to Qwen2FlashAdapter; apply query_template from runtime options in SentenceTransformerDenseAdapter; replace pip install with uv add in docker error messages
  • Performance: vectorize GTE sparse encode path; vectorize tokenization and packing for gte-multilingual-base; vectorize tokenization and packing to reduce throughput gap; switch Qwen/GTE models to flash attention adapter
  • bundles: remove florence2 and gliner standalone bundles — extraction adapters (gliner, glirel, gliclass) are now included in the default bundle
  • bench: add native MTEB reranking task support with MRR metric
  • bench: encode-dense matrix eval — 3 models × 8 tasks
  • benchmarks: add date-prefixed versioning for chronological filename ordering
  • benchmarks: reorganize and expand model size lookup table with alphabetical ordering
  • benchview: add detailed perf metrics, metric filter, and threshold selector
  • benchview: add marimo benchmark dashboard notebook
  • benchview: add perf metric selector to Model Size tab
  • benchview: detailed perf metrics, metric filter, threshold selector
  • router,bench,sdk: improve throughput with inflight tracking, batching, and connection pooling
  • server: typed request parsing with msgspec
  • sie_server: add gliner, glirel, and gliclass extraction dependencies to the default bundle
  • adapter: add dense projection support to Qwen2FlashAdapter
  • adapter: apply query_template from runtime options in SentenceTransformerDenseAdapter
  • apply CodeRabbit auto-fixes
  • bench: replace pip install with uv add in docker error messages
  • benchview: add missing statistics import and use _median helper
  • bundles: include sglang bundle in default cluster and eval-matrix configs
  • client: update websocket header parameter name from extra_headers to additional_headers
  • colbert: enable native mode fallback for non-CUDA devices and add Matryoshka truncation
  • deps: cap timm upper bound and fix lazy handler init
  • docker: clear stale apt lists before update to prevent 404s
  • docker: remove all apt cache mounts from Dockerfiles
  • docker: remove no-op /var/lib/apt cache mount from apt RUN blocks
  • docker: restore /var/cache/apt mounts, keep /var/lib/apt removed
  • helm: increase CPU worker pool memory limits for expanded default bundle
  • model: add missing query_template to stella_en_400M_v5
  • models: switch all-MiniLM-L6-v2 to SentenceTransformerDenseAdapter
  • multilingual-e5-large-instruct: use instruct query template, NFCorpus 0.3521 → 0.3567
  • replace invalid HTML entities in SVG with XML numeric entities
  • rope_flash: clear cached _rope_dummy on unload and use torch.cat for packing
  • router: also resolve pool-derived GPU names to spot variants
  • router: resolve bare GPU types to spot variants for KEDA scaling
  • sdk: resolve sync/async client inconsistencies in score() and encode()
  • server: centralize request validation to prevent 500s from malformed items
  • server: use BertFlashAdapter for e5-small-v2, resolve e5 perf anomalies
  • server: use BertFlashAdapter for intfloat/e5-small-v2 and remove stale benchmarks
  • set compute_precision to bfloat16 for stella_en_1.5B_v5
  • sie_server: add HF_TOKEN auth and config kwargs and fix stella models
  • sie_server: resolve BGE-M3 linear weights loading for HF model IDs and fix test fixtures
  • sie_server: support NV-Embed-v2 with PyTorch embedding adapter
  • splade: align special token filtering and guard empty batches
  • test: always rebuild Docker images to pick up code changes
  • typecheck: move ty type checker from mise tool to uv dependency
  • adapters: vectorize GTE sparse encode path
  • rope_flash: vectorize tokenization and packing for gte-multilingual-base
  • rope_flash: vectorize tokenization and packing to reduce throughput gap
  • server: switch Qwen/GTE models to flash attention adapter
  • splade: vectorize tokenization and sparse aggregation (1.5x throughput)
  • splade: vectorize tokenization and sparse aggregation in SPLADEFlashAdapter
  • New capabilities: add GLiNER v2.5 model configs; stream request bodies through proxy instead of buffering; add classification model configs for GLiClass-large and cross-encoder NLI
  • Reliability and operations: release pipeline cache collision and smoke test timeout; strip content-length header from streamed proxy responses
  • Performance: stream response body to eliminate bytes.join bottleneck
  • models: add GLiNER v2.5 model configs
  • router: stream request bodies through proxy instead of buffering
  • sie_server: add classification model configs for GLiClass-large and cross-encoder NLI
  • release pipeline cache collision and smoke test timeout
  • router: strip content-length header from streamed proxy responses
  • router: stream response body to eliminate bytes.join bottleneck
  • Reliability and operations: revert sharing=locked, add cache-read-only for build step
  • revert sharing=locked, add cache-read-only for build step
  • Reliability and operations: revert token lifetime extension, re-auth before push instead; revert token lifetime, re-auth before push
  • revert token lifetime extension, re-auth before push instead
  • revert token lifetime, re-auth before push
  • Reliability and operations: release image builds failing from GCP token expiry
  • release image builds failing from GCP token expiry
  • Reliability and operations: update bundle definitions to replace legacy and gte-qwen2 with gliner
  • bundles: update bundle definitions to replace legacy and gte-qwen2 with gliner
  • Breaking change: HTTP 409 dependency conflict responses are removed from all API endpoints; the DEPENDENCY_CONFLICT error code no longer exists; .beads/ issue tracking data removed from repository
  • New capabilities: add X-SIE-Worker response header for per-worker metrics tracking; add encode-image-text measurements to benchmarks dir; add encode-multivector perf measurements; add encode-multivector performance measurements; add encode-visual-document perf measurements; add encode-visual-document performance measurements
  • Reliability and operations: increase helm install timeout from 10m to 15m; add trailing empty line to gitignore; align release-images workflow with docker task flags; register GLiClass and DeBERTa models in bundles; build and deploy gliner bundle in Kind smoke tests
  • Performance: add connection pooling load test results (Feb 24); pool httpx client and add X-SIE-Worker header in router proxy; pool httpx client in router proxy to eliminate per-request TCP overhead; move transformers imports to module level
  • deps: HTTP 409 dependency conflict responses are removed from all API endpoints; the DEPENDENCY_CONFLICT error code no longer exists
  • .beads/ issue tracking data removed from repository
  • deps: model config files no longer support the dependencies field
  • add X-SIE-Worker response header for per-worker metrics tracking
  • benchmarks: add encode-image-text measurements to benchmarks dir
  • benchmarks: add encode-multivector perf measurements
  • benchmarks: add encode-multivector performance measurements
  • benchmarks: add encode-visual-document perf measurements
  • benchmarks: add encode-visual-document performance measurements
  • benchmarks: add extract-detection L4-SPOT performance measurements
  • benchmarks: add extract-kie-docvqa measurements to benchmarks dir
  • benchmarks: add extract-relation L4-SPOT performance measurement
  • benchmarks: add score-colbert perf measurements
  • benchmarks: add score-colbert performance measurements
  • models: add encode-image-text measurements
  • models: add extract-detection measurements
  • models: add extract-kie-docvqa measurements
  • models: add extract-relation measurements
  • router: add structured audit logging for API requests
  • .claude: add trailing empty line to gitignore
  • align release-images workflow with docker task flags
  • bundles: register GLiClass and DeBERTa models in bundles
  • ci: build and deploy gliner bundle in Kind smoke tests
  • colbert: remove CUDA requirement and improve device compatibility
  • eval: read ‘sie_id’ instead of ‘name’ from model configs in runner
  • extract: use dict access for Entity TypedDict in sort
  • gliner: relax stale transformers<4.52 pin
  • increase helm install timeout from 10m to 15m
  • reduce cpu-gliner resource requests for Kind CI
  • router: read ‘sie_id’ instead of ‘name’ from model configs
  • server: migrate NLI adapter to classifications and improve API consistency
  • server: migrate nli_classification adapter and improve type annotations
  • server: populate classifications instead of entities in GLiClass adapter
  • use manifest mode for release-please and reset to v0.0.0
  • use nested .gitignore for .claude/ directory
  • add connection pooling load test results (Feb 24)
  • pool httpx client and add X-SIE-Worker header in router proxy
  • pool httpx client in router proxy to eliminate per-request TCP overhead
  • pytorch-embedding: move transformers imports to module level
  • server: use uvloop as default event loop for uvicorn
  • keep CONTRIBUTING.md clone URLs pointing to sie.git
  • remove beads, agent prompts, mypy refs; consolidate ty config
  • deps: move adapter dependencies from per-adapter pyproject.toml to bundle YAML
  • deps: remove model-level dependencies feature

Contact us

Tell us about your use case and we'll get back to you shortly.