Skip to main content

11 posts tagged with "s3"

Amazon S3 storage service topics and usage

View All Tags

Spice v2.4.0-rc.1 (Oct 8, 2026)

ยท 64 min read
Sergei Grebnov
Member of Technical Staff at Spice AI

Spice v2.4.0-rc.1 is now available! ๐Ÿ”ฅ

Spice v2.4.0-rc.1 is the first release candidate for v2.4.0. It adds performance improvements, S3 event-driven ingestion, SQL results-cache warmup, and adaptive HTTP rate controls. The release also upgrades to DataFusion v55, Ballista v55, Arrow v59, Vortex v0.86, Iceberg v0.11, and Turso v0.81.

Highlights in v2.4.0-rc.1 include:

What's New in v2.4.0-rc.1โ€‹

Performance & Query Engineโ€‹

This release upgrades Apache DataFusion to the v55.2.0 dependency line and Apache Arrow to v59.3.0. It also upgrades Vortex to v0.86.1, Apache Iceberg to the v0.11.0 fork, and Apache Ballista to v55.

Apache DataFusion v55โ€‹

The DataFusion v55 release adds the following improvements:

  • Sort pushdown and TopK pruning: Parquet scans reevaluate each unread row group as the threshold for ORDER BY ... LIMIT tightens. They skip groups that cannot contribute to the result. TopK pruning also supports multiple sort columns.
  • Join planning: The optimizer converts eligible inner joins to semi joins and removes redundant sides of outer joins. It also orders filter predicates by estimated cost.
  • Aggregation and expressions: Multi-column GROUP BY uses column-oriented storage for all supported key types, such as fixed-size binary UUIDs. More string functions preserve dictionary encoding, and IN lists use specialized paths for small integer types.
  • Parquet reads: Scans skip nested fields that the declared schema does not contain. They also skip page-index reads when a file has no page index.
  • Spill handling: Sorts bound the number of streams in a merge. If memory is insufficient, they spill the largest stream again in smaller batches.
  • SQL diagnostics and functions: EXPLAIN accepts PostgreSQL-style options and FORMAT pgjson. New array functions cover element-wise addition, subtraction, scaling, sums, and averages.

Spice carries these changes through its query plans and preserves statistics across plan wrappers. Cayenne keeps Vortex scans below 10 MiB unsplit to avoid repeated footer reads. See #14612.

Vortex v0.79.0 to v0.86.1โ€‹

Spice v2.3.2 used the Vortex v0.79.0 fork. This upgrade covers the full upstream range from v0.79.0 through v0.86.1, not only the v0.86 changes:

  • Types and arithmetic: Vortex adds native Map arrays, Arrow map conversion, and operations and compression for maps. It also adds union arrays and decimal addition, subtraction, multiplication, and division.
  • Row selection: Piecewise-sequence indices represent contiguous selections without an expanded index for every row. Specialized paths handle chunked arrays, fixed-size lists, variable-length lists, and binary values.
  • Scans and expressions: Layout scans gain a physical plan and expression optimization. Filters can pass through scalar functions with multiple arguments. Filters on wide lists restrict child elements to the selected range, and constant masks can resolve from metadata.
  • Compression: Binary arrays support FSST compression with variable-length offsets. OnPair becomes a stable encoding for reads and gains a storage-backed dictionary. Vortex can convert run-end arrays of lists and decimals.
  • File metadata: Files can store custom metadata. Readers cache decoded type descriptions, and file-format editions define the supported encodings and types.
  • Memory and execution: Builders append nested values in batches, preserve list views, and propagate buffer allocators. Expression rewrites retain unchanged nodes, and row functions support batch execution.
  • Correctness and validation: NULL handling changes cover dictionary predicates, BETWEEN bounds, and empty arrays. Readers add validation for footer offsets, compression metadata, and array indices. Other corrections cover nested scalar hashes, decimal operations, and variable-length binary selections above 4 GiB.

See the complete Vortex v0.79.0 to v0.86.1 changelog for every upstream change. Changes to standalone Vortex bindings and GPU execution do not imply new Spice features.

Apache Iceberg v0.11.0โ€‹

Spice updates its Iceberg reader and catalog integrations from v0.10.1 to the v0.11.0 fork for DataFusion 55 and Arrow 59. The upstream v0.11 changes extend reads and catalog compatibility:

  • Iceberg v3 reads: The reader applies deletion vectors from Puffin files and carries row identifiers and sequence numbers through scans.
  • Metadata-only scans: Metadata-only projections do not require data-column reads. Manifest reads reuse partition types, and positional-delete processing buffers runs instead of allocating a key per row.
  • REST catalogs: Clients negotiate server-advertised endpoints and use session-scoped OAuth2 authentication.

The upgrade also aligns Iceberg storage with OpenDAL v0.58. See #14771.

Apache Ballista v55โ€‹

Spice.ai Enterprise feature. See the Enterprise documentation.

The Ballista v55 upgrade adds virtual-core resource accounting for distributed tasks. A protocol-version handshake detects incompatible schedulers and executors. Cancellation identifies tasks by their task IDs, and task-state records use an append-only model.

Schedulers and executors must run the same version. Upgrade all cluster components together.

S3 Event-Driven Ingestionโ€‹

S3 listing datasets can use refresh_mode: changes with S3 event notifications delivered through SQS:

datasets:
- from: s3://my-bucket/events/
name: events
params:
file_format: parquet
s3_region: us-east-1
s3_auth: iam_role
s3_changes_queue_url: ${ secrets:events_queue_url }
acceleration:
enabled: true
engine: cayenne
mode: file
refresh_mode: changes

New object notifications append the object's rows. A periodic listing backfill covers missed or expired notifications. Each dataset needs its own queue and permissions to read and delete SQS messages, alongside its S3 read/list permissions.

This is object ingestion: removal notifications are ignored by default. Set s3_on_object_removed: rebuild to rebuild the entire prefix when an object is removed. An overwrite of an already applied object key is not ingested again; use new object keys for incoming data. See #14121.

SQL Results-Cache Warmup and Shared Fetchesโ€‹

The SQL results cache can persist query plan shapes and replay them after a dataset's first full or append refresh:

runtime:
caching:
sql_results:
enabled: true
warmup: on_first_refresh

For example, run these queries against an accelerated orders dataset with warmup enabled:

SELECT id, status FROM orders WHERE id = 1;
SELECT id, status FROM orders WHERE id = 2;

Spice records one query shape because only the equality-filter value differs. After a restart and the dataset's first full or append refresh, warmup reruns that shape with distinct id values from the refreshed dataset, filling the cache before the dataset becomes ready. Keep the local .spice/data directory across restarts, or configure runtime.state.location, to retain recorded query shapes.

Warmup replays up to ten distinct recorded query shapes and tries up to 1,024 distinct filter-value combinations per shape, stopping when the cache is full. A dataset stays not ready until warmup completes. Later refreshes do not repeat warmup. The feature requires the default plan-based cache key; cache_key_type: sql is incompatible. See #14178.

Concurrent cache-miss fetches for the same request now share a source fetch. SQL results caching also includes changes for tables updated during query execution and for stale results served while revalidation runs. HTTP dataset caching supports RFC 5861 stale-if-error handling. See #14142, #14710, #14708, and #14134.

SQL, search, and embedding caches now use Spice's sharded cache backend. Existing engine: moka and engine: pingora values are accepted for configuration compatibility but no longer select a backend.

Adaptive HTTP Rate Controlsโ€‹

HTTP rate controls now adapt admission to upstream failures while staying within configured request limits. rate_control_acquire_timeout bounds how long a request waits for capacity and defaults to the connector's client timeout. rate_control_failure_threshold and rate_control_window control the response to upstream failures.

In Spice.ai Enterprise, instances that share a runtime.state.location coordinate per-second and per-minute limits through shared state; OSS instances keep these limits in memory independently. Concurrency limits remain local to each instance. Components sharing an upstream origin must use matching rate-control settings. See HTTP rate-control documentation and #14143.

ORC Files and Object Metadata Queriesโ€‹

Listing connectors support file_format: orc. Object-store listing also uses predicates on metadata columns, including _last_modified, to narrow eligible objects. Queries that select only partition or metadata columns can use those values without reading file contents. See #14075, #14265, #14303, and #14116.

Hugging Face Datasetsโ€‹

The new Hugging Face data connector queries and accelerates datasets from the Hugging Face Hub. It supports Parquet, CSV, TSV, JSON, and ORC files. Public datasets need no credentials:

datasets:
- from: hf://datasets/stanfordnlp/imdb/plain_text/
name: imdb
acceleration:
enabled: true

The location format is hf://datasets/<owner>/<dataset>[@<revision>][/<path>]. A path can select a file, a folder, or a glob. A revision can name a branch, a tag, or a commit. Use @~parquet to select the Hub's automatic Parquet conversion.

Set hf_token for private or gated datasets. Set hf_endpoint for a Hub mirror or proxy. Each scan reads one commit. Refreshes follow the selected branch, but the dataset keeps its registered schema until reload. See #14877.

Automatic Primary-Key Handlingโ€‹

Cayenne keeps one row per primary_key without requiring an on_conflict policy. When a dataset sets time_column, the row with the newest time wins; without it, the last arrival wins. This applies to full and append refreshes as well as writes. Existing explicit conflict policies remain accepted during the deprecation period.

datasets:
- from: s3://my-bucket/orders/
name: orders
time_column: updated_at
params:
file_format: parquet
acceleration:
enabled: true
engine: cayenne
mode: file
primary_key: id

For a read-write dataset whose writes should stay in its acceleration, set acceleration.write_mode: acceleration. Its source need not support writes. This mode cannot be combined with a dataset that refreshes by changes.

Cayenne secondary indexes also support dynamic join filters, and their write handling covers inserts, updates, deletes, and refreshes. See #14726, #14282, and #14593.

PostgreSQL and MySQL replication, and MongoDB change streams, rejected Cayenne datasets that omitted on_conflict. Their validation still required an explicit upsert policy. These sources now accept Cayenne datasets with primary_key alone. See #14880.

For file-mode Cayenne datasets, append refreshes failed with a configuration that combined primary_key, time_column, and retention_sql. The refresh selected a version-resolution path that did not support retention. The refresh now resolves each key's newest version before Cayenne applies retention. See #14878.

Cayenne maintained aggregates now share one compact index of per-key contributions across views. Rebuilds capture concurrent writes and apply them after the scan. The aggregate budget uses 10% of a bounded query pool without the former 512 MiB cap. An unbounded pool retains the 512 MiB budget. See #14762.

Read Datasets from Published Snapshotsโ€‹

Spice.ai Enterprise feature. See the Enterprise documentation.

A dataset can now read published acceleration snapshots directly, without configuring the original source connector:

datasets:
- from: s3://my-bucket/spice/snapshots/orders/
name: orders
params:
file_format: snapshot
s3_region: us-east-1

Spice reads the snapshot metadata to select the engine, restores the published data, and checks for newer snapshots. The dataset is read-only. Snapshot-mode readers can also use S3 notifications delivered through SQS, with periodic checks retained for missed notifications. Each reader process needs its own queue.

This release includes changes to snapshot retries, slow-connection bootstrap, publication metadata, and coordination between snapshot archiving and Cayenne maintenance. Cayenne datasets with a datalake tier cannot create acceleration snapshots. See snapshot documentation, #14529, and #14335.

Connector and Protocol Updatesโ€‹

  • Connector status: ADBC, Databricks Spark Connect and SQL Warehouse, FlightSQL, Glue, HTTP/HTTPS, Iceberg, Localpod, and MongoDB are now Stable data connectors.
  • MCP: support for specification version 2026-07-28, alongside the earlier protocol era. See #14043.
  • GitHub: nested GraphQL pagination and rate-limit pacing updates; the default concurrency limit is now four. See #14179 and #14431.
  • GitHub nested pages: Scans could return incomplete reviews or comments because pagination accepted a short page as complete. The connector now rejects incomplete connections and repeated cursors. It retries a failed nested page without another fetch of the outer page. Datasets with the same token also share the REST quota. See #14862.
  • Iceberg REST: Clients could read an empty dataset because the catalog synthesized metadata without snapshots. The catalog now returns the source metadata for Iceberg datasets that Spice reads unchanged. Other datasets return 400 BadRequestException. Clients need their own storage credentials. See #14588 and Breaking Changes.

Other Fixesโ€‹

The release includes fixes in the following areas; the linked PRs provide details of the changes:

  • Cayenne queries and writes: NULL-aware NOT IN, maintained aggregates, dynamic filters, partition-filter forwarding, memory-mode DML and retention, and primary-key handling across CDC checkpoints. See #14429, #14761, #14370, #14047, and #14344.
  • Startup and reloads: retry datasets with unavailable sources, serve existing accelerations during source outages, and invalidate cached plans and results on catalog replacement or dataset unload. See #14623, #14624, #13914, and #14365.
  • Federation: local evaluation of casts and functions whose source semantics differ, plus filter pushdown changes for DynamoDB, Cosmos DB, and MongoDB. See #14484, #14601, and #14419.
  • HTTP and GraphQL: response-status handling during refresh, retry-budget handling, non-JSON gateway responses, and URL redaction in HTTP errors. See #13538, #14313, #14781, and #14490.
  • Search: deletion of obsolete Elasticsearch chunks, non-finite embedding handling, and source-scan coordination during full-text refresh. See #13960, #13902, and #14663.
  • Models and tools: tool-call-only assistant turns, required tool choices, streaming tool-use completion, model-load diagnostics, and propagation of the API-key principal into MCP tool calls. See #14232, #14460, #14548, and #14828.
  • CDC shutdown and reconnects: source-position recording before accelerations close, and MySQL shared-stream reconnect handling. See #14702 and #14751.
  • SQL weekdays: date_part('dow') aligns with EXTRACT(dow), with Sunday represented as zero. See #14796.
  • Cayenne schema statistics: Decimal bounds could retain an old scale after schema evolution because maintenance published statistics from the previous schema. Cayenne now rejects statistics from an obsolete schema and keeps row counts conservative. See #14856.
  • Vector search: vector_search planning failed after the DataFusion 55 upgrade because a second optimization pass tried to reorder a join with a dynamic filter. The planner now preserves that join's input order. See #14857.

Default accelerator: Datasets and views that enable acceleration without engine now use Cayenne. Explicit engine settings keep their behavior. Storage still defaults to memory. Set mode: file for persistent acceleration. See Breaking Changes for migration guidance.

Dependency Updatesโ€‹

Dependency / ComponentVersion
DataFusionv55.2.0
Apache Arrowv59.3.0
Vortexv0.86.1
Apache Icebergv0.11.0
Apache Ballistav55.0.0
Tursov0.8.1
ADBCv0.24
Rust toolchainv1.98.1

Contributorsโ€‹

Breaking Changesโ€‹

Cayenne is the default accelerator on supported platforms. A dataset or view that omits acceleration.engine switches from Arrow to Cayenne. To retain Arrow, set engine: arrow explicitly before upgrading. Windows keeps Arrow as its default. Persistent datasets should continue to name their engine and use mode: file.

on_conflict is deprecated and scheduled for removal in v3.0. Cayenne automatically keeps one row per primary key, choosing the newest time_column value when configured, or the last arrival otherwise. Existing explicit policies remain supported during the deprecation period. Review those policies before removing them, especially drop or policies that reject conflicting rows.

on_conflict no longer routes writes to the acceleration. For read-write datasets whose writes should stay in the acceleration, use:

acceleration:
enabled: true
engine: cayenne
write_mode: acceleration

This setting cannot be used with refresh_mode: changes, including a connector's default change-stream mode. The default write_through and write_back modes require a writable source.

Cache engine selection is retired. engine: moka and engine: pingora remain accepted but are ignored. Remove the field and use caching_policy to select eviction behavior.

GitHub connector default concurrency is four. Review explicit concurrency settings if your deployment relied on the previous default.

HTTP rate-control waits are bounded by default. Requests waiting for rate-control capacity now time out after the connector's client timeout. Set rate_control_acquire_timeout to a suitable duration, or 0 to retain the previous unbounded wait behavior.

Cayenne acceleration snapshots are unavailable for datalake-tier datasets. Review snapshot settings on datasets using cayenne_datalake_location; this release disables snapshotting that configuration.

Iceberg REST no longer synthesizes metadata for unsupported datasets. GET /v1/namespaces/{namespace}/tables/{table} returns 400 BadRequestException for accelerated datasets, views, and other datasets that Spice does not read unchanged from Iceberg. If a client used this endpoint for schema discovery, use SQL DESCRIBE or information_schema.columns instead. Query these datasets through /v1/sql or Arrow Flight SQL. For eligible Iceberg datasets, clients read the source metadata and need their own storage access. See the Get a table API.

Cookbook Updatesโ€‹

The Spice Cookbook provides recipes to help you get started with Spice.

Upgradingโ€‹

To upgrade to v2.4.0-rc.1 once the release artifacts are available, use one of the following methods:

CLI:

spice upgrade v2.4.0-rc.1

Docker:

Pull the spiceai/spiceai:2.4.0-rc.1 image:

docker pull spiceai/spiceai:2.4.0-rc.1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 2.4.0-rc.1

AWS Marketplace:

Spice is available in the AWS Marketplace. Marketplace availability follows its published versions.

What's Changedโ€‹

Changelogโ€‹

  • fix(runtime): discard cached logical plans when a hot reload replaces a catalog (fixes #13910) by @claudespice in #13914
  • fix(acceleration): let a schema repair correct a checkpoint without resetting the freshness clock (fixes #13817) by @claudespice in #13894
  • fix(search): filter a chunked Elasticsearch delete on a field that can match the key (fixes #13714) by @claudespice in #13926
  • fix(search): classify a partially non-finite embedding as unindexable on every backend (fixes #13872) by @claudespice in #13902
  • fix: stabilize GitHub tests and bound GraphQL registration (fixes #13762) by @lukekim in #13939
  • fix(postgres): decode versioned JSONB binary replication values by @phillipleblanc in #13962
  • docs: require a reviewed Enhancement before any user-facing surface changes by @lukekim in #13970
  • ci: upgrade spiceio setup action to v0.9.0 by @lukekim in #13971
  • fix(postgres): preserve microseconds in timestamp writeback by @phillipleblanc in #13963
  • feat(hash-index): verify the bloom filter's block index with Verus by @lukekim in #13777
  • build(deps-dev): bump js-yaml by @dependabot in #13989
  • docs: release notes for v2.3.0 by @bjchambers in #13999
  • fix(ci): drop the dangling substrait-compliance submodule pointer by @bjchambers in #14002
  • fix(cayenne): release the keyset bytes an abandoned PK checkout accounted (fixes #13668) by @grokspice in #13925
  • fix(caching): keep a declared key from disabling eviction and stranding stale rows (fixes #13976) by @bjchambers in #13992
  • Add Substrait compliance harness (IBM TPC-H Mode A + FlightSQL Mode B stub) by @lukekim in #13879
  • ci: skip DynamoDB TPC-H benches in OSS testoperator dispatch by @phillipleblanc in #14016
  • docs: update security support and roadmap after v2.3.0 by @phillipleblanc in #14024
  • chore: post v2.3.0 release housekeeping by @bjchambers in #13969
  • test(adbc): guard BigQuery corpus offline and in release gate by @phillipleblanc in #14017
  • fix(duckdb): deny the regexp built-ins DuckDB cannot answer faithfully (fixes #13809) by @claudespice in #13871
  • fix: Update tpch benchmark snapshots for federated/adbc[bigquery].yaml by @app/github-actions in #13984
  • fix(search): drop the chunks a shortened row no longer produces from a chunked index (refs #13717) by @claudespice in #13960
  • Reduce Cayenne allocations during primary-key validation and filtering by @lukekim in #14009
  • test(forks): guard seven fork patches that had no repo-side test by @krinart in #13996
  • fix(deps): bump arrow-rs to correctly-rounded Decimalโ†’Float cast (closes #13978) by @Jeadie in #14012
  • perf(vortex): defer projection setup on filtered scans until the filter resolves by @bjchambers in #14035
  • endgame: include spiceai/skills versioned release by @lukekim in #14031
  • fix(caching): partition doomed entries at the survivor cutoff so eviction converges (closes #13994) by @Jeadie in #14021
  • fix(deps): bump arrow-rs fork pin for Decimal->Float rounding fix by @Jeadie in #14049
  • Fix subqueries with use_source acceleration by @phillipleblanc in #14022
  • fix(cayenne): make DELETE, UPDATE and INSERT work on a mode: memory acceleration (fixes #12008) by @bjchambers in #14047
  • fix: clarify OpenDAL S3 retry warnings by @lukekim in #14040
  • test(s3): run the parquet-overwrite fixtures on RustFS by @bjchambers in #14067
  • feat(cayenne): materialize multi-reference CTEs on the query path by @lukekim in #13918
  • perf(vortex): answer a constant IN list by probing a set, and falsify it by interval by @bjchambers in #14061
  • perf(vortex): skip a scan split whose zones cannot satisfy the filter by @peasee in #14064
  • fix(arrow): report an exact row count from the indexed point-lookup scan by @krinart in #13972
  • fix(cayenne): apply sort_columns with refresh_mode: full by @peasee in #14063
  • fix(smb): pad an empty CREATE buffer so Samba lists the share root (fixes #13293) by @grokspice in #14050
  • perf(cache): key the logical-plan cache on SQL text, not parameter values by @bjchambers in #14069
  • fix: harden HuggingFace E2E chat against slow Metal generation by @lukekim in #14072
  • fix(cayenne): move accelerator filesystem I/O off Tokio workers by @lukekim in #14073
  • test(chbench): enable CTE materialization and IVM on mysql/postgres adaptive HTAP by @lukekim in #14070
  • ci: run Substrait Mode A TPC-H on pull requests and the merge queue by @lukekim in #14071
  • feat(mcp): support MCP specification 2026-07-28 (dual-era) by @lukekim in #14043
  • fix(ci): call a linker that died of a signal an infrastructure failure, not a check failure (fixes #13614) by @grokspice in #14044
  • fix: Provide temporary directory in docker images by @Jeadie in #14089
  • fix: restore OSS installer, CLI and test workflow coverage by @phillipleblanc in #14025
  • Delete v2.2.0.md by @Jeadie in #14095
  • docs(release): add v2.3.1 release notes by @phillipleblanc in #14094
  • fix(test): allow DELETE in the CORS allow-methods assertion by @claudespice in #14098
  • fix(ci): stop install-protoc unzipping into a shared ~/.local by @lukekim in #14097
  • feat(connectors): add ORC listing format via in-repo FileFormat by @lukekim in #14075
  • fix(cayenne): run snapshot bootstrap check before opening the metastore by @Jeadie in #14093
  • feat(cache): verify the results-cache namespace prefix with Verus by @lukekim in #14074
  • feat(cloud-connect): add a GetDatasets command that answers the /v1/datasets document (refs #13369) by @grokspice in #14051
  • Suppress Cayenne startup logs when no Cayenne dataset is configured by @Jeadie in #14042
  • fix(cache): re-bind parameter values when revalidating a stale result (fixes #14099) by @bjchambers in #14100
  • fix(cluster): support distributed HTTP scans by @phillipleblanc in #14108
  • fix(bigquery): keep ILIKE evaluation local by @phillipleblanc in #14110
  • docs: update security support for v2.3.1 by @phillipleblanc in #14117
  • feat(cayenne): reuse ScanView until write, lag only for read-only CDC by @lukekim in #14055
  • fix(deps): remediate open Dependabot alerts by @phillipleblanc in #14111
  • fix(testoperator): validate results in every scale factor 1 TPC-H, TPC-DS and ClickBench benchmark by @lukekim in #14119
  • fix(cayenne): reject ambiguous metastore paths by @phillipleblanc in #14130
  • perf: serve results-cache hits where the request arrives and cut per-hit overhead by @lukekim in #14103
  • test(runtime): record query previews in the management export test by @lukekim in #14155
  • ci: upgrade spiceio setup action to v0.11.0 by @lukekim in #14152
  • feat(caching): Make caching_stale_if_error RFC-5861 compliant (with stale-if-error header) by @Jeadie in #14134
  • perf(runtime-table): defer cache-eviction key extraction to entries a delete actually names by @Jeadie in #14138
  • fix(runtime): report the acceleration.ready_state deprecation once per component (fixes #13749) by @claudespice in #14006
  • fix(runtime): write the inferred Arrow sort order under the prefixed key its validation accepts (fixes #14023) by @claudespice in #14032
  • fix(connectors): Fix JSON/Orca files using metadata columns by @Jeadie in #14115
  • feat(cayenne): build secondary indexes from indexes in file and memory mode by @phillipleblanc in #14149
  • fix(cayenne): round-trip decimal, binary, and time stats and drop them on scale change by @lukekim in #14139
  • feat(cayenne): cluster warm and datalake tiers, and write full refreshes as key-range files by @lukekim in #14124
  • fix(cayenne): compile the cold-tier pruning test and backtick a doc literal by @lukekim in #14175
  • fix(cayenne): make the crates own targets lint and compile by @phillipleblanc in #14200
  • test(forks): guard five more fork patches, and drop a row that is not fork state by @krinart in #14015
  • perf(cache): promote encoded SQL results to raw after the second decode by @lukekim in #14199
  • fix(turso): build a dictionary column directly so a dictionary over a list, map or boolean value reads back (fixes #13033) by @grokspice in #14181
  • fix(ci): probe the macOS toolchain before reaching for brew in the release builds by @grokspice in #14203
  • fix(ci): skip Metal kernel precompilation in the macOS release build by @grokspice in #14204
  • fix(github): paginate nested GraphQL connections and pace to GitHub's rate limits by @lukekim in #14179
  • fix(vortex): stop an IN list holding a NULL from panicking the scan by @krinart in #14163
  • bench(cayenne): use std::hint::black_box in the clustering bench by @lukekim in #14129
  • fix(runtime-table): stop rebuilding SessionContext on every cache fetch by @Jeadie in #14141
  • test(chbench): cluster order_line, oorder and customer on the adaptive HTAP arms by @lukekim in #14192
  • fix(runtime): count a first load as still loading in the Dataset load summary (fixes #13974) by @claudespice in #14020
  • fix(duckdb): push regexp_count down again at a rendering that counts as the kernel does (fixes #13870) by @claudespice in #14153
  • build: lint and test the sign-off under the same profile as the merge queue by @lukekim in #14180
  • ci: require the Verus proofs in the merge queue as one check by @lukekim in #14189
  • ci: run the longest macOS jobs on their own runner pool by @lukekim in #14229
  • fix(cayenne): build the DELETE sink inside the execution-time write lock (fixes #13828) by @claudespice in #14218
  • build(deps): bump the github-actions-dependencies group across 1 directory with 7 updates by @dependabot in #14231
  • fix(cluster): decide what a Flight message carries by its IPC header, not its body length (refs #13737) by @claudespice in #14212
  • ci: run Mode A TPC-H on merge queue and trunk/release push only by @lukekim in #14247
  • chore(deps): bump spiceai/duckdb-rs to 76655d2f by @lukekim in #14246
  • ci: stop exporting empty AWS and DuckLake endpoints to the schema test by @phillipleblanc in #14194
  • fix(runtime): reload a localpod dataset when the dataset it reads through is reloaded (fixes #3288) by @claudespice in #14208
  • build(deps): bump the aws-sdk group with 3 updates by @dependabot in #14255
  • fix(ci): resolve Homebrew prefix when brew is the spice flock wrapper by @lukekim in #14210
  • chore(deps): raise the datafusion-table-providers pin to include the NUMERIC result-column fix by @phillipleblanc in #14254
  • ci: align the DuckLake bootstrap with the embedded DuckDB, wait for Databricks startup, and stop dispatching legs that cannot pass by @phillipleblanc in #14250
  • fix(runtime): install the Spice function deny-list on the PostgreSQL catalog connector (refs #13664) by @claudespice in #14225
  • perf(cayenne): share inline-cache view entries by Arc instead of cloning them per scan by @krinart in #14191
  • build(deps): bump aws-actions/configure-aws-credentials by @dependabot in #14256
  • ci: stop dispatching the indexed turso TPC-H SF1 tests by @phillipleblanc in #14252
  • fix(catalog): keep the tables registered under an existing schema by @phillipleblanc in #14193
  • feat(cache): Spice sharded cache as the sole LruCache engine by @lukekim in #14206
  • ci: lint GitHub Actions definitions with actionlint, and fix the 73 findings it surfaced by @grokspice in #14223
  • perf(cache): tighten the Raw SQL results-cache serve path by @lukekim in #14205
  • ci: stop triggering the CUDA build on pull requests by @lukekim in #14267
  • ci: run CodeQL on pull requests and the merge queue by @lukekim in #14269
  • Release 2.3.2 release notes by @krinart in #14271
  • chore: make AGENTS.md the canonical agent instructions by @lukekim in #14237
  • ci: run CodeQL Analyze on spiceai-dev-runners by @lukekim in #14281
  • feat: TypeSafe Jev System One evaluation provider by @lukekim in #14215
  • fix(cayenne): tag file statistics bounds as the column's Arrow type (fixes #14280) by @phillipleblanc in #14283
  • ci: install spiceio when the runner has no gh (refs #14233) by @lukekim in #14288
  • test(forks): 11 repo guards by @krinart in #14261
  • docs: add Spice.ai in Action manuscript and companion labs by @lukekim in #13965
  • test(duckdb): add an integration test for the index CTE materialization by @sgrebnov in #13885
  • fix(cayenne): coalesce inline writes into one batch by @sgrebnov in #14279
  • release: Update SECURITY.md and endgame template after 2.3.2 by @peasee in #14293
  • fix(cayenne): count each Arrow allocation once in the inline-cache gauge by @krinart in #14272
  • feat(s3): SQS event-driven changes for refresh_mode: changes by @lukekim in #14121
  • feat(caching): single-flight coalesce concurrent cache-miss fetches by @Jeadie in #14142
  • fix(caching): make caching_stale_if_error detect transient HTTP failures on real schemas by @krinart in #14161
  • Use Cayenne secondary indexes for dynamic join filters by @phillipleblanc in #14282
  • fix(ci): compare DynamoDB sets without relying on element order by @bjchambers in #14328
  • docs(release): Remove QA analytics step from endgame by @peasee in #14329
  • ci: run CodeQL on the merge queue and trunk, not on pull requests by @lukekim in #14290
  • docs(cayenne): reposition the reference, add query serving, re-audit against trunk by @lukekim in #14338
  • docs(endgame): update versioned docs release steps by @ewgenius in #14277
  • test(forks): guard the ballista per-task file-scan restriction by @krinart in #14292
  • fix(cli): surface the full error chain for spice chat connection failures by @krinart in #14289
  • fix(ci): degrade the incomplete-sign-off handler when the runner has no gh (refs #14234) by @claudespice in #14304
  • fix(runtime-table): serialize a direct write against acceleration snapshot creation (fixes #13548) by @claudespice in #14310
  • fix(duckdb): screen regexp_like and regexp_replace as regexp_count is screened (fixes #14148) by @claudespice in #14321
  • fix(http): honor retry budget without an extra origin request by @phillipleblanc in #14313
  • Fix SchemaCastScanExec's schema conversion in fn partition_statistics by @Jeadie in #14258
  • fix(cluster): recognise a Flight keepalive by its empty envelope, not by what its header declares (fixes #13737) by @claudespice in #14327
  • perf(http): defer zero-TTL acceleration lookup until origin failure by @phillipleblanc in #14302
  • fix(cayenne): keep a key visible when it is re-inserted over a stale-insert tombstone by @sgrebnov in #14312
  • perf(cayenne): read only key and filter columns in filtered key deletes by @sgrebnov in #14374
  • fix(cayenne): let small protected-snapshot merges run during a long large-tier merge by @sgrebnov in #14296
  • perf(cayenne): serve primary-key lookups on a freshly loaded table in ~1 ms by @lukekim in #14314
  • fix(cayenne): skip min/max statistics for nested columns (fixes #14368) by @sgrebnov in #14392
  • test(caching): cover SchemaCastScanExec statistics projection by name, retype, and SQL filter by @Jeadie in #14146
  • fix(turso): keep a quantified comparison out of Turso SQL (fixes #14041) by @grokspice in #14393
  • fix(udfs): declare the local_embed dev-dependency the embed tests need (fixes #13092) by @claudespice in #14377
  • perf(cayenne): batch metastore manifest rewrites, upsert in place, and run every write on one writer connection by @lukekim in #14369
  • fix(runtime-table): ignore zero-row batches in stale fallback by @phillipleblanc in #14331
  • Prune object-store file listing by _last_modified predicates by @Jeadie in #14265
  • Reading only partition or metadata columns needlessly scans all file contents by @Jeadie in #14116
  • fix(duckdb): keep a concat over a binary operand out of the federated plan (fixes #13915) by @claudespice in #14333
  • fix(ci): keep the sign-off attribution inside GitHub's 140-character status cap (fixes #14076) by @grokspice in #14390
  • fix(ci): expose a present-but-unlinked cc tool on macOS runners instead of routing it through brew install (fixes #13479) by @grokspice in #14387
  • ci: re-measure integration.yml's job bounds after the archive consolidation (fixes #13429) by @grokspice in #14388
  • fix(ci): start DuckLake's local MinIO from an image that is still published by @grokspice in #14399
  • test(runtime-table): make the metric-scraping refresh tests pass under cargo test by @claudespice in #14381
  • fix(federation): keep a correlated subquery predicate above a join of two sources (refs #8220) by @claudespice in #14372
  • fix(caching): snapshot staleness before the origin fetch for stale_if_error by @Jeadie in #14263
  • perf(snapshots): skip unchanged snapshot metadata with a conditional GET by @sgrebnov in #14409
  • fix(search): prune the rest of a key group from an Elasticsearch chunked index (refs #13717) by @claudespice in #14320
  • fix(bench): derive MySQL's empty-field NULL handling from the column type (refs #13152) by @claudespice in #14345
  • fix(graphql): debit a LIMIT by the rows a page returned, not the declared page size (fixes #14308) by @claudespice in #14353
  • fix(ci): fit retention_oom's retry budget inside its workflow step, and guard the coupling (fixes #13512) by @grokspice in #14389
  • docs: say plainly what Spice is, refresh the README for v2.3, and promote connector statuses by @lukekim in #14410
  • fix(ci): pin, checksum and retry the oha download in the E2E graceful-shutdown jobs by @grokspice in #14418
  • fix(snapshots): resolve snapshot entries relative to the metadata location (#14425) by @sgrebnov in #14426
  • Lower the GitHub connector default concurrency limit to 4 by @lukekim in #14431
  • build(deps): bump nvidia/cuda in the docker-dependencies group by @dependabot in #14439
  • build(deps): bump the aws-sdk group with 3 updates by @dependabot in #14440
  • build(deps): bump the github-actions-dependencies group across 1 directory with 5 updates by @dependabot in #14441
  • test(cayenne): bound the refused-build guard by builds, not by the host's speed by @grokspice in #14424
  • fix(cache): don't report an invalidation cancelled by runtime shutdown as a failure by @grokspice in #14417
  • ci: run remote sign-off on the spiceai-macos pool by @lukekim in #14449
  • ci: run CodeQL Analyze on spiceai-macos by @lukekim in #14442
  • fix(runtime): keep the built-in date_part so both weekday spellings agree (fixes #13920) by @claudespice in #14154
  • fix(cayenne): include the in-memory CDC tier when an overwrite or a retention pass covers the whole table by @lukekim in #14428
  • fix: keep NOT IN null-aware through the join reorder and the Cayenne sort-merge rewrite by @lukekim in #14429
  • fix(cayenne): discard a compaction whose snapshot an overwrite replaced mid-pass by @lukekim in #14432
  • fix(cayenne): stop maintained views, Vortex IN lists and dynamic-filter sharing from returning wrong rows by @lukekim in #14427
  • fix(cayenne): hide spilled rows on CDC upsert fallback by @bjchambers in #14416
  • fix(ci): skip the integration, ADBC, chDB and E2E gate jobs on pull requests instead of passing them (fixes #13841) by @grokspice in #14454
  • fix(cayenne): judge a filtered key delete's captured sources by the index captured with them (refs #13913) by @claudespice in #14455
  • fix(ci): clear the macos-15 image's openssl@1.1 symlink before installing MySQL by @lukekim in #14489
  • ci: produce CodeQL SARIF in the Analyze job on spiceai-macos by @lukekim in #14500
  • fix(cayenne): correctness fixes for the goal-driven adaptive controller, with a closed-loop simulation harness by @lukekim in #14443
  • fix(cdc): keep the newest source commit timestamp on a coalesced change batch by @lukekim in #14463
  • fix(cayenne): draw a sequence for a current-snapshot append so per-key OCC can order it (fixes #13685) by @claudespice in #14360
  • fix(http): name a configured model's load failure instead of reporting it not found (fixes #13303) by @claudespice in #14395
  • fix(duckdb): keep inferred source indexes off change-stream accelerations so upserts commit under concurrent reads (refs #13929) by @claudespice in #14396
  • fix(duckdb): keep a text cast over a binary operand out of the federated plan (fixes #14355) by @claudespice in #14448
  • fix(llms): honor tool_choice required and allowed_tools on mistral.rs-hosted models instead of panicking (fixes #14230) by @claudespice in #14460
  • test(runtime): run the load-error counter test in its own process (fixes #13085) by @claudespice in #14462
  • fix(runtime): stop a replaced dataset configuration's load from registering over the new one (fixes #1458) by @claudespice in #14367
  • fix(install): stop asking for sudo on a first install into a fresh HOME (fixes #14445) by @claudespice in #14495
  • perf(cayenne): serve primary-key point lookups in half the time by @lukekim in #14433
  • fix(deps): bump DataFusion for upstream fixes to wrong results from filter pushdown, simplification and planning by @lukekim in #14430
  • fix(runtime): size every internal DataFusion session from the CPU budget by @bjchambers in #14412
  • fix(runtime): serve nested, zoned and half-float columns from the Iceberg catalog API (fixes #4815) by @claudespice in #14480
  • fix(acceleration): keep cached results when a snapshot refresh finds no newer snapshot by @sgrebnov in #14497
  • fix(runtime): sleep cron tests to the next boundary, not one that already fired (fixes #13759) by @claudespice in #14474
  • fix(ci): derive every lint-rust guard make runs, whatever its separator or recipe layout (fixes #13783) by @claudespice in #14475
  • fix(cayenne): stop the small-file compaction of a position-mode PK table from deadlocking on its own write lock (fixes #14420) by @claudespice in #14481
  • fix(runtime-table): report refresh bytes for the rows each batch holds by @Jeadie in #14469
  • fix(spark,databricks): keep Spice-only functions out of SQL sent to Spark Connect and Databricks SQL Warehouse (refs #13664) by @claudespice in #14498
  • fix(vortex): size a cached footer by what it retains, not its serialized bytes (fixes #12917) by @claudespice in #14502
  • fix(cayenne,telemetry): Register accelerated sink dataset immediately from existing acceleration by @peasee in #13955
  • Update spicepod.yml by @Jeadie in #14533
  • fix(cayenne): keep a rewrite's count inexact when it retains a late protected snapshot (fixes #14383) by @claudespice in #14385
  • fix: Update tpch benchmark snapshots for accelerated/on_zero_results/file[parquet]-cayenne[file]-on_zero_results.yaml by @app/github-actions in #14408
  • build(rust): upgrade toolchain to 1.98.1 by @lukekim in #14560
  • fix(postgres): release a shared slot's hold on a table its publication cannot drop (fixes #13032) by @claudespice in #14527
  • fix(cayenne): size the build side before sort-merging a non-Cayenne outer join by @krinart in #14520
  • Update DF Upgrade template by @krinart in #14526
  • fix(ci): wait for the refresh to invalidate the results cache, not a fixed 3s by @grokspice in #14598
  • fix(deps): move the DataFusion pin past the four unparser fixes, and guard each of them (fixes #13022) by @grokspice in #14570
  • fix(dynamodb, cosmosdb, mongodb): push filters down only where the source evaluates them as SQL does by @lukekim in #14419
  • feat(snapshots): serve a dataset from published snapshots with from: s3://โ€ฆ and file_format: snapshot by @lukekim in #14529
  • perf(cayenne): keep the per-shard PK index across checkpoint flushes and back off futile bakes; fix q17 and two wrong-results races (fixes #14235) by @lukekim in #14555
  • test(chbench): lower the MySQL adaptive pods' CDC linger to 500 ms by @lukekim in #14615
  • fix(graphql): show the parse-failure bytes in JSON decode previews by @Jeadie in #14530
  • Defer source-first cache fallback planning by @phillipleblanc in #14556
  • perf(cayenne): order whole-table rewrites per scan partition, then merge by @bjchambers in #14488
  • fix(ci): tolerate the parquet-rename race's third DuckDB error in the E2E log scan by @grokspice in #14541
  • fix(graphql): retry an inferred credential refusal instead of failing the refresh by @Jeadie in #14539
  • fix(cayenne): serve a widened table's files from their persisted statistics (refs #13829) by @claudespice in #14316
  • fix(acceleration): checkpoint DuckDB's write-ahead log before a snapshot copies its file (fixes #13912) by @claudespice in #14325
  • fix(ci): let E2E cleanup run when the job failed before creating its working directory by @grokspice in #14553
  • Replace unmaintained backoff with a workspace crate by @phillipleblanc in #14621
  • test: read Expect stand-ins with the installed shell by @phillipleblanc in #14634
  • fix(models): apply a tool_choice that forces a call to one round of the tool-use loop (fixes #14459) by @claudespice in #14635
  • fix(runtime): don't end a tool-use stream on a stale tool_calls finish (fixes #13309) by @claudespice in #14548
  • fix(smb): name listed locations under the share so directory datasets load (fixes #14060) by @claudespice in #14549
  • fix(models): name a configured model's load failure in ai() and streaming nsql (fixes #14394) by @claudespice in #14632
  • test(chbench): remove the mysql-cayenne[file]-adaptive-split HTAP arm by @sgrebnov in #14620
  • fix(cayenne): record sharded CDC keys in the table-wide PK index by @lukekim in #14603
  • fix(sqlite): keep TRY_CAST and every cast SQLite answers differently out of the federated plan (fixes #14398) by @claudespice in #14496
  • ci: move remaining Ubuntu 22.04 references to 24.04 by @lukekim in #14677
  • fix(snapshots): retry a failed snapshot attempt with backoff by @sgrebnov in #14676
  • fix(snapshots): large snapshots no longer fail bootstrap on slow connection by @sgrebnov in #14571
  • fix(refresh): start a refresh's source scan only when the sink reads it, so the full-text index keeps its rows (fixes #14619) by @claudespice in #14663
  • test(cayenne): compare suite answers cell by cell against SQLite and chDB by @lukekim in #14473
  • Evaluate any chat model via /v1/evaluate by @lukekim in #14568
  • ci: quarantine the management API integration schedule until its dev OAuth client is reactivated (refs #12376) by @grokspice in #14684
  • ci: stop scheduled Testoperator Ballista benchmarks by @phillipleblanc in #14682
  • fix(cayenne): stop snapshotting datasets with a datalake tier, whose restored copies deleted each other's files by @lukekim in #14583
  • fix(cayenne): count each Arrow allocation once for the mem-tier limit and checkpoint write sizing by @krinart in #14675
  • fix(ci): classify a sign-off that reached its step budget as signalled, so an expiry publishes no verdict (fixes #13843) by @grokspice in #14685
  • fix(runtime): retry a dataset whose source is unreachable at startup (fixes #14609) by @bjchambers in #14623
  • fix(cayenne): archive only the snapshot directories a dataset snapshot references (fixes #14605) by @sgrebnov in #14627
  • ci: route trunk-push macOS builds to the standard pool by @phillipleblanc in #14645
  • ci: install strip for the retention OOM regression test by @phillipleblanc in #14701
  • fix(smb): serve every share on a host from the one store registered for it (fixes #14550) by @grokspice in #14689
  • fix(runtime): discard cached plans and results when a dataset is unloaded (refs #14251) by @claudespice in #14365
  • feat(snapshots): reload snapshot-mode datasets from S3 event notifications (SQS) by @lukekim in #14335
  • fix(cayenne): report runs per size tier when protected-snapshot compaction declines (fixes #13622) by @claudespice in #14476
  • fix(cayenne): re-run a post-write compaction pass a concurrent append asked for (fixes #13906) by @claudespice in #14479
  • fix(http): keep the request URL out of HTTP connector errors (fixes #13534) by @claudespice in #14490
  • fix(runtime): load every localpod dataset that reads from one parent at startup (fixes #13087) by @claudespice in #14532
  • fix(build): share one cargo metadata helper across the lint guards, so a broken cargo is never a violation (fixes #13121) by @claudespice in #14534
  • fix(duckdb): push regexp_replace down again for a one-digit group reference, and refresh the Postgres ClickBench q35 plan (refs #13966) by @claudespice in #14552
  • fix(duckdb): keep a cast into binary out of the federated plan (refs #14397) by @claudespice in #14633
  • fix(runtime): keep finished async query jobs finished, and delete expired ones by @lukekim in #14585
  • fix(federation): keep DataFusion's cast built-ins out of every federated plan (fixes #14444) by @claudespice in #14484
  • test(cayenne): cover a swapped join filter through the sort-merge rewrite end to end (refs #14235) by @claudespice in #14551
  • fix(runtime): decide every Spark/built-in function collision by name, and refuse an undecided one (fixes #14361) by @grokspice in #14400
  • fix(ci): run every bin target's unit tests in the sign-off gate, and give spice connect its own --cloud-region refusal (fixes #13426) by @grokspice in #14406
  • test(data_components): compile the federation unparser guards in every scoped test run (fixes #13625) by @grokspice in #14447
  • fix(runtime-component): skip an inferred Cayenne index on a floating-point column instead of failing the load (fixes #14590) by @grokspice in #14599
  • fix(cache): serve stale cached results during frequent table updates (stale_while_revalidate_ttl) by @sgrebnov in #14708
  • Update Turso to 0.8.1 and retry write conflicts a metastore statement raises by @lukekim in #14680
  • fix(snapshots): stop a snapshot dataset's load when a reload replaces it by @lukekim in #14673
  • fix(cayenne): hand non-partition filters to every partition scan (fixes #12959) by @claudespice in #14370
  • fix(postgres-accel): resolve secret references in the sidecar connection parameters (fixes #13296) by @claudespice in #14544
  • fix(cayenne): drain the in-memory CDC tier before a full rewrite scans it (fixes #14450) by @claudespice in #14486
  • feat(caching): warm SQL results cache on first refresh from persisted plan shapes by @lukekim in #14178
  • fix(ci): retry only the container startup in the zero-retry MySQL CDC tests by @grokspice in #14727
  • feat(key-index): immutable secondary index runs over compound Arrow keys by @bjchambers in #14592
  • fix(federation): keep a fractional-to-integer cast out of the plans pushed to DuckDB, PostgreSQL, MySQL and BigQuery (fixes #14482) by @grokspice in #14601
  • fix(testoperator): give accelerated bench configs a 300s ready_wait and name unready datasets on timeout (fixes #13973) by @claudespice in #14728
  • fix(federation): keep arrow_typeof and its plan-introspection siblings local on every backend (fixes #14334) by @grokspice in #14695
  • test(bench): refresh the tpch_q16 explain snapshots for the null-aware NOT IN plan (fixes #13977) by @claudespice in #14731
  • ci: build trunk-push macOS legs on spiceai-macos-large again, keeping one-at-a-time coalescing by @grokspice in #14735
  • ci(verus): derive the verified crates from cargo metadata; pin vstd once by @bjchambers in #14709
  • docs: raise the test and evidence bar: differential first, exact assertions, performance always measured by @lukekim in #14723
  • test(vortex): wait for a dropped segment cache to be freed before asserting it is gone (fixes #13295) by @claudespice in #14470
  • ci(integration): report each integration part's own result in its required check by @grokspice in #14743
  • fix(cayenne): back off futile bakes under a violated query goal too by @lukekim in #14736
  • feat(object_store_occ): add transactional WAL and MVCC snapshots by @lukekim in #14732
  • fix(snapshots): bootstrap readers from late snapshot publications by @phillipleblanc in #14468
  • fix(cayenne): apply retention_sql to a mode: memory acceleration (fixes #14045) by @claudespice in #14344
  • fix(duckdb, sqlite): keep the first copy of a key a write repeats under on_conflict: drop (fixes #14629) by @claudespice in #14748
  • fix(cache): release a removed dataset's memory when the plan cache discards its plans (fixes #14251) by @claudespice in #14760
  • build(deps): bump the aws-sdk group with 2 updates by @dependabot in #14764
  • fix(llms): explain an Anthropic model's refusal of a sampling control instead of passing the bare 400 through (fixes #13564) by @grokspice in #14700
  • build(deps-dev): bump dompurify by @dependabot in #14613
  • fix(snapshots): allow cayenne_file_path and cayenne_metadata_dir on file_format: snapshot datasets (fixes #14696) by @sgrebnov in #14698
  • test(snapshots): make snapshot integration tests robust under parallel CI runs by @sgrebnov in #14765
  • feat(cayenne): keep the secondary index current across every write by @bjchambers in #14593
  • fix: Update tpch benchmark snapshots for accelerated/file[parquet]-duckdb[memory].yaml by @app/github-actions in #14674
  • Isolate Docker integration fixtures across concurrent processes by @phillipleblanc in #14639
  • ci: run Substrait Mode A TPC-H at SF 1 on spiceai-macos in every merge-queue entry by @lukekim in #14522
  • ci: stop repeating scheduled benchmarks: one source per accelerator, hosted sources weekly, a release commit once by @lukekim in #14681
  • build(deps): bump nvidia/cuda by @dependabot in #14763
  • fix(http): stop a non-2xx response body replacing an accelerated table's rows (fixes #13515) by @claudespice in #13538
  • fix(chat-api): answer a tool-call-only assistant turn instead of panicking (fixes #13207) by @grokspice in #14232
  • fix(sqlite, duckdb): keep a decimal AVG, and on SQLite a decimal SUM, out of the federated plan (fixes #14492) by @claudespice in #14670
  • fix(spiceai, duckdb, sharepoint): name the Spicepod key a missing-parameter error asks for (refs #14446) by @claudespice in #14769
  • Add table-bound ChangeSink ownership and backends by @phillipleblanc in #14704
  • fix(ci): serialize rustup installs on shared macOS runners by @lukekim in #14717
  • fix(cayenne): keep the per-shard PK index when the table-wide index is discarded by @lukekim in #14604
  • Upgrade to DataFusion 55.1 and Arrow 59.3 by @krinart in #14612
  • Update openapi.json by @app/github-actions in #14742
  • build(deps): bump rustls from 0.23.40 to 0.23.45 by @dependabot in #14773
  • Upgrade to iceberg-rust to 0.11.0 and DF 55.2 by @sgrebnov in #14771
  • fix(cayenne): write an append into an empty unkeyed table as one load by @phillipleblanc in #14784
  • fix(mysql_replication): don't re-send a live member's delivered commits after a reconnect by @lukekim in #14751
  • ci(codeql): don't fail the SARIF upload when the merge queue already deleted its ref by @grokspice in #14793
  • fix(acceleration): accept time_format iso8601 when the time_column is already a timestamp by @lukekim in #14777
  • fix(cache): store SQL results that go stale mid-query by @lukekim in #14710
  • fix(cayenne): serve filtered and global maintained aggregates by @lukekim in #14761
  • perf(cayenne): build the checkpoint's tombstone union after releasing the capture locks by @lukekim in #14759
  • ci: keep no artifacts or caches from PR and merge-queue checks, and run gating macOS jobs on spiceai-macos-large by @lukekim in #14804
  • fix(cayenne): keep maintenance from deleting files a snapshot is archiving by @sgrebnov in #14789
  • fix(snapshots): never overwrite shared snapshot metadata, and publish more than once to a file:// location by @lukekim in #14582
  • perf(cdc): build deferred change rows in the reader while the apply runs by @lukekim in #14746
  • fix: Update tpch benchmark snapshots for accelerated/on_zero_results/file[parquet]-cayenne[file]-on_zero_results.yaml by @app/github-actions in #14786
  • test(testoperator): add a cold-start time-to-ready regression test by @phillipleblanc in #14785
  • test(testoperator): make append tests more robust by @sgrebnov in #14817
  • fix(runtime,data_components): wait for change-data-capture sources to record their positions before shutdown closes the accelerations (fixes #14523) by @grokspice in #14702
  • fix(graphql): detect non-JSON responses and treat gateway errors as retryable by @lukekim in #14781
  • fix(runtime): require keys for s3_auth: key, honor gs:// state location params, and leave newer rate-control state alone by @lukekim in #14584
  • build(deps): bump hickory-resolver from 0.26.1 to 0.26.2 by @dependabot in #14774
  • fix(listing): skip zero-byte objects (S3 folder markers) in format-selected listings by @sgrebnov in #14822
  • Avoid unnecessary repartitioning in indexed dynamic-filter joins by @bjchambers in #14799
  • feat(cayenne): keep one row per primary key automatically and deprecate on_conflict by @bjchambers in #14726
  • fix(object-store): stop reading a response body at its first error by @phillipleblanc in #14831
  • fix(ci): use byte-order collation in the benchmark Postgres container (refs #14815) by @sgrebnov in #14836
  • fix(listing): keep a partition predicate as a residual filter on the _location fast path by @grokspice in #14790
  • fix(ci): search only the unixodbc keg, not all of Homebrew's lib, in the macOS release builds by @grokspice in #14846
  • fix(federation): keep date, timestamp and interval values local on SQLite reached through ADBC or ODBC (fixes #14753) by @claudespice in #14840
  • chore: pin the fork at spiceai/datafusion#249 and guard the DataFusion fixes backported from 54 by @krinart in #14800
  • Revert "fix(sqlite, duckdb): keep a decimal AVG, and on SQLite a decimal SUM, out of the federated plan (fixes #14492) (#14670)" by @krinart in #14825
  • perf(runtime): skip EnsureRequirements while planning a point lookup by @lukekim in #14807
  • fix(testoperator): report CH-benCH queries with no rows on either side as vacuous, not as matches by @lukekim in #14810
  • feat(acceleration): make Cayenne the default acceleration engine by @phillipleblanc in #14837
  • fix(cayenne): keep a join's LIMIT when the oversized-join rewrite makes it a sort-merge join by @lukekim in #14805
  • test(bigquery): exclude subtrees an empty join build side never runs from the corpus job count (fixes #14848) by @sgrebnov in #14849
  • fix(cayenne): fail contradictory write settings once, and order NULL times below every time by @bjchambers in #14847
  • fix(runtime): serve an existing acceleration while its source is unavailable (fixes #14610) by @bjchambers in #14624
  • Prune object-store file listing by metadata columns (#14264) by @Jeadie in #14303
  • fix(vortex): fold partition values into the file-pruning predicate; test null-equal joins under mode:file by @lukekim in #14797
  • feat(rate-control): adaptive, bounded and cluster-coordinated HTTP rate controls by @Jeadie in #14143
  • fix(sql): align date_part('dow') with EXTRACT(dow) Sunday=0 by @lukekim in #14796
  • fix(runtime): carry API-key principal into MCP tools/call by @lukekim in #14828
  • perf(cayenne): hold maintained-aggregate retraction state in a compact shared index by @lukekim in #14762
  • fix(deps): keep a hash join's order once it has a dynamic filter, so vector_search plans again by @Jeadie in #14857
  • fix(cayenne): fence schema statistics and control maintenance tests by @bjchambers in #14856
  • fix(github): retry nested GraphQL pages in place and fail incomplete nested connections by @sgrebnov in #14862
  • fix(cayenne): load an append dataset that has retention_sql and orders versions by time by @phillipleblanc in #14878
  • fix(cdc): Cayenne replication needs only a primary key, not on_conflict by @phillipleblanc in #14880
  • feat(connector-huggingface): Hugging Face datasets data connector (hf://datasets/...) by @lukekim in #14877
  • fix(iceberg): Iceberg REST clients read the real table or get an error, never an empty one by @lukekim in #14588

Full Changelog: https://github.com/spiceai/spiceai/compare/v2.3.2...v2.4.0-rc.1

Spice v2.0.1 (Jun 17, 2026)

ยท 4 min read
Phillip LeBlanc
Co-Founder and CTO of Spice AI

Spice v2.0.1 is now available! ๐Ÿ› ๏ธ

Spice v2.0.1 is a patch release focused on reliability and performance. It speeds up Apache Iceberg reads and fixes bugs across AWS S3 and object-store datasets, data acceleration, distributed query, and authenticated access.

What's New in v2.0.1โ€‹

Faster Iceberg Reads with Parallel File Scanningโ€‹

The Apache Iceberg reader now scans data files in parallel (#11331), improving read throughput and latency for Iceberg tables that span many files.

AWS S3 & Object-Store Reliabilityโ€‹

Three fixes improve S3 and object-store dataset behavior:

  • Refresh-skip restored (#11339): ETag/Version-based refresh-skip works reliably again, so unchanged S3 objects are no longer re-downloaded on every refresh.
  • Retry when source files are not yet available (#11342): an object-store dataset whose source files are not present at startup now retries and becomes ready once the data appears, instead of failing permanently.
  • Path-style addressing for dotted bucket names (#11347): on standard AWS, buckets whose names contain dots now default to path-style addressing, avoiding TLS wildcard certificate errors under virtual-hosted-style HTTPS.

Data Acceleration & Distributed Query Fixesโ€‹

Two fixes ensure accelerated datasets behave correctly in more configurations:

  • Acceleration endpoints (#11345): /v1/datasets/{name}/acceleration/refresh (and the related update-refresh-sql, partition-filters, and snapshots endpoints) now work for all accelerated datasets, fixing cases where some incorrectly reported Table is not accelerated.
  • Distributed clusters (#11226): the distributed query coordinator now serves accelerated data from executors for all accelerated datasets, instead of falling back to reading from the source for some.

Authenticated Query Fixesโ€‹

With authentication enabled, queries now consistently run as the requesting user (#11253), so per-user behavior such as results caching is correctly scoped to each user.

Contributorsโ€‹

Breaking Changesโ€‹

No breaking changes.

Cookbook Updatesโ€‹

No new cookbook recipes.

The Spice Cookbook includes more than 100 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v2.0.1, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:2.0.1 image:

docker pull spiceai/spiceai:2.0.1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 2.0.1

AWS Marketplace:

Spice is available in the AWS Marketplace.

What's Changedโ€‹

Changelogโ€‹

Full Changelog: https://github.com/spiceai/spiceai/compare/v2.0.0...v2.0.1

Spice v1.11.5 (Apr 1, 2026)

ยท 4 min read
Sergei Grebnov
Member of Technical Staff at Spice AI

Announcing the release of Spice v1.11.5! ๐Ÿ› ๏ธ

Spice v1.11.5 is a patch release improving on_zero_results: use_source fallback performance, Delta Lake timestamp predicate data skipping, S3 Parquet read performance, PostgreSQL partitioned table support, Cayenne target file size handling, and preparing the CLI for v2.0 runtime upgrades.

What's New in v1.11.5โ€‹

on_zero_results: use_source Fallback Performance Improvementโ€‹

Improved the on_zero_results: use_source fallback path to run DataFusion's physical optimizer on the federated scan plan (#9927). The fallback path now runs SessionState::physical_optimizers() rules on the federated scan plan before execution, enabling parallel file group scanning and other optimizations. This results in significantly faster fallback queries on multi-core machines, particularly for file-based data sources like Delta Lake.

Delta Lake: Improved Data Skipping for >= Timestamp Predicatesโ€‹

Delta Lake table scans with >= timestamp filters now correctly prune files that do not match the predicate (#9932), improving query performance through more effective data skipping (file-level pruning).

PostgreSQL: Partitioned Tables Supportโ€‹

The PostgreSQL data connector now supports partitioned tables (#9997) for both federated and accelerated queries.

S3 Parquet Read Performance Improvementโ€‹

Improved parquet read performance from S3 and other object stores (#10064), particularly for tables with many columns. Column data ranges are now coalesced into fewer, larger requests instead of being fetched individually, reducing the number of HTTP round-trips.

Cayenne: Ensure Target File Size is Respectedโ€‹

The Cayenne accelerator now correctly respects the configured target file size (#10071). Previously, Cayenne could produce many small, fragmented Vortex files; with this fix, files are written at the expected target size, improving storage efficiency and query performance.

CLI: Support for v2.0 Runtime Upgradesโ€‹

The Spice CLI can now upgrade to v2.0 runtime versions. This enables upgrading to v2.0 release candidates and, once released, the v2.0 stable runtime.

spice upgrade v2.0.0-rc.1

Running spice upgrade without a version will upgrade to the latest stable version, including v2.0 once released.

Note: Native Windows runtime builds will no longer be provided in v2.0. Use WSL for local development instead.

Contributorsโ€‹

Breaking Changesโ€‹

No breaking changes.

Cookbook Updatesโ€‹

No new cookbook recipes.

The Spice Cookbook includes 86 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v1.11.5, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.11.5 image:

docker pull spiceai/spiceai:1.11.5

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 1.11.5

AWS Marketplace:

Spice is available in the AWS Marketplace.

What's Changedโ€‹

Changelogโ€‹

  • fix(runtime): Run physical optimizer on FallbackOnZeroResultsScanExec fallback plan by @sgrebnov in #9927
  • fix(delta_lake): Fix data skipping for >= timestamp predicates by @sgrebnov in #9932
  • fix(PostgreSQL): Fix schema discovery for PostgreSQL partitioned tables by @sgrebnov in #9997
  • fix(cli): Skip models variant download for v2+ in upgrade/install by @lukekim and @sgrebnov in #10052
  • perf(s3): Improve Parquet read performance by @sgrebnov in #10064
  • fix(cayenne): Ensure Cayenne respects target file size by @krinart in #10071

Full Changelog: https://github.com/spiceai/spiceai/compare/v1.11.4...v1.11.5

Spice v1.11.4 (Mar 12, 2026)

ยท 5 min read
Sergei Grebnov
Member of Technical Staff at Spice AI

Announcing the release of Spice v1.11.4! โšก

Spice v1.11.4 is a patch release improving S3 metadata column query robustness and enabling on_zero_results: use_source for accelerated views.

What's New in v1.11.4โ€‹

Accelerated Views: on_zero_results: use_source Supportโ€‹

Accelerated views now support the on_zero_results: use_source configuration (#9699). Previously, accelerated views only supported on_zero_results: return_empty, which returned an empty result set when the accelerated data contained no matching rows. With this change, views can fall back to querying the source data when the accelerated query returns zero results, matching the behavior already available for accelerated datasets.

Example configuration:

views:
- name: sales_summary
sql: |
SELECT region, SUM(amount) as total
FROM sales
GROUP BY region
acceleration:
enabled: true
on_zero_results: use_source

How the Fallback Worksโ€‹

When an accelerated view is configured with on_zero_results: use_source, the following happens at query time:

  1. The accelerated store is queried first. The query runs against the view's accelerated data (e.g., Spice Cayenne, Arrow, DuckDB, or SQLite).

  2. If the accelerated query returns zero rows, the runtime falls back to re-executing the view's SQL query against the datasets it references.

  3. Referenced datasets are queried according to their own configuration. The view's SQL is re-executed against each referenced dataset as it is configured. This means:

    • If a referenced dataset is accelerated, the query hits that dataset's accelerated store โ€” not the raw data source.
    • If a referenced dataset is accelerated with on_zero_results: use_source and its accelerated store also returns zero rows, it will independently fall back to its own federated data source (e.g., Postgres, S3, etc.).
    • If a referenced dataset is federated (not accelerated), the query goes directly to the data source.

This means the fallback can chain through multiple layers: first the view's acceleration, then each referenced dataset's acceleration, and finally the original data source โ€” each layer independently applying its own on_zero_results behavior.

Example: Multi-layer fallback

datasets:
- from: postgres:orders
name: orders
acceleration:
enabled: true
refresh_sql: "SELECT * FROM orders WHERE created_at > now() - interval '7 days'"
on_zero_results: use_source # Falls back to Postgres if accelerated data has no matches

views:
- name: recent_orders_summary
sql: |
SELECT status, COUNT(*) as order_count
FROM orders
GROUP BY status
acceleration:
enabled: true
on_zero_results: use_source # Falls back to re-running the SQL against referenced datasets

In this example, a query like SELECT * FROM recent_orders_summary WHERE status = 'cancelled' follows this path:

  1. Queries recent_orders_summary in the view's accelerated store (DuckDB/SQLite).
  2. If zero rows are returned, re-executes SELECT status, COUNT(*) ... FROM orders GROUP BY status against the orders dataset.
  3. Since orders is accelerated, the query hits the orders accelerated store.
  4. If orders also returns zero rows (e.g., the refresh_sql excluded cancelled orders), it falls back to querying Postgres directly.

S3 Data Connector: More Robust Metadata Column Handlingโ€‹

Improved the robustness of metadata column (location, last_modified, size) handling for S3 datasets. Building on the v1.11.3 release, this update addresses an additional edge case where the query optimizer's projection swap could cause an index out of bounds panic when metadata columns are referenced in projections with filters or scalar functions.

Contributorsโ€‹

Breaking Changesโ€‹

No breaking changes.

Cookbook Updatesโ€‹

No new cookbook recipes.

The Spice Cookbook includes 86 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v1.11.4, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.11.4 image:

docker pull spiceai/spiceai:1.11.4

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 1.11.4

AWS Marketplace:

Spice is available in the AWS Marketplace.

What's Changedโ€‹

Changelogโ€‹

  • fix(s3): Make metadata column handling more robust by @sgrebnov in #9714
  • feat(views): Enable on_zero_results: use_source for accelerated views by @krinart in #9699

Full Changelog: https://github.com/spiceai/spiceai/compare/v1.11.3...v1.11.4

Spice v1.11.3 (Mar 9, 2026)

ยท 3 min read
Phillip LeBlanc
Co-Founder and CTO of Spice AI

Announcing the release of Spice v1.11.3! ๐Ÿ› ๏ธ

Spice v1.11.3 is a patch release fixing schema consistency issues in the S3 and FlightSQL data connectors, improving CDC cache invalidation, and enhancing the HTTP data connector's error handling and response metadata.

What's New in v1.11.3โ€‹

S3 Data Connector Fixโ€‹

Fixed an issue where queries using metadata columns (location, last_modified, size) on S3 datasets produced Input field name does not match with the projection expression errors (#9647). This occurred when projecting metadata columns with filters or scalar functions (e.g., SELECT lower(location) FROM table WHERE location = '...'), and when projection returned no matching files.

FlightSQL Schema Consistencyโ€‹

Fixed an issue where the Flight SQL JDBC driver returned Unsupported ArrowType Utf8View errors when performing ::TEXT type casts (#9253). The FlightSQL endpoint now maps view types (e.g., Utf8View, BinaryView) to their non-view equivalents, ensuring compatibility with JDBC and ODBC clients.

CDC Cache Invalidationโ€‹

Fixed an issue where the SQL results cache was invalidated on every change stream poll, even when zero records were returned (#9472). This caused near-total cache miss rates for datasets using refresh_mode: changes (e.g., DynamoDB Streams), effectively rendering the cache useless. Cache invalidation now only occurs when a change batch contains actual data changes.

HTTP Data Connector Improvementsโ€‹

  • HTTP error responses (e.g., 5xx) are now excluded from the cache, preventing transient server errors from polluting cached results.
  • Added a response_headers column (Map type) to HTTP responses, providing access to response header metadata in query results.

Contributorsโ€‹

Breaking Changesโ€‹

No breaking changes.

Cookbook Updatesโ€‹

No new cookbook recipes.

The Spice Cookbook includes 86 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v1.11.3, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.11.3 image:

docker pull spiceai/spiceai:1.11.3

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 1.11.3

AWS Marketplace:

Spice is available in the AWS Marketplace.

What's Changedโ€‹

Changelogโ€‹

  • fix(s3): Fix metadata column schema mismatches in projected queries by @sgrebnov in #9664
  • s3_metadata_columns tests: include test for location outside table prefix by @sgrebnov in #9676
  • Fix Flight SQL schema consistency: expand view types and verify field names by @sgrebnov in #9438
  • Improve CDC cache invalidation by @krinart in #9651
  • Skip caching http error response + add response_headers by @krinart in #9670

Full Changelog: https://github.com/spiceai/spiceai/compare/v1.11.2...v1.11.3

Spice v1.10.3 (Dec 29, 2025)

ยท 2 min read
Phillip LeBlanc
Co-Founder and CTO of Spice AI

Announcing the release of Spice v1.10.3! ๐Ÿš€

v1.10.3 is a patch release with improved startup reliability, fixes for Azure BlobFS versioned containers, S3 custom endpoint query resolution, and a fix for the OpenAI Responses API.

What's New in v1.10.3โ€‹

Additional Improvements & Bug Fixesโ€‹

  • Reliability: Telemetry exporter initialization now runs asynchronously, preventing blocked startup in environments with network restrictions (e.g., Kubernetes with restrictive network policies).
  • Reliability: Fixed an issue where queries on Azure Blob containers with versioning enabled would fail with "Azure does not support suffix range requests" error in distributed query mode.
  • Reliability: Fixed S3 location-based queries against custom S3 endpoints (e.g., MinIO, LocalStack). Queries with location predicates on datasets using s3_endpoint and s3_region parameters now correctly route to the configured endpoint instead of defaulting to AWS S3.
  • Reliability: Fixed "project index out of bounds" errors in the query optimizer when union children have mismatched schemas. The optimizer now validates schema compatibility before applying projection pushdown.
  • Reliability: Fixed an issue where the OpenAI Responses API (/v1/responses) was not working correctly.

Contributorsโ€‹

Breaking Changesโ€‹

No breaking changes.

Cookbook Updatesโ€‹

No major cookbook updates.

The Spice Cookbook includes 84 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v1.10.3, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.10.3 image:

docker pull spiceai/spiceai:1.10.3

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

๐ŸŽ‰ Spice is now available in the AWS Marketplace!

What's Changedโ€‹

Changelogโ€‹

Spice v1.5.0 (July 21, 2025)

ยท 14 min read
Evgenii Khramkov
Member of Technical Staff at Spice AI

Announcing the release of Spice v1.5.0! ๐Ÿ”

Spice v1.5.0 brings major upgrades to search and retrieval. It introduces native support for Amazon S3 Vectors, enabling petabyte scale vector search directly from S3 vector buckets, alongside SQL-integrated vector and tantivy-powered full-text search, partitioning for DuckDB acceleration, and automated refreshes for search indexes and views. It includes the AWS Bedrock Embeddings Model Provider, the Oracle Database connector, and the now-stable Spice.ai Cloud Data Connector, and the upgrade to DuckDB v1.3.2.

What's New in v1.5.0โ€‹

Amazon S3 Vectors Support: Spice.ai now integrates with Amazon S3 Vectors, launched in public preview on July 15, 2025, enabling vector-native object storage with built-in indexing and querying. This integration supports semantic search, recommendation systems, and retrieval-augmented generation (RAG) at petabyte scale with S3โ€™s durability and elasticity. Spice.ai manages the vector lifecycleโ€”ingesting data, creating embeddings with models like Amazon Titan or Cohere via AWS Bedrock, or others available on HuggingFace, and storing it in S3 Vector buckets.

Spice integration with Amazon S3 Vectors

Example Spicepod.yml configuration for S3 Vectors:

datasets:
- from: s3://my_data_bucket/data/
name: my_vectors
params:
file_format: parquet
acceleration:
enabled: true
vectors:
engine: s3_vectors
params:
s3_vectors_aws_region: us-east-2
s3_vectors_bucket: my-s3-vectors-bucket
columns:
- name: content
embeddings:
- from: bedrock_titan
row_id:
- id

Example SQL query using S3 Vectors:

SELECT *
FROM vector_search(my_vectors, 'Cricket bats', 10)
WHERE price < 100
ORDER BY score

For more details, refer to the S3 Vectors Documentation.

SQL-integrated Search: Vector and BM25-scored full-text search capabilities are now natively available in SQL queries, extending the power of the POST v1/search endpoint to all SQL workflows.

Example Vector-Similarity-Search (VSS) using the vector_search UDTF on the table reviews for the search term "Cricket bats":

SELECT review_id, review_text, review_date, score
FROM vector_search(reviews, "Cricket bats")
WHERE country_code="AUS"
LIMIT 3

Example Full-Text-Search (FTS) using the text_search UDTF on the table reviews for the search term "Cricket bats":

SELECT review_id, review_text, review_date, score
FROM text_search(reviews, "Cricket bats")
LIMIT 3

DuckDB v1.3.2 Upgrade: Upgraded DuckDB engine from v1.1.3 to v1.3.2. Key improvements include support for adding primary keys to existing tables, resolution of over-eager unique constraint checking for smoother inserts, and 13% reduced runtime on TPC-H SF100 queries through extensive optimizer refinements. The v1.2.x release of DuckDB was skipped due to a regression in indexes.

Partitioned Acceleration: DuckDB file-based accelerations now support partition_by expressions, enabling queries to scale to large datasets through automatic data partitioning and query predicate pruning. New UDFs, bucket and truncate, simplify partition logic.

New UDFs useful for partition_by expressions:

  • bucket(num_buckets, col): Partitions a column into a specified number of buckets based on a hash of the column value.
  • truncate(width, col): Truncates a column to a specified width, aligning values to the nearest lower multiple (e.g., truncate(10, 101) = 100).

Example Spicepod.yml configuration:

datasets:
- from: s3://my_bucket/some_large_table/
name: my_table
params:
file_format: parquet
acceleration:
enabled: true
engine: duckdb
mode: file
partition_by: bucket(100, account_id) # Partition account_id into 100 buckets

Full-Text-Search (FTS) Index Refresh: Accelerated datasets with search indexes maintain up-to-date results with configurable refresh intervals.

Example refreshing search indexes on body every 10 seconds:

datasets:
- from: github:github.com/spiceai/docs/pulls
name: spiceai.doc.pulls
params:
github_token: ${secrets:GITHUB_TOKEN}
acceleration:
enabled: true
refresh_mode: full
refresh_check_interval: 10s
columns:
- name: body
full_text_search:
enabled: true
row_id:
- id

Scheduled View Refresh: Accelerated Views now support cron-based refresh schedules using refresh_cron, automating updates for accelerated data.

Example Spicepod.yml configuration:

views:
- name: my_view
sql: SELECT 1
acceleration:
enabled: true
refresh_cron: '0 * * * *' # Every hour

For more details, refer to Scheduled Refreshes.

Multi-column Vector Search: For datasets configured with embeddings on more than one column, POST v1/search and similarity_search perform parallel vector search on each column, aggregating results using reciprocal rank fusion.

Example Spicepod.yml for multi-column search:

datasets:
- from: github:github.com/apache/datafusion/issues
name: datafusion.issues
params:
github_token: ${secrets:GITHUB_TOKEN}
columns:
- name: title
embeddings:
- from: hf_minilm
- name: body
embeddings:
- from: openai_embeddings

AWS Bedrock Embeddings Model Provider: Added support for AWS Bedrock embedding models, including Amazon Titan Text Embeddings and Cohere Text Embeddings.

Example Spicepod.yml:

embeddings:
- from: bedrock:cohere.embed-english-v3
name: cohere-embeddings
params:
aws_region: us-east-1
input_type: search_document
truncate: END
- from: bedrock:amazon.titan-embed-text-v2:0
name: titan-embeddings
params:
aws_region: us-east-1
dimensions: '256'

For more details, refer to the AWS Bedrock Embedding Models Documentation.

Oracle Data Connector: Use from: oracle: to access and accelerate data stored in Oracle databases, deployed on-premises or in the cloud.

Example Spicepod.yml:

datasets:
- from: oracle:"SH"."PRODUCTS"
name: products
params:
oracle_host: 127.0.0.1
oracle_username: scott
oracle_password: tiger

See the Oracle Data Connector documentation.

GitHub Data Connector: The GitHub data connector supports query and acceleration of members, the users of an organization.

Example Spicepod.yml configuration:

datasets:
- from: github:github.com/spiceai/members # General format: github.com/[org-name]/members
name: spiceai.members
params:
# With GitHub Apps (recommended)
github_client_id: ${secrets:GITHUB_SPICEHQ_CLIENT_ID}
github_private_key: ${secrets:GITHUB_SPICEHQ_PRIVATE_KEY}
github_installation_id: ${secrets:GITHUB_SPICEHQ_INSTALLATION_ID}
# With GitHub Tokens
# github_token: ${secrets:GITHUB_TOKEN}

See the GitHub Data Connector Documentation

Spice.ai Cloud Data Connector: Graduated to Stable.

spice-rs SDK Release: The Spice Rust SDK has updated to v3.0.0. This release includes optimizations for the Spice client API, adds robust query retries, and custom metadata configurations for spice queries.

Contributorsโ€‹

Breaking Changesโ€‹

  • Search HTTP API Response: POST v1/search response payload has changed. See the new API documentation for details.
  • Model Provider Parameter Prefixes: Model Provider parameters use provider-specific prefixes instead of openai_ prefixes (e.g., hf_temperature for HuggingFace, anthropic_max_completion_tokens for Anthropic, perplexity_tool_choice for Perplexity). The openai_ prefix remains supported for backward compatibility but is deprecated and will be removed in a future release.

Cookbook Updatesโ€‹

The Spice Cookbook now includes 72 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v1.5.0, download and install the specific binary from github.com/spiceai/spiceai/releases/tag/v1.5.0 or pull the v1.5.0 Docker image (spiceai/spiceai:1.5.0).

What's Changedโ€‹

Dependenciesโ€‹

Changelogโ€‹

  • fix: openai model endpoint (#6394) by @Sevenannn in #6394
  • Enable configuring otel endpoint from spice run (#6360) by @Advayp in #6360
  • Enable Oracle connector in default build configuration (#6395) by @sgrebnov in #6395
  • fix llm integraion test (#6398) by @Sevenannn in #6398
  • Promote spice cloud connector to stable quality (#6221) by @Sevenannn in #6221
  • v1.5.0-rc.1 release notes (#6397) by @lukekim in #6397
  • Fix model nsql integration tests (#6365) by @Sevenannn in #6365
  • Fix incorrect UDTF name and SQL query (#6404) by @lukekim in #6404
  • Update v1.5.0-rc.1.md (#6407) by @sgrebnov in #6407
  • Improve error messages (#6405) by @lukekim in #6405
  • build(deps): bump Jimver/cuda-toolkit from 0.2.25 to 0.2.26 (#6388) by @app/dependabot in #6388
  • Upgrade dependabot dependencies (#6411) by @phillipleblanc in #6411
  • Fix projection pushdown issues for document based file connector (#6362) by @Advayp in #6362
  • Add a PartitionedDuckDB Accelerator (#6338) by @kczimm in #6338
  • Use vector_search() UDTF in HTTP APIs (#6417) by @Jeadie in #6417
  • add supported types (#6409) by @kczimm in #6409
  • Enable session time zone override for MySQL (#6426) by @sgrebnov in #6426
  • Acceleration-like indexing for full text search indexes. (#6382) by @Jeadie in #6382
  • Provide error message when partition by expression changes (#6415) by @kczimm in #6415
  • Add support for Oracle Autonomous Database connections (Oracle Cloud) (#6421) by @sgrebnov in #6421
  • prune partitions for exact and in list with and without UDFs (#6423) by @kczimm in #6423
  • Fixes and reenable FTS tests (#6431) by @Jeadie in #6431
  • Upgrade DuckDB to 1.3.2 (#6434) by @phillipleblanc in #6434
  • Fix issue in limit clause for the Github Data connector (#6443) by @Advayp in #6443
  • Upgrade iceberg-rust to 0.5.1 (#6446) by @phillipleblanc in #6446
  • v1.5.0-rc.2 release notes (#6440) by @lukekim in #6440
  • Oracle: add automated TPC-H SF1 benchmark tests (#6449) by @sgrebnov in #6449
  • fix: Update benchmark snapshots (#6455) by @app/github-actions in #6455
  • Preserve ArrowError in arrow_tools::record_batch (#6454) by @mach-kernel in #6454
  • fix: Update benchmark snapshots (#6465) by @app/github-actions in #6465
  • Add option to preinstall Oracle ODPI-C library in Docker image (#6466) by @sgrebnov in #6466
  • Include Oracle connector (federated mode) in automated benchmarks (#6467) by @sgrebnov in #6467
  • Update crates/llms/src/bedrock/embed/mod.rs by @lukekim in #6468
  • v1.5.0-rc.3 release notes (#6474) by @lukekim in #6474
  • Add integration tests for S3 Vectors filters pushdown (#6469) by @sgrebnov in #6469
  • check for indexedtableprovider when finding tables to search on (#6478) by @Jeadie in #6478
  • Parse fully qualified table names in UDTFs (#6461) by @Jeadie in #6461
  • Add integration test for S3 Vectors to cover data update (overwrite) (#6480) by @sgrebnov in #6480
  • Add 'Run all tests' option for models tests and enable Bedrock tests (#6481) by @sgrebnov in #6481
  • Add support for a members table type for the GitHub Data Connector (#6464) by @Advayp in #6464
  • S3 vector data cannot be null (#6483) by @Jeadie in #6483
  • Don't infer FixedSizeList size during indexing vectors. (#6487) by @Jeadie in #6487
  • Add support for retention_sql acceleration param (#6488) by @sgrebnov in #6488
  • Make dataset refresh progress tracing less verbose (#6489) by @sgrebnov in #6489
  • Use RwLock on tantivy index in FullTextDatabaseIndex for update concurrency (#6490) by @Jeadie in #6490
  • Add tests for dataset retention logic and refactor retention code (#6495) by @sgrebnov in #6495
  • Upgade dependabot dependencies (#6497) by @phillipleblanc in #6497
  • Add periodic tracing of data loading progress during dataset refresh (#6499) by @sgrebnov in #6499
  • Promote Oracle Data Connector to Alpha (#6503) by @sgrebnov in #6503
  • Use AWS SDK to provide credentials for Iceberg connectors (#6498) by @phillipleblanc in #6498
  • Add integration tests for partitioning (#6463) by @kczimm in #6463
  • Use top-level table in full-text search JOIN ON (#6491) by @Jeadie in #6491
  • Use accelerated table in vector_search JOIN operations when appropriate (#6516) by @Jeadie in #6516
  • Fix 'additional_column' for quoted columns (fix for qualified columns broke it) (#6512) by @Jeadie in #6512
  • Also use AWS SDK for inferring credentials for S3/Delta/Databricks Delta data connectors (#6504) by @phillipleblanc in #6504
  • Add per-dataset availability monitor configuration (#6482) by @phillipleblanc in #6482
  • Suppress the warning from the AWS SDK if it can't load credentials (#6533) by @phillipleblanc in #6533
  • Change default value of check_availability from default to auto (#6534) by @lukekim in #6534
  • README.md improvements for v1.5.0 (#6539) by @lukekim in #6539
  • Temporary disable s3_vectors_basic (#6537) by @sgrebnov in #6537
  • Ensure binder errors show before query and other (#6374) by @suhuruli in #6374
  • Update spiceai/duckdb-rs -> DuckDB 1.3.2 + index fix (#6496) by @mach-kernel in #6496
  • Update table-providers to latest version with DuckDB fixes (#6535) by @phillipleblanc in #6535
  • S3: default to public access if no auth is provided (#6532) by @sgrebnov in #6532

Spice v1.0-stable (Jan 20, 2025)

ยท 11 min read
William Croxson
Member of Technical Staff at Spice AI

๐ŸŽ‰ After 47 releases, Spice.ai OSS has reached production readiness with the 1.0-stable milestone!

The core runtime and features such as query federation, query acceleration, catalog integration, search and AI-inference have all graduated to stable status along with key component graduations across data connectors, data accelerators, catalog connectors, and AI model providers.

Highlights in v1.0-stableโ€‹

Breaking Changesโ€‹

  • Default Runtime Version: The CLI will install the GPU accelerated AI-capable Runtime by default (if supported), when running spice install or spice run. To force-install the non-GPU version, run spice install ai --cpu.

  • Default OpenAI Model: The default OpenAI model has updated to gpt-4o-mini.

  • Identifier Normalization: Unquoted identifiers such as table names are no longer normalized to lowercase. Identifiers will now retain their exact case as provided.

  • Sandboxed Docker Image: The Runtime Docker Image now runs the spiced process as the nobody user in a minimal chroot sandbox.

  • Insecure S3 and ABFS endpoints: The S3 and ABFS connectors now enforce insecure endpoint checks, preventing HTTP endpoints unless allow_http is explicitly enabled. Refer to the documentation for details.

Dependenciesโ€‹

No major dependency changes.

Upgradingโ€‹

To upgrade to v1.0.0, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.0.0 image:

docker pull spiceai/spiceai:1.0.0

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

Contributorsโ€‹

  • @peasee
  • @ewgenius
  • @Jeadie
  • @Sevenannn
  • @lukekim
  • @phillipleblanc
  • @sgrebnov

What's Changedโ€‹

- feat: Update load test criteria, testoperator updates by @peasee in <https://github.com/spiceai/spiceai/pull/4311>
- Update helm for v1.0.0-rc.5 by @ewgenius in <https://github.com/spiceai/spiceai/pull/4313>
- Update spicepod.schema.json by @github-actions in <https://github.com/spiceai/spiceai/pull/4318>
- Bump version to v1.0.0, update SECURITY.md by @ewgenius in <https://github.com/spiceai/spiceai/pull/4314>
- Initial criteria for models, embeddings by @Jeadie in <https://github.com/spiceai/spiceai/pull/4223>
- Update benchmark snapshots by @github-actions in <https://github.com/spiceai/spiceai/pull/4321>
- Add dremio param for running load test by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4315>
- Promote Databricks (mode: delta_lake) connector to stable by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4328>
- Handle failed query in load test by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4327>
- feat: Use load test hours for baseline query sets by @peasee in <https://github.com/spiceai/spiceai/pull/4334>
- Fix typo in 1.0.0-rc.5 release notes by @ewgenius in <https://github.com/spiceai/spiceai/pull/4329>
- feat: add testoperator data consistency by @peasee in <https://github.com/spiceai/spiceai/pull/4319>
- docs: Release DuckDB connector stable by @peasee in <https://github.com/spiceai/spiceai/pull/4335>
- Fix DocumentDB -> DynamoDB by @lukekim in <https://github.com/spiceai/spiceai/pull/4339>
- Update benchmark snapshots by @github-actions in <https://github.com/spiceai/spiceai/pull/4337>
- fix: Download hits.parquet from MinIO for benchmark by @peasee in <https://github.com/spiceai/spiceai/pull/4338>
- Update openapi.json by @github-actions in <https://github.com/spiceai/spiceai/pull/4341>
- Remove evil averages by @lukekim in <https://github.com/spiceai/spiceai/pull/4343>
- Don't run builds on non-code changes by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4344>
- Remove streaming requirement from Databricks spark Beta and Spark connector Beta by @ewgenius in <https://github.com/spiceai/spiceai/pull/4345>
- Update s3 tpcds spicepods by @ewgenius in <https://github.com/spiceai/spiceai/pull/4346>
- Explicitly set required scale factor for throughput and load tests by @ewgenius in <https://github.com/spiceai/spiceai/pull/4347>
- Fix s3 tpcds dataset name by @ewgenius in <https://github.com/spiceai/spiceai/pull/4348>
- Promote Iceberg Catalog Connector to Beta by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4350>
- Update s3 clickbench benchmark snapshots by @ewgenius in <https://github.com/spiceai/spiceai/pull/4351>
- fix: DuckDB clickbench on zero results by @peasee in <https://github.com/spiceai/spiceai/pull/4349>
- Add integration test with snapshots for databricks catalog connector by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4353>
- refactor: Remove on zero results from benchmarks, add data consistency workflow by @peasee in <https://github.com/spiceai/spiceai/pull/4354>
- Fix Bug: No field named body_embedding when do vector search with refresh sql containing subset of columns by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4297>
- docs: Update roadmap by @peasee in <https://github.com/spiceai/spiceai/pull/4364>
- feat: Release accelerators stable by @peasee in <https://github.com/spiceai/spiceai/pull/4361>
- Add TPCH/TPCDS test spicepods for MySQL by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4365>
- Catch when an insecure (http) S3 and ABFS data connectors endpoint is used without specifying the `allow_http` parameter by @ewgenius in <https://github.com/spiceai/spiceai/pull/4363>
- Update ROADMAP - Iceberg catalog alpha for v1.0 by @ewgenius in <https://github.com/spiceai/spiceai/pull/4367>
- Promote databricks catalog and databricks (spark_connect) connector to beta by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4369>
- Update Roadmap - Iceberg beta by @ewgenius in <https://github.com/spiceai/spiceai/pull/4373>
- Build CUDA binaries for Linux by @Jeadie in <https://github.com/spiceai/spiceai/pull/4320>
- Promote Nvidia NIM as Alpha by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4380>
- Promote xai to alpha by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4381>
- Update stable criteria for object store based connectors by @ewgenius in <https://github.com/spiceai/spiceai/pull/4383>
- Testoperator: http consistency and overhead tests, fixes and ci by @ewgenius in <https://github.com/spiceai/spiceai/pull/4382>
- Promote S3 Data Connector to Stable by @ewgenius in <https://github.com/spiceai/spiceai/pull/4385>
- Download platform-supported CUDA binary version on Linux by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4356>
- Fix http consistency test workflow, add overhead workflow by @ewgenius in <https://github.com/spiceai/spiceai/pull/4387>
- feat: Add Postgres test spicepods by @peasee in <https://github.com/spiceai/spiceai/pull/4388>
- Fix typos + specific in model criteria; Make explicit alpha/beta tests for LLMS in `crates/llms/tests`. by @Jeadie in <https://github.com/spiceai/spiceai/pull/4377>
- Fix federation bug for correlated subqueries of deeply nested Dremio tables by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4389>
- Fix http overhead workflow by @ewgenius in <https://github.com/spiceai/spiceai/pull/4390>
- Tweak model tests, fix embedding input by @ewgenius in <https://github.com/spiceai/spiceai/pull/4391>
- Promote Dremio to Stable quality by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4392>
- Add beta functionality tests for embedding models. by @Jeadie in <https://github.com/spiceai/spiceai/pull/4352>
- docs: Release postgres connector stable by @peasee in <https://github.com/spiceai/spiceai/pull/4398>
- Increase timeout for model response in E2E tests by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4399>
- Disable ident normalization (i.e. `SELECT MyColumn from table` works) by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4400>
- Preserve schema metadata by @ewgenius in <https://github.com/spiceai/spiceai/pull/4402>
- Make models integration tests tracing less verbose by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4403>
- Fix `cuda` feature build on Windows by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4404>
- Promote MySQL to Stable by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4406>
- docs: Release Delta Lake and Unity catalog by @peasee in <https://github.com/spiceai/spiceai/pull/4405>
- Use `gpt-4o-mini` as a default model for openai provider by @ewgenius in <https://github.com/spiceai/spiceai/pull/4410>
- Fix streaming for Openai and Anthropic by @Jeadie in <https://github.com/spiceai/spiceai/pull/4409>
- Tweak model loading and missing tool errors messages by @ewgenius in <https://github.com/spiceai/spiceai/pull/4412>
- Spice CLI: fallback to CPU build for unsupported GPU Compute Capability by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4407>
- Build Windows CUDA binaries as part of `build_and_release` workflow by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4386>
- Update docs link by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4416>
- feat: Add CPU models install escape hatch by @peasee in <https://github.com/spiceai/spiceai/pull/4419>
- Handle OpenAI API Errors by @ewgenius in <https://github.com/spiceai/spiceai/pull/4417>
- Update spice cli to use `GH_TOKEN` or `GITHUB_TOKEN` env variables when calling releases api by @ewgenius in <https://github.com/spiceai/spiceai/pull/4175>
- Implement secure sandboxing for Docker image by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4411>
- Automatically install supported CUDA binary on Windows by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4420>
- Metrics for LLMs+ embeddings by @Jeadie in <https://github.com/spiceai/spiceai/pull/4418>
- Jeadie/25 01 17/beta perf by @Jeadie in <https://github.com/spiceai/spiceai/pull/4397>
- Pass GitHub token to all CI steps calling spice run by @ewgenius in <https://github.com/spiceai/spiceai/pull/4423>
- Run the models integration tests on PRs by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4421>
- Run CUDA builds in a separate workflow by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4430>
- Promote OpenAI models and embeddings providers to RC by @ewgenius in <https://github.com/spiceai/spiceai/pull/4432>
- Update link to retrieval-augmented generation (RAG) details by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4433>
- Unity catalog should strip parameter prefix before passing parameters to delta lake factory by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4436>
- Update quickstart traces to match current version by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4435>
- Update Supported Embeddings Providers Readme section by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4434>
- Local models can stream tools by @Jeadie in <https://github.com/spiceai/spiceai/pull/4429>
- fix: Use MetricsCollector::show() for HTTP testoperator commands by @peasee in <https://github.com/spiceai/spiceai/pull/4442>
- Fix run query action by @ewgenius in <https://github.com/spiceai/spiceai/pull/4444>
- Default to AI-enabled runtime for `spice run`/`spice install` by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4443>
- Change no spicepod.yaml log to warning by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4447>
- refactor: Update Catalog Connector error messages by @peasee in <https://github.com/spiceai/spiceai/pull/4441>
- Fix panic when converting OTel metrics by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4449>
- refactor: Update model errors by @peasee in <https://github.com/spiceai/spiceai/pull/4446>
- Update spiceai/mistral.rs to silence metadata logs by @ewgenius in <https://github.com/spiceai/spiceai/pull/4452>
- fix xAI; don't use openai defaults by @Jeadie in <https://github.com/spiceai/spiceai/pull/4450>
- Improves the UX of using huggingface models by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/4451>
- Add GH Workflow to test `spice ai` runtime installation by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4448>
- fix: Use specific model errors where available by @peasee in <https://github.com/spiceai/spiceai/pull/4454>
- Detect and report unsupported embedding column type during dataset registration by @sgrebnov in <https://github.com/spiceai/spiceai/pull/4456>
- Handle Errors by @Jeadie in <https://github.com/spiceai/spiceai/pull/4455>
- Catch and report negative openai_temperature error by @Sevenannn in <https://github.com/spiceai/spiceai/pull/4453>
- Clarify release check error message if it is caused by wrong GH token by @ewgenius in <https://github.com/spiceai/spiceai/pull/4458>

**Full Changelog**: <https://github.com/spiceai/spiceai/compare/v1.0.0-rc.5...v1.0.0>

Resourcesโ€‹

Communityโ€‹

Spice.ai started with the vision to make AI easy for developers. We are building Spice.ai in the open and with the community. Reach out on Slack or by email to get involved.

Spice v0.20-beta (Nov 4, 2024)

ยท 4 min read
Phillip LeBlanc
Co-Founder and CTO of Spice AI

Announcing the release of Spice v0.20-beta ๐Ÿงฉ

Spice v0.20.0-beta improves federated query performance with column pruning and adds support for Metal (Apple Silicon) and CUDA (NVidia) accelerators. The S3, PostgreSQL, MySQL, and GitHub Data Connectors have graduated from Beta to Release Candidates. The Arrow, DuckDB, and SQLite Data Accelerators have graduated from Alpha to Beta.

Highlights in v0.20.0-betaโ€‹

Data Connectors: The S3, PostgreSQL, MySQL, and GitHub Data Connectors have graduated from beta to release candidate.

Data Accelerators: The Arrow, DuckDB, and SQLite Data Accelerators have graduated from alpha to beta.

Metal and CUDA Support: Added support for Metal (Apple Silicon) and CUDA (NVidia) for AI/ML workloads including embeddings and local LLM inference.

For instructions on compiling a Meta or CUDA binary, see the Installation Docs.

Breaking Changesโ€‹

  • The ODBC Data Connector now requires ODBC drivers specified in connection strings are registered in the system ODBC driver manager.

Example invalid connection string:

DRIVER={/path/to/driver.so};SERVER=localhost;DATABASE=master

Example valid connection string:

DRIVER={My ODBC Driver};SERVER=localhost;DATABASE=master

Where My ODBC Driver is the name of an ODBC driver registered in the ODBC driver manager.

Contributorsโ€‹

  • @ewgenius
  • @peasee
  • @phillipleblanc
  • @sgrebnov
  • @Jeadie
  • @barracudarin
  • @Sevenannn

What's Changedโ€‹

- Update Helm for v0.19.4-beta and add release notes by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3310>
- Update spicepod.schema.json by @github-actions in <https://github.com/spiceai/spiceai/pull/3311>
- `metal` & `cuda` flags for spice by @Jeadie in <https://github.com/spiceai/spiceai/pull/3212>
- Promote postgres connector to RC quality by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3305>
- docs: Update ROADMAP.md by @peasee in <https://github.com/spiceai/spiceai/pull/3322>
- feat: Enable federation for in-memory accelerators by @peasee in <https://github.com/spiceai/spiceai/pull/3325>
- fix: Only allow env files from the current dir by @peasee in <https://github.com/spiceai/spiceai/pull/3327>
- Always read TimezoneTZ from PostgreSQL as UTC by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3330>
- For multi-sink acceleration refreshes, ensure parent table completes before the children. by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3329>
- Update TPC-DS Q49 (Decimal to Float) to match SQLite's type system by @sgrebnov in <https://github.com/spiceai/spiceai/pull/3323>
- Enable parquet pushdown in Spice by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3245>
- Use spice object_store fork to fix S3 ambiguous error by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3304>
- Don't mix commented out queries for s3 connectors and accelerators by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3331>
- Allow only valid WHERE conditions in vector searches by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3335>
- fix: Allow only ODBC profiles by @peasee in <https://github.com/spiceai/spiceai/pull/3324>
- Track how many times an acceleration falls back during initialization by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3339>
- Anthropic model regex and fix tool parsing aggregation bug by @Jeadie in <https://github.com/spiceai/spiceai/pull/3334>
- Upgrade runtime along with CLI on `spice upgrade` by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3341>
- Update upcoming Roadmap by @phillipleblanc in <https://github.com/spiceai/spiceai/pull/3343>
- fix: Prevent acceleration files outside of working directory by @peasee in <https://github.com/spiceai/spiceai/pull/3340>
- Document S3 connector limitations by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3333>
- Update Object Store Patch by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3361>
- Promote SQLite Data Accelerator to Beta by @sgrebnov in <https://github.com/spiceai/spiceai/pull/3365>
- Promote S3 connector to RC quality by @Sevenannn in <https://github.com/spiceai/spiceai/pull/3362>
- Revert "fix: Only allow env files from the current dir" by @peasee in <https://github.com/spiceai/spiceai/pull/3368>
- docs: Fix typo for S3 release status in README.md by @peasee in <https://github.com/spiceai/spiceai/pull/3370>
- Include unnecessary columns pruning step during federated plan creation by @sgrebnov in <https://github.com/spiceai/spiceai/pull/3363>

**Full Changelog**: <https://github.com/spiceai/spiceai/compare/v0.19.4-beta...v0.20.0-beta>

Resourcesโ€‹

Communityโ€‹

Spice.ai started with the vision to make AI easy for developers. We are building Spice.ai in the open and with the community. Reach out on Slack or by email to get involved.

Spice v0.18-beta (Sep 16, 2024)

ยท 7 min read
Sergei Grebnov
Member of Technical Staff at Spice AI

Announcing the release of Spice v0.18-beta.

The v0.18.0-beta release adds new Sharepoint and File data connectors, introduces AWS Identity and Access Management (IAM) support for the S3 Data Connector, improves performance of the GitHub connector, and increases the overall reliability of all data accelerators. The /ready API endpoint was enhanced to report as ready only when all components, including loaded data, have successfully reported readiness.

Highlights in v0.18.0-betaโ€‹

Sharepoint Data Connector: Use from: sharepoint: to access and accelerate documents stored in Microsoft 365 OneDrive for Business (Sharepoint). The CLI also includes a new spice login sharepoint to aid in local development and testing.

Example spicepod.yml:

datasets:
- from: sharepoint:drive:Documents/path:/important_documents/
name: important_documents
params:
sharepoint_client_id: ${secrets:SPICE_SHAREPOINT_CLIENT_ID}
sharepoint_tenant_id: ${secrets:SPICE_SHAREPOINT_TENANT_ID}
sharepoint_client_secret: ${secrets:SPICE_SHAREPOINT_CLIENT_SECRET}

See the Sharepoint Data Connector documentation.

AWS Identity and Access Management (IAM) for S3: A new s3_auth parameter for the s3 data connector to configure the authentication method to use when connecting to S3. Supported values are public, key, and iam_role. Use s3_auth: iam_role to assume the instance IAM role.

Example spicepod.yml:

datasets:
- from: s3://my-bucket
name: bucket
params:
s3_auth: iam_role # Assume IAM role of instance

See the S3 Data Connector documentation.

File Data Connector Use from: file: to query files stored by locally accessible filesystems.

Example spicepod.yml:

datasets:
- from: file://path/to/customer.parquet
name: customer
params:
file_format: parquet

See the File Data Connector documentation.

Improved /ready Api Now includes the initial data load for accelerated datasets in addition to component readiness to ensure readiness is only reported when data has loaded and can be successfully queried.

Breaking Changesโ€‹

  • GitHub Data Connector: The data type for time-related columns has changed from Utf8 to Timestamp. To upgrade, data type references to timestamp. For example, if using time_format:, change uses of time_format: ISO8601 to time_format: timestamp.

  • Ready API: The /ready API reports ready only when all components have reported ready and data is fully loaded. To upgrade, evaluate uses of the Ready API (such as Kubernetes readiness probes) and consider how it might affect system behavior.

Dependenciesโ€‹

No major dependencies updates.

Contributorsโ€‹

  • @phillipleblanc
  • @Jeadie
  • @lukekim
  • @sgrebnov
  • @peasee
  • @eltociear
  • @Sevenannn
  • @ewgenius
  • @karifabri

New Contributorsโ€‹

What's Changedโ€‹

- Update spicepod.schema.json by @github-actions in https://github.com/spiceai/spiceai/pull/2585
- Set helm to v0.17.4-beta by @ewgenius in https://github.com/spiceai/spiceai/pull/2595
- Bump to next v0.18.0-beta version by @ewgenius in https://github.com/spiceai/spiceai/pull/2596
- Add snapshot test docs / Update beta criteria for data accelerators by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2594
- Enable federation for accelerated queries (sqlite, duckdb, postgres) by @sgrebnov in https://github.com/spiceai/spiceai/pull/2598
- spelling updates on v0.17.4 release notes by @karifabri in https://github.com/spiceai/spiceai/pull/2601
- Update endgame template by @ewgenius in https://github.com/spiceai/spiceai/pull/2591
- fix: Re-attach DuckDB attachments on each query by @peasee in https://github.com/spiceai/spiceai/pull/2602
- Speed up sqlite accelerator benchmark test with indexes by @Sevenannn in https://github.com/spiceai/spiceai/pull/2597
- Fix refresh API using `refresh_mode: append` by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2609
- Tweak `/ready` to only report ready when components have all reported Ready by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2600
- Add `s3_auth` parameter to configure IAM role authentication by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2611
- Bump fundu from 2.0.0 to 2.0.1 by @dependabot in https://github.com/spiceai/spiceai/pull/2576
- fix: Remove comments from SQL files by @peasee in https://github.com/spiceai/spiceai/pull/2627
- Utilize runtime.status().is_ready() to check acceleration dataset readiness in benchmark test by @Sevenannn in https://github.com/spiceai/spiceai/pull/2614
- Allow for prefix to be kept in internal Parameters by @Jeadie in https://github.com/spiceai/spiceai/pull/2603
- Bump itertools from 0.12.1 to 0.13.0 by @dependabot in https://github.com/spiceai/spiceai/pull/2572
- Bump golang.org/x/mod from 0.20.0 to 0.21.0 by @dependabot in https://github.com/spiceai/spiceai/pull/2571
- Add initial threat model using OWASP Threat Dragon by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2599
- fix: Explicitly error for duplicate duckdb file accelerators by @peasee in https://github.com/spiceai/spiceai/pull/2628
- Benchmark test binary can parse command line option by @Sevenannn in https://github.com/spiceai/spiceai/pull/2626
- Snapshot tests shouldn't crash the Spice benchmark test by @Sevenannn in https://github.com/spiceai/spiceai/pull/2613
- Bump anyhow from 1.0.86 to 1.0.87 by @dependabot in https://github.com/spiceai/spiceai/pull/2573
- Upgrade datafusion to improve SQLite subquery tables aliasing support by @sgrebnov in https://github.com/spiceai/spiceai/pull/2634
- Run benchmark separately using workflow by @Sevenannn in https://github.com/spiceai/spiceai/pull/2631
- Sharepoint UX changes by @Jeadie in https://github.com/spiceai/spiceai/pull/2633
- Improve `/ready` to only mark a dataset ready iff the initial refresh completed by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2630
- Support relative paths for file connector by @Jeadie in https://github.com/spiceai/spiceai/pull/2637
- Fix `error decoding response body` GitHub file connector bug by @sgrebnov in https://github.com/spiceai/spiceai/pull/2645
- GraphQL pagination and robustness. by @Jeadie in https://github.com/spiceai/spiceai/pull/2632
- docs: Update bug template by @peasee in https://github.com/spiceai/spiceai/pull/2629
- Define GitHub `issues` data connector schema upfront by @sgrebnov in https://github.com/spiceai/spiceai/pull/2646
- Add support for loading from Sharepoint Group's default drive. by @Jeadie in https://github.com/spiceai/spiceai/pull/2642
- Fix typo in workflow, fix the postgres connector container readiness check by @Sevenannn in https://github.com/spiceai/spiceai/pull/2654
- Fix check all features by @Sevenannn in https://github.com/spiceai/spiceai/pull/2653
- Enable Warn/Error traces from dependency components by @sgrebnov in https://github.com/spiceai/spiceai/pull/2655
- Use lower case iso8601 for time_column by @Sevenannn in https://github.com/spiceai/spiceai/pull/2551
- Add basic integration test for Spice spill-to-disk and re-hydration scenario by @sgrebnov in https://github.com/spiceai/spiceai/pull/2643
- Add 'RefreshOverrides::max_jitter' to 'POST /v1/datasets/:name/acceleration/refresh' by @Jeadie in https://github.com/spiceai/spiceai/pull/2641
- Bump rustls-pemfile from 1.0.4 to 2.1.3 by @dependabot in https://github.com/spiceai/spiceai/pull/2575
- Update dependencies to support querying postgres enum types by @Sevenannn in https://github.com/spiceai/spiceai/pull/2657
- Upgrade table-providers by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2659
- Improve `spill_to_disk_and_rehydration` integration test by @sgrebnov in https://github.com/spiceai/spiceai/pull/2658
- Enhance GitHub connector robustness with explicit table schema definitions by @sgrebnov in https://github.com/spiceai/spiceai/pull/2661
- Rename sharepoint fields by @Jeadie in https://github.com/spiceai/spiceai/pull/2668
- Disable dataset checkpoint for DuckDB acceleration by @phillipleblanc in https://github.com/spiceai/spiceai/pull/2676
- Revert "Enable federation for accelerated queries (sqlite, duckdb, postgres) (#2598) by @Sevenannn in https://github.com/spiceai/spiceai/pull/2683

**Full Changelog**: https://github.com/spiceai/spiceai/compare/v0.17.4-beta...v0.18.0-beta

Resourcesโ€‹

Communityโ€‹

Spice.ai started with the vision to make AI easy for developers. We are building Spice.ai in the open and with the community. Reach out on Slack or by email to get involved.