Skip to content
4 changes: 2 additions & 2 deletions apps/web/content/docs/core/filters.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ title: Filters
description: Reference for the filter builders and the native syntax each adapter compiles them to.
---

You build a filter once with the exported helpers. Each adapter compiles the same filter to its provider's native syntax. The compilers are total: every filter compiles on every provider.
You build a filter once with the exported helpers. Each adapter compiles the same filter to its provider's native syntax. Five compilers accept every filter. The Vectorize and Redis compilers return a `Result`, because Vectorize has no `or` and no `exists`, and Redis matches only fields its index schema declares.

```ts
import { and, eq, gt, isIn, not, or } from "vecstore-sdk";
Expand Down Expand Up @@ -85,4 +85,4 @@ Supabase has no compiler to import. The filter travels to Postgres as JSON and `

`compileRedisFilter` takes a second argument, the metadata fields the index schema declares, because a Redis query names a field the schema has to hold.

Where a provider lacks an operator, the compiler rewrites the filter instead of failing. Three providers are exceptions. Upstash's filter is a string, so the compiler rejects a field name or a string value it cannot write safely and the verb returns an `invalid_argument` error. Vectorize has no OR and no presence test, so `compileVectorizeFilter` returns a `Result` and the verb returns an `unsupported` error before it sends a request. Redis matches nothing on a field its schema does not declare, so `compileRedisFilter` returns a `Result` and the verb returns an `invalid_argument` error instead of an empty page. See [Provider differences](/docs/providers/differences) for the two cases where the semantics diverge.
Where a provider lacks an operator, the compiler rewrites the filter instead of failing. Three providers are exceptions. Upstash's filter is a string, so the compiler rejects a field name or a string value it cannot write safely and the verb returns an `invalid_argument` error. Vectorize has no OR and no presence test, so `compileVectorizeFilter` returns a `Result` and the verb returns an `unsupported` error before it sends a request, or `invalid_argument` when two clauses set the same operator on one field, or mix a membership and a comparison operator on one field. Redis matches nothing on a field its schema does not declare, so `compileRedisFilter` returns a `Result` and the verb returns an `invalid_argument` error instead of an empty page. See [Provider differences](/docs/providers/differences) for the two cases where the semantics diverge.
2 changes: 1 addition & 1 deletion apps/web/content/docs/core/store.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ type MetadataValue = string | number | boolean | string[];
type Metadata = Readonly<Record<string, MetadataValue>>;
```

Metadata values are `string`, `number`, `boolean`, or `string[]`. That is the intersection of what the seven providers accept, so a record that upserts on one provider upserts on all of them.
Metadata values are `string`, `number`, `boolean`, or `string[]`. That is the intersection of what the seven providers accept. Each adapter adds a few rules of its own. Qdrant, Vectorize, and Upstash in metadata mode reject metadata under `_id` or `_namespace` and a namespace that holds `/`. Redis rejects a value whose type contradicts the declared field type, and an index name that holds `:`.

## Scores

Expand Down
2 changes: 1 addition & 1 deletion apps/web/content/docs/guides/design.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ Three constraints applied:

**Compilers.** One pure function per provider from `Filter` to the provider's native filter type. Five of the seven are total on the operators: no filter fails to compile for want of one. Qdrant has no float `match`, so `eq` on a float becomes a closed `range`. Pinecone has no `$not`, so the compiler pushes negation to the leaves with De Morgan's laws. Pinecone's `$in` rejects booleans, so those expand to `$or` of `$eq`. Upstash has no `NOT` either and takes the same pushdown. Upstash can also reject its input, because its filter is a string: a field name or a string value its grammar cannot hold safely throws. Vectorize is the one provider whose filter language is smaller than the AST. Its filter is a flat object of fields joined with AND, so `or` has no rewrite and `exists` has no operator, and `compileVectorizeFilter` returns a `Result` instead. Redis expresses every operator, but a Redis query names a field its index schema has to declare, so `compileRedisFilter` takes the declared fields as a second argument and returns a `Result` too. A filter on a field Redis has no index entry for fails at compile time rather than coming back as an empty page. Supabase has no compiler at all. Its filter is already JSON, so it travels as data and `vecstore_filter_sql` emits the pgvector predicates inside Postgres.

**Adapters.** Most adapters take a client that satisfies a structural `*ClientLike` interface, `RedisClientLike` included: it names the eight `ft`, `json`, and `unlink` calls the adapter makes, so a `node-redis` client and a cluster client both fit. The Vectorize adapter names the `cloudflare` type instead, because two of the responses it needs are typed `unknown` there and a structural interface would have to restate that `unknown` in its own signatures. The store is generic over the client type either way, so `raw` keeps the caller's concrete type. Every verb runs inside one `run` helper that converts a thrown SDK error to a `VecstoreError` and returns a `Result`.
**Adapters.** Most adapters take a client that satisfies a structural `*ClientLike` interface, `RedisClientLike` included: it names the seven `ft`, `json`, and `unlink` calls the adapter makes, so a `node-redis` client and a cluster client both fit. The Vectorize adapter names the `cloudflare` type instead, because two of the responses it needs are typed `unknown` there and a structural interface would have to restate that `unknown` in its own signatures. The store is generic over the client type either way, so `raw` keeps the caller's concrete type. Every verb runs inside one `run` helper that converts a thrown SDK error to a `VecstoreError` and returns a `Result`.

**Emulation.** Qdrant gets namespaces through a `_namespace` payload key with a tenant index, and arbitrary ids through a deterministic UUID. Vectorize has native namespaces but scopes ids to the whole index and caps them at 64 bytes, so it borrows the same id hashing. pgvector and Supabase get namespaces through a column in the primary key. Upstash spends its one namespace level on the index name, so by default it keeps the namespace in a `_namespace` metadata key and prefixes stored ids with it. `namespaceMode: "native"` spends a real Upstash namespace per index and namespace pair instead. Redis puts the namespace in a tag on the document and in the key, which leaves the id as the last segment of the key and the metadata untouched. All of them round-trip exactly.

Expand Down
6 changes: 3 additions & 3 deletions apps/web/content/docs/guides/migration.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -68,9 +68,9 @@ Moving to Supabase needs a setup step the others do not: install `sql/supabase.s

## Copy the data

The SDK does not move records for you. Read from the old index with `fetch` or your own export, then `upsert` into the new one. The adapter batches upserts, so pass as many records per call as fit in memory.
The SDK does not move records for you. The SDK has no verb that lists records, so export ids with the provider's own scroll or list call through `store.raw` (Qdrant `scroll`, Pinecone `listPaginated`, Upstash `range`, a keyset `SELECT` on pgvector or Supabase, a paged `FT.SEARCH` on Redis, `listVectors` on Vectorize), then `fetch` them in pages of 1,000 and `upsert` into the new index. The adapter batches upserts, so pass as many records per call as fit in memory.

Record ids, vectors, and metadata round-trip unchanged. Every provider accepts only `string`, `number`, `boolean`, and `string[]` metadata values, so a record that upserts on one provider upserts on all of them.
Record ids, vectors, and metadata round-trip unchanged. Every provider accepts only `string`, `number`, `boolean`, and `string[]` metadata values. Each adapter adds a few rules of its own. Qdrant, Vectorize, and Upstash in metadata mode reject metadata under `_id` or `_namespace` and a namespace that holds `/`. Redis rejects a value whose type contradicts the declared field type, and an index name that holds `:`.

## Check what each provider stores

Expand Down Expand Up @@ -108,4 +108,4 @@ Before you cut over, run the live conformance suite against the new backend:
VECSTORE_LIVE=1 QDRANT_URL=http://localhost:6333 bun run test:live
```

The suite creates a temporary index, exercises every verb and filter operator, and deletes the index.
The suite creates a temporary index, runs the store verbs, the record verbs, and the `eq`, `gt`, `isIn`, `notIn`, and `not` builders, and deletes the index. The pgvector and Supabase adapters also run on PGlite in the unit suite.
2 changes: 1 addition & 1 deletion apps/web/content/docs/guides/testing.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,6 @@ REDIS_URL=redis://localhost:6379 \
bun run test:live
```

The suite skips providers without a variable. For each provider it creates a temporary index, exercises every verb and filter operator, and deletes the index. `QDRANT_API_KEY` is optional for a local Qdrant and required for Qdrant Cloud.
The suite skips providers without a variable. For each provider it creates a temporary index, runs the store verbs, the record verbs, and the `eq`, `gt`, `isIn`, `notIn`, and `not` builders, and deletes the index. The pgvector and Supabase adapters also run on PGlite in the unit suite. `QDRANT_API_KEY` is optional for a local Qdrant and required for Qdrant Cloud.

Three providers ask for more. Supabase needs the SQL functions installed before the suite runs, and the service role key because the suite creates and drops tables. Upstash cannot create an index at all, so the adapter creates a namespace inside the index your URL and token point at. Point them at a scratch index with dimension 3 and the cosine similarity function, which is what the conformance suite asks for. The Vectorize token needs the Vectorize permission, and that run skips the `delete({ all: true })` case, because Vectorize has no call that empties a namespace. `REDIS_URL` has to point at a server that carries the query engine and JSON, such as Redis 8 or Redis Stack; the Redis run adds cases for the default namespace, `exists`, list fields, and delete by filter.
2 changes: 1 addition & 1 deletion apps/web/content/docs/providers/pinecone.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ The adapter needs `@pinecone-database/pinecone` 7 or later, which introduced the

## Delete by filter

Serverless indexes reject `delete({ filter })`. The adapter returns an `unsupported` error with `feature: "deleteByFilter"`. To delete matching records, delete by ids instead, or query first and delete the returned ids.
Serverless indexes reject `delete({ filter })`. The adapter returns an `unsupported` error with `feature: "deleteByFilter"`. To delete matching records, delete by ids instead, or query first and delete the returned ids, repeating until a query returns nothing, since one query returns at most `topK` matches.

## What the adapter stores

Expand Down
4 changes: 3 additions & 1 deletion apps/web/content/docs/providers/qdrant.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,9 @@ bun add vecstore-sdk @qdrant/js-client-rest

## Ids

Qdrant accepts only UUID (universally unique identifier) or integer point ids. The adapter hashes any other id to a deterministic UUID and stores your id in the `_id` payload key. A lowercase UUID id in the default namespace passes through unchanged, unless it is a version 8 UUID, the version the adapter's own hashed ids use; those are hashed too, so a raw id can never collide with a hashed one.
Qdrant accepts only UUID (universally unique identifier) or integer point ids. The adapter hashes any other id to a deterministic UUID and stores your id in the `_id` payload key. A lowercase UUID id in the default namespace passes through unchanged, unless it is a version 8 UUID, the version the adapter's own hashed ids use. Those are hashed too, so a raw id can never collide with a hashed one.

A namespace cannot contain `/`, because the adapter joins the namespace and the id when it hashes a point id, and `upsert` rejects a record whose metadata sets `_id` or `_namespace`. Both return `invalid_argument`. `upsert` sends 500 points per request, and `fetch` and `delete({ ids })` send 1,000 ids per request.

## Namespaces

Expand Down
2 changes: 2 additions & 0 deletions apps/web/content/docs/providers/redis.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -78,6 +78,8 @@ A record is one JSON document:
| Vector | A JSON array of numbers, stored at full precision and indexed as `FLOAT32` |
| Metric | `COSINE`, `L2`, or `IP`, fixed when the index is created |

An index name cannot contain `:`, because the adapter joins the key prefix, the index name, and the namespace with it. `createIndex` and every verb return `invalid_argument` for one that does. `upsert` writes 500 documents per batch, one batch after another.

Pass `keyPrefix` to move the keyspace off `vecstore:`, and `algorithm: "HNSW"` to trade exact search for scale. The default, `FLAT`, searches every vector and is what Redis recommends below a million of them.

Because ids live in the key rather than in a reserved metadata field, nothing is hidden from the metadata you read back.
Expand Down
7 changes: 4 additions & 3 deletions apps/web/content/docs/providers/upstash.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -59,8 +59,9 @@ Choose `"native"` when you can name your namespaces up front, and `"metadata"` w
- Records carry the namespace in the `_namespace` metadata key, and the adapter adds that key to every query and filtered delete.
- The adapter stores ids as `{namespace}/{length}/{id}`, so two namespaces can hold the same id. The original id goes in the `_id` metadata key and comes back on `query` and `fetch`.
- Records in the default namespace keep their id and carry neither key.
- An id in the default namespace cannot have the shape `{namespace}/{length}/{id}` with a matching length, because it would collide with a stored namespaced record; `upsert` returns `invalid_argument` for one.
- The adapter strips `_id` and `_namespace` from returned metadata. Do not write metadata under those names.
- An id in the default namespace cannot have the shape `{namespace}/{length}/{id}` with a matching length, because it would collide with a stored namespaced record. `upsert` returns `invalid_argument` for one.
- The adapter strips `_id` and `_namespace` from returned metadata. `upsert` returns `invalid_argument` for a record whose metadata sets either key.
- A namespace cannot contain `/` in this mode, because the adapter joins the namespace and the id with it. The verbs return `invalid_argument` for one that does.

One Upstash namespace holds every vecstore namespace under that index, so the number of tenants you can have is unbounded.

Expand Down Expand Up @@ -102,7 +103,7 @@ Because the filter is a string, the compiler checks every part it writes into th

## Score normalization

Upstash normalizes every score to the range 0 to 1. Higher is always more similar, including for euclidean distance. The other three adapters return the provider's raw score instead. Cosine returns `(1 + cosine_similarity) / 2` and euclidean returns `1 / (1 + squared_distance)`.
Upstash normalizes every score to the range 0 to 1. Higher is always more similar, including for euclidean distance. Qdrant, pgvector, Supabase, Pinecone, and Vectorize return the provider's raw score, and Redis returns a distance. Cosine returns `(1 + cosine_similarity) / 2` and euclidean returns `1 / (1 + squared_distance)`.

## What the adapter stores

Expand Down
26 changes: 17 additions & 9 deletions apps/web/content/docs/providers/vectorize.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ Vectorize splits its control plane from its data plane. Creating, listing, and d

The adapter takes the HTTP client, because that is the only shape where the whole store contract works. A binding has no call that creates, lists, or drops an index, no way to enumerate namespaces, and no way to list vectors, so `createIndex`, `deleteIndex`, and `listIndexes` would all be unsupported and `index(name)` would name nothing.

The HTTP API works anywhere `fetch` does, including inside a Worker. Reaching your own account from a Worker costs a subrequest and an API token, so keep the binding for hot read paths and use this adapter where the store contract matters.
The HTTP API works anywhere `fetch` does, including inside a Worker. The adapter also imports `node:crypto` to hash ids, so a Worker needs the `nodejs_compat` compatibility flag. Reaching your own account from a Worker costs a subrequest and an API token, so keep the binding for hot read paths and use this adapter where the store contract matters.

## Metadata indexes

Expand All @@ -60,7 +60,9 @@ A Vectorize id is unique across the whole index rather than within a namespace,
| In any other namespace | A deterministic UUID, original in `_id` |
| Id longer than 64 bytes | A deterministic UUID, original in `_id` |

The same id in the same namespace always hashes to the same UUID, so `upsert` still replaces. `query` and `fetch` read `_id` back and strip it from the metadata they return. Do not write metadata under that name.
A namespace cannot contain `/`, because the adapter joins the namespace and the id when it hashes them. The verbs return `invalid_argument` for one that does.

The same id in the same namespace always hashes to the same UUID, so `upsert` still replaces. `query` and `fetch` read `_id` back and strip it from the metadata they return. `upsert` returns `invalid_argument` for a record whose metadata sets `_id`.

## What Vectorize cannot do

Expand All @@ -73,16 +75,22 @@ Four things return `{ kind: "unsupported" }` with a `feature` you can match on:
| `orFilter` | `or(...)`, and `not` over `and(...)` | A Vectorize filter joins every clause with AND. |
| `existsFilter` | `exists(...)` | Vectorize has no operator for a missing property. |

To clear records without a delete-by-filter, query for the ids you want gone and delete those:
To clear records without a delete-by-filter, query for the ids you want gone and delete those in a loop:

```ts
const matches = await docs.query({ vector, topK: 50, filter });

if (matches.ok && matches.value.length > 0) {
await docs.delete({ ids: matches.value.map((match) => match.id) });
}
const clear = async () => {
for (;;) {
const matches = await docs.query({ vector, topK: 50, filter });
if (!matches.ok || matches.value.length === 0) {
return matches;
}
await docs.delete({ ids: matches.value.map((match) => match.id) });
}
};
```

Vectorize caps a query that returns metadata at 50 matches and applies writes asynchronously, so loop until a query returns nothing, and expect a short delay before deleted records stop appearing.

To drop everything, delete the index and create it again.

## Filter compilation
Expand All @@ -100,7 +108,7 @@ compileVectorizeFilter(
// { genre: { $eq: "drama" }, year: { $gt: 2000, $lt: 2010 } }
```

It is the one compiler that returns a `Result`, because Vectorize is the one provider whose filter language is smaller than the filter AST. `not` is pushed to the leaves with De Morgan's laws, the same way the Pinecone and Upstash compilers do it, so `not(eq(...))` becomes `$ne` and `not(or(a, b))` becomes two AND clauses. What is left over fails:
It is one of two compilers that return a `Result` (the Redis compiler is the other), because Vectorize is the one provider whose filter language is smaller than the filter AST. `not` is pushed to the leaves with De Morgan's laws, the same way the Pinecone and Upstash compilers do it, so `not(eq(...))` becomes `$ne` and `not(or(a, b))` becomes two AND clauses. What is left over fails:

- `or(...)`, and any `not` that De Morgan turns into an OR, return `unsupported` with `feature: "orFilter"`.
- `exists(...)` returns `unsupported` with `feature: "existsFilter"`.
Expand Down
2 changes: 1 addition & 1 deletion apps/web/lib/landing-content.ts
Original file line number Diff line number Diff line change
Expand Up @@ -387,7 +387,7 @@ export const highlights = [
title: "Errors as values, never thrown.",
},
{
body: "Pinecone, Upstash, and Vectorize have them. Qdrant, pgvector, Supabase, and Redis get them emulated with the same API.",
body: "Pinecone and Vectorize have them. Qdrant, pgvector, Supabase, Upstash, and Redis get them emulated with the same API.",
title: "Namespaces on every provider.",
},
] as const;
Expand Down
Loading
Loading