Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions docs/best_practices.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,3 +34,14 @@ Effective data management ensures clean operations and provides flexibility when
- Use the `reprocess_chunks()` function to queue existing chunks that are missing embeddings.
- Use the `recreate_chunks()` function for a complete chunk regeneration, which deletes all existing chunks first.
- Each column gets independent chunk tables and triggers, so you can disable them selectively as needed.
- Settle on an embedding model before enabling vectorization, because the
model's dimension is fixed into the chunk table when the table is
created.
- Rebuild the vectorizer with `disable_vectorization(...,
drop_chunk_table => TRUE)` and then `enable_vectorization()` after
changing to a model of a different dimension, repeating this for every
vectorized column. Neither `recreate_chunks()` nor a
`disable_vectorization()` that keeps the chunk table alters the column,
so neither resolves the mismatch.
- Budget for the provider cost of re-embedding an entire table before
changing the model on a populated one.
14 changes: 14 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,20 @@ These settings configure the connection to your embedding provider, including th
| `pgedge_vectorizer.model` | `text-embedding-3-small` | Model name | No | No | No |
| `pgedge_vectorizer.extra_headers` | (empty) | Semicolon-separated `key: value` HTTP headers added to all API requests | No | No | No |

!!! warning "The model fixes the vector dimension of a chunk table"

The model determines how many dimensions the provider returns, and
`enable_vectorization()` fixes that number into the chunk table's
`embedding vector(N)` column when the table is created. Changing
`pgedge_vectorizer.model` to a model with a different dimension does
not migrate an existing chunk table. The background worker compares
the dimensions before writing, so nothing is corrupted, but it marks
the affected queue items `failed` with the message
`Dimension mismatch: model=N, table=M` and no new embeddings are
stored for that table until you act. The
[Troubleshooting](troubleshooting.md) document describes how to
recover.

## Worker Settings

These settings control the background workers that process the embedding queue, including concurrency, batch sizes, and retry behavior.
Expand Down
92 changes: 91 additions & 1 deletion docs/troubleshooting.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,13 @@ SHOW shared_preload_libraries;

## Workers Not Processing After CREATE EXTENSION

Background workers start with PostgreSQL — before the extension is created. Workers automatically detect the extension using exponential backoff (checking every 5s, 10s, 20s, ... up to 5 minutes). After running `CREATE EXTENSION pgedge_vectorizer`, workers should discover it within seconds.
Background workers start with PostgreSQL, before the extension is
created. A worker checks for the extension on an exponential backoff,
waiting 5s, then 10s, then 20s, and so on up to a ceiling of 5 minutes.
After running `CREATE EXTENSION pgedge_vectorizer`, a worker discovers it
on its next check, so a database configured long before the extension was
created can wait up to 5 minutes. Reload the configuration to have the
workers check immediately.

If workers don't start processing:

Expand Down Expand Up @@ -64,3 +70,87 @@ SELECT * FROM pgedge_vectorizer.failed_items;
```sql
SELECT pgedge_vectorizer.retry_failed();
```

## Dimension Mismatch After Changing the Model

Each chunk table stores its vectors in an `embedding vector(N)` column,
where N is fixed when `enable_vectorization()` creates the table. If you
change `pgedge_vectorizer.model` to a model returning a different number
of dimensions, the worker cannot write the new vectors into the existing
column, and embeddings stop being produced for that table.

Nothing is corrupted when this happens. The worker compares the two
dimensions before it writes, so the existing embeddings are left intact
and no vector of the wrong size is ever stored.

The affected queue items move to `failed` rather than being retried,
because retrying cannot succeed. Run the following query to identify
them:

```sql
SELECT chunk_table, error_message, count(*)
FROM pgedge_vectorizer.queue
WHERE status = 'failed'
GROUP BY chunk_table, error_message;
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

An affected item reports `Dimension mismatch: model=N, table=M`, where N
is the dimension the configured model returned and M is the dimension
the chunk table expects. The server log carries a matching warning that
names the table.

Restoring the previous model is the quicker of the two remedies, and is
the right one if the change was accidental. Substitute the model that
built the table rather than the name shown here, because setting any
other model leaves the dimensions mismatched:

```sql
ALTER SYSTEM SET pgedge_vectorizer.model = 'the-previous-model';
SELECT pg_reload_conf();
SELECT pgedge_vectorizer.retry_failed();
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

The `table=M` figure in the error gives the dimension the chunk table
expects, so the model you restore must be one that returns M dimensions.

Rebuilding the vectorizer keeps the new model and re-embeds the table
under it. Note that `recreate_chunks()` does not resolve a dimension
change, because that function deletes the rows of a chunk table without
altering the type of the column. Follow these steps instead:

1. Set the new model and reload the configuration so that the dimension
detection uses the model you want.

```sql
ALTER SYSTEM SET pgedge_vectorizer.model = 'text-embedding-3-large';
SELECT pg_reload_conf();
```

2. Drop the vectorizer together with its chunk table, which discards the
embeddings of the old dimension.

```sql
SELECT pgedge_vectorizer.disable_vectorization(
'docs', 'body', drop_chunk_table => TRUE);
```

3. Enable vectorization again, which detects the new dimension, creates
the chunk table to match, and queues every source row.

```sql
SELECT pgedge_vectorizer.enable_vectorization('docs', 'body');
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

Repeat those steps for every vectorized column, not just the one you
noticed. The model is a single global setting while chunk tables are
independent per column, so changing it affects every vectorizer whose
dimension no longer matches. The following query lists them:

```sql
SELECT source_table, source_column, chunk_table
FROM pgedge_vectorizer.vectorizers
ORDER BY source_table, source_column;
```

Re-embedding calls the provider for every chunk in every table rebuilt,
so confirm the cost against your provider's pricing before starting.
Loading