Most teams meet vector databases through a retrieval-augmented chatbot tutorial: embed some documents, store the vectors, search by similarity, paste the results into a prompt. The tutorial works, and then the real questions start. Which results is the user allowed to see? Which policy version is current? Why did the right document come fourth? Those are database questions, and they decide what to use far more than raw vector speed does.

What a vector is doing for you

An embedding model turns a piece of text into a list of numbers, say 768 of them, arranged so that texts with similar meaning land close together. "Refund within 14 days" and "money back if you cancel in two weeks" end up near each other even though they share almost no words. Similarity search finds the stored vectors closest to the query's vector, usually by cosine distance.

Two things follow. Similarity is only as good as the embedding model's idea of "similar", which was learned from someone else's data. And the nearest vectors are the most similar texts, which is a different property from being the right ones. A superseded policy can be the best match in the collection.

Exact search is often enough

With a few thousand chunks you don't need an index at all. Compute the distance to every vector and sort. A modern laptop does this for tens of thousands of 768-dimensional vectors in milliseconds, with perfect recall. Plenty of internal tools never outgrow this.

Approximate nearest-neighbour (ANN) indexes exist because exact search grows linearly with the collection. They trade a little recall for a lot of speed:

  • HNSW builds a layered graph of neighbours and walks it greedily from a coarse layer to a fine one. Queries are fast and recall is high; building and memory are the costs.
  • IVF clusters the vectors and searches only the clusters nearest to the query. It builds faster and uses less memory, and it's more sensitive to how many clusters you probe.

Either way, "approximate" means some true nearest neighbours are missed. How many depends on settings you choose, and you should measure it on your own data instead of trusting a benchmark chart.

Filters are where it gets hard

Real queries almost always carry conditions: this tenant, this language, this product, documents the user may read, policies in effect today. The naive order is to find the 40 nearest vectors and then drop the ones that fail the filter. If only 10% of the collection passes, you end up with about four results, and possibly none of the right ones.

pgvector's documentation spells out exactly this case: with HNSW's default hnsw.ef_search of 40 and a filter matching 10% of rows, about four rows come back on average. Since version 0.8.0 it supports iterative index scans, which keep scanning until enough filtered rows are found, in a strict or relaxed ordering mode (pgvector 0.8.0 release notes summarize the feature; the pgvector README documents the settings). Dedicated vector databases handle filtering in their own ways. Whatever you pick, test your most selective filter, not your average one, because the users behind the narrowest permissions are the ones who get the worst answers.

Metadata is half the retrieval system

For every chunk, store more than the text and the vector:

  • source document, version and section;
  • status (draft, approved, superseded) and effective dates;
  • who may see it (tenant, role, region);
  • a content hash, so re-indexing unchanged text is a no-op.

Then treat eligibility as a hard filter and similarity as ranking within the eligible set. In a support-drafting design I wrote up, that single change removes the failure where an assistant quotes a well-matched but superseded refund policy (the design study walks through it).

pgvector or a dedicated vector database

If your application already runs on Postgres, pgvector keeps vectors next to the rows they describe. Filters are ordinary WHERE clauses, updates are transactional with the metadata, and backups, permissions and monitoring already exist. For many AI features that's the deciding argument, and the same database can also run the job queue and outbox that keeps the index in step. A policy approval and the vector index update can happen in one transaction, so there's no window where the index disagrees with the database.

A dedicated vector database earns its place when the collection is very large, when vector queries dominate the load and need their own scaling, or when you need features such as built-in hybrid search or multi-tenancy at a scale your Postgres setup can't handle. Those are real needs. They're rarely the first need.

Retrieval needs its own evaluation

Before tuning index parameters, build a small test set: 50 real questions, each with the chunk IDs a person says should be retrieved. Measure recall at k (did the right chunks appear in the top k?) for exact search, then for your ANN index, then with filters. The gaps between those three numbers tell you whether to fix the embedding model, the index settings or the filtering. Add a reranker (a small cross-encoder that scores each query and passage pair) when the right chunks are retrieved but ranked too low.

A starting checklist

  1. Start with exact search if you have fewer than tens of thousands of chunks.
  2. Store eligibility metadata with every chunk and filter before ranking.
  3. Test the most selective filter your users will hit.
  4. Keep vectors in the same database as the data they describe until you have a measured reason not to.
  5. Measure recall on your own questions before and after every change.

Vectors find text that sounds similar. Everything that decides whether that text may be used is ordinary data modelling, and it's worth doing first.

The practical next step

Map one real execution and one failure.

That will reveal more about the right architecture than a tool comparison or model demo.

Let's build something real