Knowledgebase

Vector Storage and Retrieval Print

  • dataengineering, data, database, woocommerce, performance, permissions, guide, howto
  • 0

Storing embeddings.

WHAT A VECTOR STORE DOES

Holds embeddings and finds the most similar to a query embedding.

WHAT SIMILARITY SEARCH IS

Finding nearest neighbours in the representation space.

WHAT EXACT SEARCH COSTS

Comparison against everything, which does not scale.

WHAT APPROXIMATE SEARCH PROVIDES

Near-nearest results, far faster.

WHAT IT TRADES

A small proportion of correct results, for large speed gains.

WHAT THE OPTIONS ARE

Extensions to existing databases Purpose-built vector databases Libraries embedded in an application

WHAT TO START WITH

An extension to a database you already run.

WHY

It avoids another system, and suffices at modest scale.

WHAT TO ESTABLISH BEFORE ADOPTING A SEPARATE SYSTEM

The actual corpus size and query rate.

WHAT FILTERING MATTERS FOR

Restricting results by attributes: permissions, date, source.

WHY IT IS A KEY SELECTION CRITERION

Filtering applied after retrieval returns too few results.

WHAT TO PREFER

Filtering applied during search.

WHAT HYBRID RETRIEVAL COMBINES

Keyword matching with similarity search.

WHY IT PERFORMS BETTER

Keyword search catches exact terms, identifiers and rare words that embeddings miss.

WHAT TO MEASURE

Whether the right passage appears in the results.


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot