Home / Chapter 4 · Context and knowledge
    Last edited · 9 min read

    Use with AI

    Embeddings and vector search

    An embedding model turns a whole passage of text into a single vector, so that texts with similar meaning get nearby vectors. Retrieval in RAG rests on this: the question becomes a vector too, and the database returns the passages with the nearest vectors, even when they share no words with it.

    In plain wordsA map where every passage of text has its own pin, and texts with similar content sit close together. Searching means pushing in a pin for the question and collecting its nearest neighbours. A contract number barely moves a pin, so contracts 48213 and 48231 sit almost on the same spot.

    Pick a question or type your own and see which passage each method puts on top

    The whole knowledge base: 10 passages

      A simulation, not a real model: an embedding model can’t run here. Each passage has hand-assigned weights for eight named concepts, and the words of the question take their weights from a small dictionary, so words outside it change nothing. A real vector has hundreds of unnamed dimensions. Cosine, BM25 (a simplified stemmer strips common endings) and RRF are computed live. The companies and contract numbers are made up.

      One vector for the whole passage

      Searching millions of vectors

      Weak spots, hybrid search and reranking

      Decisions when building the index

      Check yourself

      How does vector search work, and why is it not enough on its own?

      An embedding model turns a whole passage into one vector, trained so that texts with similar meaning land close together. The query is embedded with the same model, and search returns the nearest vectors by cosine similarity. At millions of passages an approximate index such as HNSW trades a little recall for milliseconds and costs memory, which quantisation reduces. Vectors catch paraphrases but miss identifiers, numbers, rare names and negation, and they always return something. So fuse them with BM25 via RRF, rerank the top candidates with a cross-encoder, and measure recall@k separately from answer quality.

      Po polsku

      Model embeddingowy zamienia cały fragment w jeden wektor, wytrenowany tak, by teksty o podobnym znaczeniu leżały blisko. Pytanie liczy się tym samym modelem, a wyszukiwanie zwraca najbliższe wektory po cosinusie. Przy milionach fragmentów indeks przybliżony, np. HNSW, oddaje ułamek recall za milisekundy i kosztuje pamięć, którą zmniejsza kwantyzacja. Wektory łapią parafrazy, ale gubią identyfikatory, liczby, rzadkie nazwy i negację, a do tego zawsze coś zwracają. Dlatego łącz je z BM25 przez RRF, czołówkę porządkuj cross-encoderem, a recall@k mierz osobno od jakości odpowiedzi.

      Follow-up questions (5)
      You are changing the embedding model. What happens to the index?
      You re-embed the whole corpus, because vectors from different models are not comparable, even with the same number of dimensions. You build the new index next to the old one, compare recall@k on the same question set, and only then switch traffic. That is why you store the source text and the model version with every vector.
      HNSW or IVF?
      HNSW gives high recall at low latency, as long as the vectors and the graph fit in RAM. IVF splits the vectors into clusters and searches a few of them. With PQ compression it fits far more vectors into the same memory, at the cost of recall, which you win back with a larger nprobe or by rescoring the top results on full vectors.
      How do you drop passages when the answer is not in the database?
      You tune a cosine threshold on your own test set, which also includes questions with no answer in the database, because the scale depends on the model. The reranker’s score is a more reliable signal. And the prompt explicitly allows the model to say “this is not in the sources”.
      How do you combine vector search with filters such as permissions or dates?
      A filter applied after the search cuts results out of the top k, so with a narrow filter too few are left, or none. It is better to filter while traversing the index, which many vector databases support, or to keep separate indexes, e.g. one per customer. Enforce permissions at retrieval time, not in the prompt.
      Recall@10 is high, but the answers are still weak. Where do you look?
      First, whether the relevant passage lands at ranks 8–10 while only the top three go into the prompt: a reranker helps there. Next, whether the passage, once cut out of its document, actually contains the answer (chunking). Finally, whether the model distorts it, which is a job for generation evals.

      Sources

      Report an error · Suggest a fix