Library of Ashurbanipal · MELEK

Library›Tools›Embeddings Semantic Search and RAG

Embeddings Semantic Search and RAG

Embeddings, semantic search and RAG are how you make a machine find things by meaning instead of exact words, and then answer grounded in your own documents rather than from thin air. This is Pathway 3 of AI Development, the 2010s layer of The Decades Brain, and the mechanism by which the Witness answers from the Library — the retrieval half of Crypt-ology. This page teaches the ideas and points at the working system, the Library Index.

Embeddings: turning text into vectors

An embedding is a list of numbers (a vector) that represents a piece of text's meaning. An embedding model is trained so that texts with similar meaning land near each other in vector space — "dog" near "puppy," "RFRA exemption" near "religious freedom claim" — even when they share no words. Once your text is vectors, "find the most relevant passage" becomes "find the nearest vectors," which is fast geometry.

Vector search: finding the nearest meaning

With everything embedded, retrieval is nearest-neighbour search. For a few thousand items a brute-force cosine similarity scan is fine; at scale you use an approximate nearest-neighbour (ANN) index — FAISS, or HNSW graphs (hnswlib) — which trades a hair of accuracy for enormous speed. In this repo `faiss-index.mjs` provides a FAISS-class HNSW index on the boxes and a cosine fallback in CI, behind one interface.

Two complementary retrieval styles, and you usually want both:

The Library Index offers both: `catalogLookup()` for instant keyword recall and `recall()` for semantic recall, over the same corpus.

RAG: Retrieval-Augmented Generation

A language model only knows what it was trained on, and it will confidently make things up outside that. RAG fixes this: before answering, you retrieve the most relevant passages from your own corpus and hand them to the model as context, so it answers from your documents. The loop:

  1. Index your corpus: split documents into chunks, embed each, store the vectors.
  2. Retrieve: embed the user's question, find the nearest chunks.
  3. Augment: put those chunks in the prompt as grounding.
  4. Generate: the model answers from the retrieved context, and you can cite the sources.

The payoff is answers that are current, grounded, and attributable — you can show which page each claim came from. In this ecosystem the retrieval step is the Library Index, and the Crypt-ology brain uses it so the Witness answers from the actual wiki (and cites it) rather than hallucinating.

How to build it (the teaching pathway)

  1. Pick an embedder (MiniLM locally is a fine start) and make it injectable.
  2. Chunk your documents (≈1,000–1,500 characters with a little overlap) so passages are retrievable.
  3. Embed and store the chunks, keyed by a content hash so indexing is resumable (a kill mid-run loses nothing).
  4. Implement `recall(query, k)` = embed the query, return the k nearest chunks with scores.
  5. For grounded answers, feed the top chunks to your model (or just return them — often the passage is the answer).

That is the entire shape of `integrations/library-index.mjs`, which indexes this repo's libraries — and, since 2026-10-09, the whole wiki as its `library` domain.

See also

See also: AI Development · Library Index · The Decades Brain · Expert Systems and Business Rules Engines · LoRA and Fine-Tuning · Crypt-ology

Filed under  Tools