← All posts

Engineering · AI

RAG vs Long Context: Do You Actually Need a Vector Database?

Mihajlo Petrović5 min read

Million-token context windows and prompt caching changed the maths. How to tell whether your corpus needs a real retrieval pipeline or just needs to be pasted into the prompt - and how to build retrieval properly if it does.

Every "chat with your documents" tutorial follows the same script: chunk the documents, embed the chunks, store them in a vector database, retrieve the top-k, stuff them into the prompt. It's presented as the architecture.

For a lot of real products, it's an unnecessary distributed system standing between you and a working feature.

Context windows around a million tokens changed the calculus. Here's how I decide now.


First, Measure the Corpus

Before any architecture discussion, answer one question: how many tokens is the entire body of knowledge?

Not "how many documents" — tokens. Use the provider's token-counting endpoint rather than guessing from word counts, then compare:

Corpus size Approach
Fits comfortably in context Put it all in the prompt. No retrieval.
A few times the window Cheap filtering (metadata, folders, dates), then stuff
Much larger, or growing daily Real retrieval

An internal handbook, a product's docs, a legal policy set, one codebase's public API — these are often smaller than people assume. A 200-page manual is perhaps 100k tokens. That fits.

If your corpus fits, retrieval is not an optimisation. It's a lossy filter you've added between the model and the answer, plus an embedding pipeline, plus a database, plus a re-index job, plus a whole class of bugs where the right chunk simply wasn't retrieved.


The Objection: "Isn't That Expensive?"

It would be, if you re-sent the whole corpus fresh on every request. That's what prompt caching is for.

Put the stable corpus at the front of the prompt, mark it as cacheable, and put the varying question at the end. Repeat requests read the cached prefix at a fraction of normal input cost. The economics of "just stuff it in" changed more than most architecture posts have caught up with.

Two rules make or break it:

  1. The cached prefix must be byte-identical. A timestamp, an unsorted JSON key order, a differently-ordered tool list — any of these silently invalidate the cache. Nothing errors. Your bill just goes up.
  2. Order is stable-first. Tools, then system prompt, then messages; corpus before question.

Check the cache-read token count in the response to confirm it's actually working. Assuming is how people conclude "caching didn't help".


When You Genuinely Need Retrieval

Not a small list — these cases are real and common:

The corpus is far bigger than the window. Every ticket you've ever closed. All of Confluence. A monorepo. No amount of caching fixes physics.

Freshness matters per-request. Prices, inventory, account state. You don't want a cached snapshot; you want a lookup at request time.

You need citations you can trust. Retrieval gives you provenance for free: you know which chunks went in, so "according to §4.2" is verifiable rather than plausible.

Per-user access control. Different users may see different documents. Filtering at retrieval time is a clean boundary; slicing a shared cached prefix per user is not.

Latency and cost at real volume. At thousands of requests an hour, sending 100k tokens each time is measurably slower and more expensive than sending 4k well-chosen ones — cache or no cache.


If You Do Build Retrieval, Build It Properly

The naive pipeline underperforms, and then people blame the model.

Chunk on structure, not character count. Split on headings and sections. A chunk that starts mid-sentence and ends mid-table is a chunk that retrieves badly and reads worse.

Hybrid search beats pure vectors. Keyword search (BM25) and embeddings fail differently: vectors handle "how do I cancel" → "termination procedure"; keywords handle exact identifiers, error codes, product names, people's names. Run both, merge the results. This is usually the single biggest quality jump.

Rerank the shortlist. Retrieve 30-50 candidates cheaply, then use a reranker to pick the best 5. Recall from the cheap stage, precision from the expensive one.

Filter on metadata first. Product, language, date, tenant. Cheap, deterministic, and it eliminates entire categories of wrong-but-similar results.

Retrieve generously. People agonise over top-k=3 to save tokens. With modern context windows, passing 10-20 good chunks is affordable and much more forgiving of imperfect ranking.


The Hybrid That Usually Wins

For most products I'd now build:

  1. Metadata filter to the relevant slice (tenant, product, language).
  2. If the slice fits in context, send all of it, cached. Done.
  3. If it doesn't, hybrid search + rerank within the slice.

Note what this makes possible: for a small tenant, the "retrieval system" is a WHERE clause. You only pay the complexity for the tenants that actually need it.


How to Decide in Practice

Start with the simplest thing that could work — stuff the corpus in, cache the prefix, ship it to real users. Measure quality with a proper eval set rather than by feel, which I wrote about separately in the post on evals.

Then add retrieval when the numbers force you to: the corpus outgrows the window, latency exceeds your budget, cost per request exceeds its value, or you need citations and access control.

The failure mode I keep seeing isn't teams choosing wrong. It's teams building a full RAG pipeline in week one — for a 40-page FAQ — and then spending months debugging retrieval quality for a problem they never actually had.

  • #rag
  • #vector database
  • #llm
  • #prompt caching
  • #architecture

Written by

Mihajlo Petrović

Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.

Have a task that repeats every week?

Tell me about it. If it can be automated well, I will show you how. If it cannot, I will say that too.

Tell me what to automate