We took the same ten thousand documents, the same fifty questions, and ran them through six different RAG setups. Same corpus, same eval — only the stack changed.
What actually moved the numbers
Chunking strategy and reranking mattered more than the vector database. The cheapest store with a good reranker beat the premium one without. Latency, on the other hand, was dominated by embedding calls, not retrieval.
6
stacks tested
10k
documents indexed
50
eval questions
The full results table — accuracy, latency, and cost per query for each setup — is in the community workspace. The short version: spend your effort on chunking and reranking before you argue about databases.
Written by Tomas V.
Systems engineer at AI Builders Stack. Cares a little too much about retrieval quality, latency budgets, and failing loudly.