Mastering RAG : Retrieval Deep Dive: From Korean Morphology to Hybrid Search - The Empirical Guide to Making Your RAG Actually Find the Right Chunk
Overview
The One Retrieval Book That Actually Measured Every Combination
Your RAG chunks are perfect. Your embeddings are the latest SOTA. Yet users say the answers are "close but wrong." The problem isn't chunking - it's retrieval. This is the deep-dive book on the ONE thing that decides whether the right chunk even reaches your LLM.
Why an Entire Book on RetrievalBecause "just use vector search" is why your queries about exact codes, part numbers, and proper nouns keep failing. This volume, Book 2 of the Mastering RAG series, is 20 chapters, 6 morphological analyzers, 4 index algorithms, 5 vector DBs, and a real production benchmark dedicated to answering one question: which retrieval combination actually wins, and why?
The Retrieval Stack, Layer by Layer- Morphological analysis: Mecab-ko, Kiwi, Khaiii, Okt, spaCy, kuromoji - how each tokenizer changes recall
- BM25 tuning: k1, b, field weighting - what the parameters actually do to your scores
- Learned Sparse: SPLADE, BGE-M3 lexical weights - when they beat BM25
- Vector search: HNSW vs IVF vs PQ vs SCANN - parameter recipes for each
- Vector DBs: Milvus, Qdrant, Weaviate, pgvector - which fits which scale
- Hybrid fusion: RRF vs Weighted Sum vs Convex Combination - why RRF wins in practice
- Reranking: BGE-Reranker, Cohere Rerank, Jina - the last mile from top-100 to top-10
- 50,000 chunks from a real production project
- Golden set with facts, procedures, comparisons, overviews
- Recall@10 lifts measured per stage - chunking baseline → +BM25 → +Dense → +Hybrid → +Rerank
- Per-tokenizer comparison: which Korean morpheme analyzer moves Recall by how much
- HNSW parameter sweep (M, efConstruction, efSearch) with actual latency numbers
- Vector DB comparison at 5M vectors: throughput, latency, ops burden
- Why Reranking is worth its 200ms - with the exact recall lift
- Korean-specific: full chapter on the four major morpheme analyzers and when each wins
- Multilingual: Japanese (MeCab/Kuromoji), Chinese (Jieba), English (spaCy) coverage
- Why "just use multilingual-E5" is the naive answer - and what to do instead
- Nori vs Kiwi in Elasticsearch: production tradeoffs
- Engineers whose RAG plateaued around 65% Recall and don't know which knob to turn
- Tech leads deciding between vector DBs at production scale
- ML engineers who need to justify Hybrid retrieval vs Dense-only
- Anyone who read "Chunking Deep Dive" (Vol 1) and wants the full retrieval story
Every retrieval blog post says "hybrid is better." This book shows you how much better, at what cost, with which reranker, on real data. Python and Java implementations included, along with a production deployment checklist and appendix cheat sheets covering every major retrieval library and benchmark suite.
The Mastering RAG SeriesVolume 1: Chunking Deep Dive - Volume 2: Retrieval Deep Dive (this book) - Volume 3: Embeddings and Production Stack. Each volume goes as deep on one topic as most books go across all of RAG.
PrerequisitesBasic RAG knowledge (or read "AI RAG for Web Developers" first). Comfortable with Python; Java examples optional. No Vol 1 dependency for retrieval concepts, though Vol 1 gives useful context on chunking baselines.
Stop guessing at retrievers. Read the benchmark, ship the hybrid.
This item is Non-Returnable
Customers Also Bought
Details
- ISBN-13: 9798191423845
- ISBN-10: 9798191423845
- Publisher: Independently Published
- Publish Date: August 2026
- Dimensions: 9 x 6 x 0.38 inches
- Shipping Weight: 0.55 pounds
- Page Count: 180
Related Categories
