menu
{ "item_title" : "Mastering RAG", "item_author" : [" Yoo Byung Hong "], "item_description" : "The One Retrieval Book That Actually Measured Every CombinationYour RAG chunks are perfect. Your embeddings are the latest SOTA. Yet users say the answers are close but wrong. The problem isn't chunking - it's retrieval. This is the deep-dive book on the ONE thing that decides whether the right chunk even reaches your LLM.Why an Entire Book on RetrievalBecause just use vector search is why your queries about exact codes, part numbers, and proper nouns keep failing. This volume, Book 2 of the Mastering RAG series, is 20 chapters, 6 morphological analyzers, 4 index algorithms, 5 vector DBs, and a real production benchmark dedicated to answering one question: which retrieval combination actually wins, and why?The Retrieval Stack, Layer by LayerMorphological analysis: Mecab-ko, Kiwi, Khaiii, Okt, spaCy, kuromoji - how each tokenizer changes recallBM25 tuning: k1, b, field weighting - what the parameters actually do to your scoresLearned Sparse: SPLADE, BGE-M3 lexical weights - when they beat BM25Vector search: HNSW vs IVF vs PQ vs SCANN - parameter recipes for eachVector DBs: Milvus, Qdrant, Weaviate, pgvector - which fits which scaleHybrid fusion: RRF vs Weighted Sum vs Convex Combination - why RRF wins in practiceReranking: BGE-Reranker, Cohere Rerank, Jina - the last mile from top-100 to top-10The Empirical Experiment 50,000 chunks from a real production projectGolden set with facts, procedures, comparisons, overviewsRecall@10 lifts measured per stage - chunking baseline → +BM25 → +Dense → +Hybrid → +RerankPer-tokenizer comparison: which Korean morpheme analyzer moves Recall by how muchHNSW parameter sweep (M, efConstruction, efSearch) with actual latency numbersVector DB comparison at 5M vectors: throughput, latency, ops burdenWhy Reranking is worth its 200ms - with the exact recall liftKorean and BeyondKorean-specific: full chapter on the four major morpheme analyzers and when each winsMultilingual: Japanese (MeCab/Kuromoji), Chinese (Jieba), English (spaCy) coverageWhy just use multilingual-E5 is the naive answer - and what to do insteadNori vs Kiwi in Elasticsearch: production tradeoffsWho This Book Is ForEngineers whose RAG plateaued around 65% Recall and don't know which knob to turnTech leads deciding between vector DBs at production scaleML engineers who need to justify Hybrid retrieval vs Dense-onlyAnyone who read Chunking Deep Dive (Vol 1) and wants the full retrieval storyWhat Makes This Book DifferentEvery retrieval blog post says hybrid is better. This book shows you how much better, at what cost, with which reranker, on real data. Python and Java implementations included, along with a production deployment checklist and appendix cheat sheets covering every major retrieval library and benchmark suite.The Mastering RAG SeriesVolume 1: Chunking Deep Dive - Volume 2: Retrieval Deep Dive (this book) - Volume 3: Embeddings and Production Stack. Each volume goes as deep on one topic as most books go across all of RAG.PrerequisitesBasic RAG knowledge (or read AI RAG for Web Developers first). Comfortable with Python; Java examples optional. No Vol 1 dependency for retrieval concepts, though Vol 1 gives useful context on chunking baselines.Stop guessing at retrievers. Read the benchmark, ship the hybrid.", "item_img_path" : "https://covers2.booksamillion.com/covers/bam/9/79/819/142/9798191423845_b.jpg", "price_data" : { "retail_price" : "29.99", "online_price" : "29.99", "our_price" : "29.99", "club_price" : "29.99", "savings_pct" : "0", "savings_amt" : "0.00", "club_savings_pct" : "0", "club_savings_amt" : "0.00", "discount_pct" : "10", "store_price" : "" } }
Mastering RAG|Yoo Byung Hong

Mastering RAG : Retrieval Deep Dive: From Korean Morphology to Hybrid Search - The Empirical Guide to Making Your RAG Actually Find the Right Chunk

local_shippingShip to Me
In Stock.
FREE Shipping for Club Members help

Overview

The One Retrieval Book That Actually Measured Every Combination

Your RAG chunks are perfect. Your embeddings are the latest SOTA. Yet users say the answers are "close but wrong." The problem isn't chunking - it's retrieval. This is the deep-dive book on the ONE thing that decides whether the right chunk even reaches your LLM.

Why an Entire Book on Retrieval

Because "just use vector search" is why your queries about exact codes, part numbers, and proper nouns keep failing. This volume, Book 2 of the Mastering RAG series, is 20 chapters, 6 morphological analyzers, 4 index algorithms, 5 vector DBs, and a real production benchmark dedicated to answering one question: which retrieval combination actually wins, and why?

The Retrieval Stack, Layer by Layer
  • Morphological analysis: Mecab-ko, Kiwi, Khaiii, Okt, spaCy, kuromoji - how each tokenizer changes recall
  • BM25 tuning: k1, b, field weighting - what the parameters actually do to your scores
  • Learned Sparse: SPLADE, BGE-M3 lexical weights - when they beat BM25
  • Vector search: HNSW vs IVF vs PQ vs SCANN - parameter recipes for each
  • Vector DBs: Milvus, Qdrant, Weaviate, pgvector - which fits which scale
  • Hybrid fusion: RRF vs Weighted Sum vs Convex Combination - why RRF wins in practice
  • Reranking: BGE-Reranker, Cohere Rerank, Jina - the last mile from top-100 to top-10
The Empirical Experiment
  • 50,000 chunks from a real production project
  • Golden set with facts, procedures, comparisons, overviews
  • Recall@10 lifts measured per stage - chunking baseline → +BM25 → +Dense → +Hybrid → +Rerank
  • Per-tokenizer comparison: which Korean morpheme analyzer moves Recall by how much
  • HNSW parameter sweep (M, efConstruction, efSearch) with actual latency numbers
  • Vector DB comparison at 5M vectors: throughput, latency, ops burden
  • Why Reranking is worth its 200ms - with the exact recall lift
Korean and Beyond
  • Korean-specific: full chapter on the four major morpheme analyzers and when each wins
  • Multilingual: Japanese (MeCab/Kuromoji), Chinese (Jieba), English (spaCy) coverage
  • Why "just use multilingual-E5" is the naive answer - and what to do instead
  • Nori vs Kiwi in Elasticsearch: production tradeoffs
Who This Book Is For
  • Engineers whose RAG plateaued around 65% Recall and don't know which knob to turn
  • Tech leads deciding between vector DBs at production scale
  • ML engineers who need to justify Hybrid retrieval vs Dense-only
  • Anyone who read "Chunking Deep Dive" (Vol 1) and wants the full retrieval story
What Makes This Book Different

Every retrieval blog post says "hybrid is better." This book shows you how much better, at what cost, with which reranker, on real data. Python and Java implementations included, along with a production deployment checklist and appendix cheat sheets covering every major retrieval library and benchmark suite.

The Mastering RAG Series

Volume 1: Chunking Deep Dive - Volume 2: Retrieval Deep Dive (this book) - Volume 3: Embeddings and Production Stack. Each volume goes as deep on one topic as most books go across all of RAG.

Prerequisites

Basic RAG knowledge (or read "AI RAG for Web Developers" first). Comfortable with Python; Java examples optional. No Vol 1 dependency for retrieval concepts, though Vol 1 gives useful context on chunking baselines.

Stop guessing at retrievers. Read the benchmark, ship the hybrid.

This item is Non-Returnable

Details

  • ISBN-13: 9798191423845
  • ISBN-10: 9798191423845
  • Publisher: Independently Published
  • Publish Date: August 2026
  • Dimensions: 9 x 6 x 0.38 inches
  • Shipping Weight: 0.55 pounds
  • Page Count: 180

Related Categories

You May Also Like...

    1

BAM Customer Reviews