Hybrid Retrieval Benchmark
Build lexical, dense, and hybrid baselines against adjudicated queries and hard negatives. Report relevance, latency, and slice evidence.
Required evidence
Submit runnable commands, immutable source and index versions, stage-specific measurements, representative failures, one rejected alternative, and a decision with an explicit stop condition. Screenshots and untraceable generated prose do not pass.
Oral defense
Trace one output backward to its source, explain the first failing stage in one counterexample, justify the baseline and metric, and demonstrate update or deletion behavior.
Source backbone
Use Designing Data-Intensive Applications, official engine documentation, primary retrieval research, and NIST AI RMF selectively.