Evaluating embeddings and LLMs on classical texts: what the benchmarks showed
While building a semantic search and chatbot over classical Chinese literature, I ran retrieval benchmarks across embedding models and a three-mode ablation across eight LLMs. Here is what the success rates, failure modes, and tables revealed.