01
5,115,734bibliographic recordsCase file / 001
Book IDSearch
A large-scale bibliographic search system for tracing books across imperfect metadata, obscure identifiers, and personal reading history.
02
06core searchable fields03
1,586personal reading snapshots04
0.0056%source parse failure rateA finding instrument, not a digital library.
Book ID Search began with a practical problem: millions of Chinese bibliographic records existed locally, but they were difficult to query, inconsistent in structure, and often discoverable only through identifiers such as SSID or DXID.
The project turns that material into a fast, legible research index. It searches metadata—not book contents—and keeps the original record available as evidence. It does not call an external book API or fabricate descriptions when the source is incomplete.
Six working layers
01
Identifier search
Search by title, author, publisher, ISBN, SSID, or DXID—even when the available record is partial.
02
Evidence-first results
Every result stays grounded in stored bibliographic fields and the original source record instead of invented summaries.
03
AI-assisted discovery
Query cleaning, intent recognition, evidence-aware reranking, and book insights help turn an imprecise memory into a traceable result.
04
Personal overlays
Private reading history can be matched against the public index without exposing personal notes or weakening the shared search layer.
05
Notes retrieval
A dedicated personal library adds full-text search across reading notes while keeping public and private records visibly distinct.
06
Quality regression
Scheduled checks monitor service health, AI answer quality, search relevance, weak records, and duplicate behavior over time.
From raw record to checkable result
Ingest
Stream large, irregular TXT records without loading the full source into memory.
Normalize
Parse stable fields while preserving raw evidence for incomplete or weak records.
Index
Build a Chinese fuzzy-search index across identifiers and bibliographic metadata.
Interpret
Clean the query, infer intent, and produce multiple evidence-bearing candidates.
Rerank
Unify lexical, identifier, and AI signals without hiding uncertainty.
Return
Show a copyable, exportable result with fields that can be checked independently.
Designed around uncertainty
Preserve raw records and show exact identifiers so every match can be checked.
Stream imports, resume interrupted jobs, and keep the entire index searchable on modest infrastructure.
Keep personal reading history and notes as a private overlay rather than blending them into the public index.
Treat relevance, weak metadata, and AI quality as recurring checks—not one-time launch tasks.
Current state / Live
Search the index.
The full 5,115,734-record index is online. Search by title, author, publisher, ISBN, SSID, or DXID.
Visit Book ID Search