Retail Search

ARCH-0.5 · cross-domain core

Live · Cranfield baseline

Mission-driven search experiments

Each phase gets its own focused workspace. The home page stays small: status, purpose, metrics, and links to the phase details.

Phase 1Live

Cranfield Foundation

Production-shaped OpenSearch BM25 baseline with public search, explain, and evaluation evidence.

Phase 2Live

Cross-Domain Validation (BEIR)

Test whether Phase 1 techniques transfer across domains. Result: BGE hybrid is universal (the ARCH-0.5 core); keyword rerankers are domain-conditional; learning-to-rank is portfolio.

Phase 3Live

Retail Relevance (Amazon ESCI)

Product search over 1.2M Amazon products with structured fields. The BGE hybrid reaches parity with the published baseline (0.8468 vs 0.8503) and understands intent keyword search cannot.

Phase 4Planned

Behavioral Ranking

Introduce behavior signals only after earlier relevance gates justify the complexity.

Dataset references

Cranfield collection Glasgow Cranfield test collection

Classic information retrieval collection with aeronautics documents, queries, and relevance judgments. Phase 1 indexes the 1,400-document collection.

BEIR benchmark BEIR project

Heterogeneous IR benchmark. Phase 2 baselined all 15 public datasets (BM25 avg nDCG@10 0.4261, on the published ~0.43) and classified techniques cross-domain: the BGE hybrid is universal (ARCH-0.5 core), keyword rerankers are domain-conditional, learning-to-rank is portfolio.

Amazon ESCI Shopping Queries Amazon Science esci-data

Product-search benchmark with Exact, Substitute, Complement, and Irrelevant relevance labels for query-product pairs.

Behavior signals Dataset not selected

Future phase. No public behavior dataset is selected yet; this phase will document the source, privacy boundary, and evaluation protocol before implementation.