Part Search Benchmark Lab
Synthetic benchmark cases enrich FindMyPart's search process without contaminating customer, revenue, or solved-case truth. Results remain evidence leads until reviewed.
Pass 1 → Pass 2
Pass 1 intentionally used underspecified identity. Pass 2 adds decisive model/YMM/position/drive/spec identity before search. This is process evidence, not a claim that every case is solved.
The candidate rejection rate applies only to the candidate sets explicitly instrumented in Pass 2. It is not yet a portfolio-wide false-positive rate. Each rejection now carries a reason code and the identity field that killed it.
Run ledger
Frozen machine-readable summaries preserve the denominator and scope of every measured run, so later UI or benchmark growth cannot silently rewrite historical metrics.
Regression queue
Every challenger search strategy must beat a frozen baseline on fixed cases before production promotion. A prettier answer is not an improvement metric.
Measured performance telemetry
Only directly observed timings belong here. Token and dollar fields stay empty unless the underlying provider reports them. Modeled work units live in the candidate microscope and are not treated as measured cost.
Candidate evaluation microscope
Search actions can hide several compatibility decisions inside one result page. This ledger counts those decisions separately. Work-unit values are experimental accounting until calibrated against runtime, tokens and tool cost.
Adversarial intake tests
Near-miss identities are first-class tests. The goal is not only to recognize a known configuration, but to refuse memory when one decisive field changes.
Advisory memory
Remembered evidence is typed by how quickly it can rot. Identity-bound fitment can suppress a redundant branch only on an exact fingerprint; supersession and other changing catalog facts can force revalidation instead.