Search provenance
A search returns a ranked list and, normally, no reason for it. If a query returns less than you expected, you cannot tell from the response whether the query was too specific, a word was analyzed into something you did not intend, or the matching rule required more terms than your query could satisfy.
Provenance answers that. Send explain: true and the response carries a provenance block
describing what we searched for, how many of those terms had to match, and what we filtered
on.
Request
Add explain: true to any Catalog API search.
curl -sX POST "https://enterprise.shoppable.com/catalog" \ -H "Authorization: Bearer $SHOPPABLE_TOKEN" \ -H "x-shoppable-secret: $SHOPPABLE_SECRET" \ -H "Content-Type: application/json" \ -d '{ "query": "a soft cream-colored oversized cashmere crewneck sweater with ribbed cuffs", "page": 1, "rows": 25, "explain": true }'Response
Your results are unchanged. A provenance object is added alongside them.
{ "products": [ "..." ], "totalCount": 193, "provenance": { "query": { "received": "a soft cream-colored oversized cashmere crewneck sweater with ribbed cuffs", "searchedFor": ["soft", "cream", "color", "overs", "cashmer", "crewneck", "sweater", "rib", "cuff"], "termsRequired": 4, "termsTotal": 9 }, "matching": { "parser": "ExtendedDismaxQParser", "rule": "Three terms or fewer require every term. Four to six require all but one. Seven or more require half. Products matching the terms adjacently rank higher." }, "filters": { "merchantScope": "open pool", "inStockOnly": false, "onSale": false, "brands": 0, "categories": 0 }, "timingMs": 462 }}Reading it
searchedFor
The terms we actually searched, after analysis. This is the most useful field on the object and it rarely matches your input word for word.
In the example above, eleven words became nine terms:
| Your word | Searched as | Why |
|---|---|---|
a, with | dropped | Common words carry no signal and are removed |
cashmere | cashmer | Stemmed, so it also matches “cashmeres” |
oversized | overs | Stemmed |
cream-colored | cream, color | Hyphens split into separate terms |
ribbed | rib | Stemmed |
If a term you consider essential is missing here, or was stemmed into something unexpected, that explains the result set far faster than changing the query and guessing.
termsRequired and termsTotal
How many of those terms a product had to contain to qualify. Four of nine, above.
The requirement scales with query length, so a short query stays precise while a long descriptive one still returns something:
| Terms in query | Required to match |
|---|---|
| 1–3 | all of them |
| 4–6 | all but one |
| 7 or more | half |
Products that contain the terms adjacently rank above products that merely contain them, so relaxing the requirement widens what is eligible without pushing loose matches to the top.
filters
What narrowed the search besides the query itself. merchantScope is the one worth
checking first: if it names fewer merchants than you expected, the query may be fine and
the catalog you are searching may not be.
Using it to tune a generated query
If you build queries programmatically, the loop is:
- Send the generated query with
explain: true. - Compare
searchedForagainst the words that actually identify the product. - Drop anything that appears in
searchedForbut does not narrow the product — colors and materials usually help, styling adjectives usually do not. - Watch
termsRequiredfall as the query shortens. Fewer, better terms beats more terms.
Notes
- Provenance never affects your results. It describes a search that already ran.
- If the explanation cannot be produced, the field is omitted and your results are returned as normal. It will not fail your request.
timingMsis the search engine’s own time and excludes network transit.