Skip to content

Search provenance

A search returns a ranked list and, normally, no reason for it. If a query returns less than you expected, you cannot tell from the response whether the query was too specific, a word was analyzed into something you did not intend, or the matching rule required more terms than your query could satisfy.

Provenance answers that. Send explain: true and the response carries a provenance block describing what we searched for, how many of those terms had to match, and what we filtered on.

Request

Add explain: true to any Catalog API search.

Terminal window
curl -sX POST "https://enterprise.shoppable.com/catalog" \
-H "Authorization: Bearer $SHOPPABLE_TOKEN" \
-H "x-shoppable-secret: $SHOPPABLE_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "a soft cream-colored oversized cashmere crewneck sweater with ribbed cuffs",
"page": 1,
"rows": 25,
"explain": true
}'

Response

Your results are unchanged. A provenance object is added alongside them.

{
"products": [ "..." ],
"totalCount": 193,
"provenance": {
"query": {
"received": "a soft cream-colored oversized cashmere crewneck sweater with ribbed cuffs",
"searchedFor": ["soft", "cream", "color", "overs", "cashmer", "crewneck", "sweater", "rib", "cuff"],
"termsRequired": 4,
"termsTotal": 9
},
"matching": {
"parser": "ExtendedDismaxQParser",
"rule": "Three terms or fewer require every term. Four to six require all but one. Seven or more require half. Products matching the terms adjacently rank higher."
},
"filters": {
"merchantScope": "open pool",
"inStockOnly": false,
"onSale": false,
"brands": 0,
"categories": 0
},
"timingMs": 462
}
}

Reading it

searchedFor

The terms we actually searched, after analysis. This is the most useful field on the object and it rarely matches your input word for word.

In the example above, eleven words became nine terms:

Your wordSearched asWhy
a, withdroppedCommon words carry no signal and are removed
cashmerecashmerStemmed, so it also matches “cashmeres”
oversizedoversStemmed
cream-coloredcream, colorHyphens split into separate terms
ribbedribStemmed

If a term you consider essential is missing here, or was stemmed into something unexpected, that explains the result set far faster than changing the query and guessing.

termsRequired and termsTotal

How many of those terms a product had to contain to qualify. Four of nine, above.

The requirement scales with query length, so a short query stays precise while a long descriptive one still returns something:

Terms in queryRequired to match
1–3all of them
4–6all but one
7 or morehalf

Products that contain the terms adjacently rank above products that merely contain them, so relaxing the requirement widens what is eligible without pushing loose matches to the top.

filters

What narrowed the search besides the query itself. merchantScope is the one worth checking first: if it names fewer merchants than you expected, the query may be fine and the catalog you are searching may not be.

Using it to tune a generated query

If you build queries programmatically, the loop is:

  1. Send the generated query with explain: true.
  2. Compare searchedFor against the words that actually identify the product.
  3. Drop anything that appears in searchedFor but does not narrow the product — colors and materials usually help, styling adjectives usually do not.
  4. Watch termsRequired fall as the query shortens. Fewer, better terms beats more terms.

Notes

  • Provenance never affects your results. It describes a search that already ran.
  • If the explanation cannot be produced, the field is omitted and your results are returned as normal. It will not fail your request.
  • timingMs is the search engine’s own time and excludes network transit.