Price Comparison Agent for Google, Amazon and TikTok Shop
Build a price comparison agent API workflow: Google product cards, Amazon and TikTok Shop listings, model-number matching in Python, GPT-6 Luna for edge cases.

I ran the same Stanley Quencher 40 oz comparison twice, four minutes apart. In the first run, 24 of the 27 Google product cards my code accepted as the right tumbler came from eBay, Mercari, Poshmark or Whatnot sellers, and the Google median was $50. In the second run Google returned 6 accepted cards, none from those marketplaces, with a median of $47.50. Amazon’s median stayed at about $64 both times. The hard part of price comparison is knowing which listings are the product, and where they came from.
This tutorial builds a price comparison agent for three branded products across Google’s shopping product cards, Amazon and TikTok Shop (all US). For each product it fetches listings, decides in code which listings are the same item (title normalization, model numbers, size and capacity), sends only the undecided titles to a cheap LLM, flags variants and odd prices, and prints min, median and max USD price per source. It’s for developers and pricing analysts who want a price comparison agent API workflow they can rerun, rather than a one-off search.
Key takeaway
- Two of the five candidate endpoints didn’t work live. Both Google Shopping routes on SandBase (
serp/google/shopping/live/advancedandmerchant/google/products/live/advanced) completed but carried DataForSEO status 40402 “Invalid Path”. The agent reads Google’spopular_productscards from the live organic SERP instead.- On a hand check of one full run, 93 of the 99 listings the agent priced as matches were the right product (94% precision). All 8 LLM “match” calls were correct. The 6 misses were two “Lite” titles, a Bluetooth-only package, and three resale titles for a different Stanley line, lid or unit count.
- Prices moved less than the samples did. Logitech MX Master 3S medians were $88.99 on Amazon, $99 to $104.50 in Google cards and $109.74 on TikTok Shop across two runs.
- The LLM step cost $0.00042 to $0.00068 per full run on
openai/gpt-6-luna. The data endpoints were listed as Free on 2026-10-03, during an “API Free Week” banner.
What the agent compares
Three products, chosen because each has a clear model name and well-known look-alike variants:
| Product | Variants the code has to reject |
|---|---|
| Anker Nano Power Bank, 10,000 mAh, 30 W, built-in USB-C cable | 45 W Nano, 5,000 mAh Nano, Nano 3-in-1, MagGo, Zolo, 2-packs |
| Logitech MX Master 3S | MX Master 3, 4, 2S, “for Mac”, “for Business”, keyboard combos |
| Stanley Quencher H2.0 FlowState 40 oz | 14, 20, 30 and 64 oz cups, ProTour, Fluted, replacement lids and straws |
| Source | Endpoint | What the code reads |
|---|---|---|
| Google product cards | dataforseo/v3/serp/google/organic/live/advanced | popular_products items: title, seller, price.current |
| Amazon search | dataforseo/v3/merchant/amazon/products/live/advanced | amazon_serp and amazon_paid items: title, ASIN, price_from, URL slug |
| TikTok Shop US search | tiktok/shop-web/search-products-list | products: title, brand, sale_price_decimal |
| Undecided titles | openai/gpt-6-luna via /v1/chat/completions | match, variant or unclear per title |
A full run made 15 to 17 HTTP calls and took 148 to 171 seconds in my two final runs. Amazon was the slow one (10 to 41 seconds per call).
Data boundary
These are public listings at fetch time, read-only, for the United States (DataForSEO location_code 2840, TikTok region US). You need a SandBase API key. SandBase isn’t an official partner of Google, Amazon or TikTok, and nothing here touches seller accounts, orders, private data or account actions. Prices are what each listing showed when fetched and say nothing else about a retailer. Google card seller names are reported only in aggregate.
| Need | Use |
|---|---|
| Public listing prices across several stores for research | SandBase public-data endpoints (this tutorial) |
| Amazon product data for an Associates site | Amazon Product Advertising API |
| Your own Google Merchant Center products | Google Merchant API |
| Your own TikTok Shop catalog and orders | TikTok Shop Partner Center |
Which endpoints held up
All five candidates are enabled in the SandBase catalog. I called each live on 2026-10-03 (UTC):
| Endpoint | What I saw | Used |
|---|---|---|
serp/google/shopping/live/advanced | HTTP 200, completed, but tasks[0].status_code 40402 “Invalid Path”, no result | No |
merchant/google/products/live/advanced | Same 40402 “Invalid Path”, with location_code and with location_name | No |
serp/google/organic/live/advanced | 2 or 3 popular_products blocks, 16 to 40 product cards per query | Yes |
merchant/amazon/products/live/advanced | 100 to 123 listings per query | Yes |
tiktok/shop-web/search-products-list | 25 products when has_more is true; often a short list | Yes, with retries |
DataForSEO documents Google Shopping products as a task-based flow (post a task, collect it later). SandBase doesn’t list a task_post model for it: GET /v1/models/dataforseo/v3/merchant/google/products/task_post returned 404. So the agent uses the product cards on Google’s regular results page. Each carries a seller and a price, but there are fewer of them and they’re noisier than a Shopping tab.
Note the envelope too. The SandBase reference shows a completed response with exactly one outputs item holding data:
![SandBase endpoint reference for dataforseo/v3/merchant/amazon/products/live/advanced showing the POST route and a response schema where outputs[0] holds only data](https://static.sandbase.ai/blog/screenshots/price-comparison-agent-google-shopping-amazon-tiktok-shop/amazon-ref-00693ba0d287.webp)
Caption: The Amazon products reference documents POST /v1/api/dataforseo/v3/merchant/amazon/products/live/advanced and an outputs[0].data payload, while live calls returned the DataForSEO envelope directly in outputs[0] (captured 2026-10-03).
Live, the three DataForSEO endpoints returned the DataForSEO envelope (status_code, tasks, …) directly in outputs[0] with no data key, while TikTok returned outputs[0].data. The program’s reader prefers outputs[0].data, falls back to outputs[0] and then a top-level output, and raises on anything not completed. DataForSEO errors also need a second check: a 40402 arrives inside a completed SandBase run, so the code raises unless tasks[0].status_code is 20000.
Tested on 2026-10-03 (UTC)
Inputs: the three branded products above, no personal data. Each request is POST https://api.sandbase.ai/v1/api/<model> with Authorization: Bearer $SANDBASE_API_KEY and only that endpoint’s fields. Field names below are observed in these responses, not documented guarantees, which is why the code reads them with .get().
curl -s https://api.sandbase.ai/v1/api/dataforseo/v3/merchant/amazon/products/live/advanced \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"keyword": "Anker Nano Power Bank 10000mAh 30W", "language_code": "en_US", "location_code": 2840}'
Run 4bea800c-c5d8-49e1-b2d9-e599c02ded92 returned 113 items. Shown: the first of them, with the other item keys (URL, image, rating, delivery info and more) and the other envelope keys omitted.
{"id": "4bea800c-c5d8-49e1-b2d9-e599c02ded92", "status": "completed",
"outputs": [{"status_code": 20000, "status_message": "Ok.",
"tasks": [{"status_code": 20000, "result": [{"datetime": "2026-10-03 13:53:30 +00:00",
"items": [{"type": "amazon_serp", "rank_absolute": 1,
"title": "Nano Power Bank,10,000mAh 30W Portable Charger,Built-in USB-C Cable",
"data_asin": "B0C9CJKCH3", "price_from": 49.99, "currency": "USD"}]}]}]}]}
Amazon titles in this feed often drop the brand, so the code also checks the brand in the URL slug (Anker-Portable-Charger-... here).
The Google request has the same body shape with "language_code": "en" and "depth": 20. Run 49caa77e-9417-425a-aa04-4235cb615d6b carried 16 cards. One card from the brand’s own Newegg storefront, with image_url, rating, product_identifiers and other keys omitted:
{"type": "popular_products_element",
"title": "Anker Nano Power Bank, 10,000mAh Portable Charger with Built-in USB-C Cable, 30W Recharging, 30W Max Output with USB-C&A, for iPhone 17/16/15 Series,",
"seller": "Newegg.com - Anker Official Store",
"price": {"currency": "USD", "current": 54.99, "displayed_price": "$54.99",
"is_price_range": false, "max_value": null, "regular": null}}
The same body sent to serp/google/shopping/live/advanced (run e8c93b6a-8421-4508-8c32-e987b66d71d9) came back completed with this task, other keys omitted:
{"status_code": 40402, "status_message": "Invalid Path.", "cost": 0, "result": null}
TikTok Shop with {"search_word": "Anker Nano Power Bank 30W", "region": "US", "offset": 0, "page_token": ""}: the first call (run 67542af5-da6e-405c-8402-b953bc9f14f0) returned a short list of 15 with has_more: false, and the next one (0c087924-c820-47c5-9925-81edf24cf946) a full page of 25 with has_more: true. One product from the full page, other keys omitted:
{"product_id": "1732594588728267232",
"title": "Anker Nano Power Bank (30W, Built-In USB-C Cable) Phone Chargeable Smartphone Charging",
"product_price_info": {"currency_name": "USD", "sale_price_decimal": "49.99"},
"sold_info": {"sold_count": 0}}

Caption: The TikTok Shop search reference lists search_word as required plus offset, page_token and region, the four fields the agent sends (captured 2026-10-03).
The complete program
One file, standard library plus requests. Run it as python3 price_compare.py with SANDBASE_API_KEY set. It writes listings.csv with every listing and its verdict, and prints the price table.
#!/usr/bin/env python3
"""Cross-platform price comparison agent: Google product cards, Amazon and TikTok Shop (US).
Usage: python3 price_compare.py
Writes listings.csv (every listing with its verdict) and prints a min/median/max price table per source.
"""
import csv
import json
import os
import re
import time
from statistics import median
import requests
BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
US = {"location_code": 2840} # DataForSEO location code for the United States
PRODUCTS = [
{"name": "Anker Nano Power Bank (10K, 30W, built-in USB-C cable)", "query": "Anker Nano Power Bank 10000mAh 30W",
"tiktok": "Anker Nano Power Bank 30W", "brand": "anker", "family": r"\bnano\b.*\b(power ?bank|charger)\b",
"attrs": {"mah": 10000, "w": 30}, "model_no": r"\ba1259", # Anker's model number for this power bank
"other_models": r"3-in-1|maggo|magnetic|zolo|prime|powercore|connector|instacord|retractable"},
{"name": "Logitech MX Master 3S", "query": "Logitech MX Master 3S", "tiktok": "Logitech MX Master 3S",
"brand": "logitech", "family": r"\bmx master\b", "attrs": {"mx": "3s"},
"other_models": r"for mac|for business|\bbiz\b|combo|keys"},
{"name": "Stanley Quencher H2.0 FlowState 40 oz", "query": "Stanley Quencher H2.0 40 oz",
"tiktok": "Stanley Quencher 40 oz", "brand": "stanley", "family": r"\bquencher\b", "attrs": {"oz": 40, "gen": 2},
"other_models": r"fluted|protour|pro tour|iceflow|travel tumbler"},
]
ACCESSORY = r"\b(for|fits?|fit for|compatible with)\s+(the\s+)?(stanley|logitech|anker)\b|\breplacement\b|" \
r"\bcase for\b|\bboot\b|\bprotector\b|\bstraw cover\b|\blid kit\b"
BUNDLE = r"\b\d+\s*(pack|pcs|pieces|count)\b|\bpack of \d+|\bbulk\b|\bbundle\b|\bpromo(tional)?\b|\bcustom\b|" \
r"\bimprinted\b|\bpersonalized\b|\bset of\b|&\s*anker|\bwith anker\b"
USED = r"\b(used|renewed|refurbished|pre-owned|open box|excellent condition)\b"
def call(model: str, params: dict) -> dict:
"""POST /v1/api/<model>. Prefer outputs[0].data, fall back to outputs[0], then output."""
resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
body = resp.json()
if resp.status_code != 200 or body.get("status") != "completed":
raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
out = (body.get("outputs") or [None])[0]
data = out.get("data", out) if isinstance(out, dict) else body.get("output")
if data is None:
raise RuntimeError(f"{model}: completed without a payload")
return data
def dataforseo_items(model: str, params: dict) -> list[dict]:
"""DataForSEO wraps results in tasks[0].result[0].items; status_code 20000 means OK."""
task = (call(model, params).get("tasks") or [{}])[0]
if task.get("status_code") != 20000:
raise RuntimeError(f"{model}: DataForSEO {task.get('status_code')} {task.get('status_message')}")
return ((task.get("result") or [{}])[0] or {}).get("items") or []
def google_shopping(query: str) -> list[dict]:
"""Product cards ('popular_products') from a live Google SERP; each card carries seller and price."""
items = dataforseo_items("dataforseo/v3/serp/google/organic/live/advanced",
{"keyword": query, "language_code": "en", "depth": 20, **US})
cards = [c for i in items if i.get("type") == "popular_products" for c in i.get("items") or []]
return [{"source": "google", "id": (c.get("product_identifiers") or {}).get("data_docid"),
"title": c.get("title") or "", "seller": c.get("seller") or "", "extra": "",
"price": (c.get("price") or {}).get("current"), "currency": (c.get("price") or {}).get("currency")}
for c in cards]
def amazon(query: str) -> list[dict]:
items = dataforseo_items("dataforseo/v3/merchant/amazon/products/live/advanced",
{"keyword": query, "language_code": "en_US", **US})
return [{"source": "amazon", "id": i.get("data_asin"), "title": i.get("title") or "", "seller": "",
"extra": (i.get("url") or "").split("/")[3] if (i.get("url") or "").count("/") > 3 else "",
"price": i.get("price_from"), "currency": i.get("currency")}
for i in items if i.get("type") in ("amazon_serp", "amazon_paid")]
def tiktok_shop(query: str, tries: int = 4) -> list[dict]:
"""First page of TikTok Shop US search. A short fallback list (has_more false) is retried."""
for _ in range(tries):
page = call("tiktok/shop-web/search-products-list",
{"search_word": query, "region": "US", "offset": 0, "page_token": ""}).get("data") or {}
if page.get("has_more"):
break
time.sleep(2)
return [{"source": "tiktok", "id": p.get("product_id"), "title": p.get("title") or "", "seller": "",
"extra": (p.get("brand_info") or {}).get("brand_name") or "",
"price": float((p.get("product_price_info") or {}).get("sale_price_decimal") or 0) or None,
"currency": (p.get("product_price_info") or {}).get("currency_name")}
for p in page.get("products") or []]
def normalize(title: str) -> str:
t = title.lower().replace("™", "").replace("®", "").replace("™", "").replace("h2.o", "h2.0")
t = re.sub(r"(\d),(\d{3})", r"\1\2", t) # 10,000 -> 10000
t = re.sub(r"\b(\d+)k\b(?!\s*dpi)", lambda m: f"{int(m[1]) * 1000}mah", t) # 10K -> 10000mah
t = re.sub(r"\b1\.(18|2)\s*l\b", "40 oz", t) # 1.18 L / 1.2 L is the 40 oz cup
return re.sub(r"\s+", " ", t)
def attributes(t: str) -> dict:
return {"mah": {int(x) for x in re.findall(r"(\d{4,5})\s*mah", t)},
"w": {float(x) for x in re.findall(r"(\d+(?:\.\d+)?)\s*w\b", t)},
"oz": {int(x) for x in re.findall(r"(\d+)\s*(?:fl\.?\s*)?(?:oz|ounce)", t)},
"gen": {int(x) for x in re.findall(r"\bh(\d)\.0\b", t)},
"mx": set(re.findall(r"mx master\s*(\d+s?)\b", t))}
def classify(row: dict, spec: dict) -> tuple[str, str]:
"""Rule-based verdict: match, variant, accessory, other or ambiguous (sent to the LLM)."""
t = normalize(row["title"])
if re.search(ACCESSORY, t):
return "accessory", "accessory wording"
if spec["brand"] not in f"{t} {row['extra'].lower()} {row['seller'].lower()}":
return "other", "brand missing"
if not re.search(spec["family"], t):
return "other", "model family missing"
if m := re.search(BUNDLE, t):
return "variant", f"bundle or multi-pack ({m[0].strip()})"
if m := re.search(USED, t):
return "variant", f"condition ({m[0]})"
if m := re.search(spec["other_models"], t):
return "variant", f"different model ({m[0]})"
if spec.get("model_no") and re.search(spec["model_no"], t):
return "match", "model number"
found, missing = attributes(t), []
for key, want in spec["attrs"].items():
if not found[key]:
missing.append(key)
elif want not in found[key]:
return "variant", f"{key} {sorted(found[key])} vs {want}"
return ("ambiguous", "missing " + ", ".join(missing)) if missing else ("match", "all attributes present")
def llm_judge(spec: dict, rows: list[dict]) -> tuple[list[str], dict]:
"""Ask a cheap model about listings the rules could not decide. Titles only, one call per product."""
if not rows:
return [], {}
lines = "\n".join(f"{i}. {r['title']} | seller: {r['seller'] or r['extra'] or 'n/a'}" for i, r in enumerate(rows))
prompt = (f"Reference product: {spec['name']}.\nFor each numbered listing, answer 'match' only if the title "
"clearly identifies this exact product (any color is fine), 'variant' if it is a different size, "
"capacity, model or a bundle, and 'unclear' if the title does not say enough. Return JSON "
'{"verdicts": ["match", ...]} in the same order.\n' + lines)
resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=120,
json={"model": "openai/gpt-6-luna", "max_tokens": 2000, # leaves room for reasoning
"response_format": {"type": "json_object"},
"messages": [{"role": "user", "content": prompt}]})
resp.raise_for_status()
body = resp.json()
try:
verdicts = json.loads(body["choices"][0]["message"]["content"] or "{}").get("verdicts") or []
except json.JSONDecodeError: # truncated or malformed answer: treat every listing as unclear
verdicts = []
verdicts = [v if v in ("match", "variant", "unclear") else "unclear" for v in verdicts]
return (verdicts + ["unclear"] * len(rows))[:len(rows)], {"id": body.get("id"), **(body.get("usage") or {})}
def compare(spec: dict) -> tuple[list[dict], dict]:
rows, seen = [], set()
for fetch, query in ((google_shopping, spec["query"]), (amazon, spec["query"]), (tiktok_shop, spec["tiktok"])):
for r in fetch(query):
if (r["source"], r["id"], r["title"], r["seller"]) not in seen: # Amazon repeats ASINs as ads
seen.add((r["source"], r["id"], r["title"], r["seller"]))
rows.append(r)
for r in rows:
r["product"] = spec["name"]
r["verdict"], r["reason"] = classify(r, spec)
unsure = [r for r in rows if r["verdict"] == "ambiguous"]
verdicts, usage = llm_judge(spec, unsure)
for r, v in zip(unsure, verdicts):
r["verdict"], r["reason"] = v, f"LLM: {v} ({r['reason']})"
matched = [r["price"] for r in rows if r["verdict"] == "match" and r["price"]]
mid = median(matched) if matched else 0
for r in rows: # a matched title at under half or over twice the cross-source median needs a human look
if r["verdict"] == "match" and r["price"] and not 0.5 * mid <= r["price"] <= 2 * mid:
r["verdict"], r["reason"] = "price_check", f"price {r['price']} vs cross-source median {mid}"
return rows, usage
def price_table(rows: list[dict]) -> list[dict]:
out = []
for source in ("google", "amazon", "tiktok"):
got = [r for r in rows if r["source"] == source]
prices = [r["price"] for r in got if r["verdict"] == "match" and r["price"] and r["currency"] == "USD"]
out.append({"source": source, "listings": len(got), "matched": len(prices),
"variants_flagged": sum(r["verdict"] == "variant" for r in got),
"price_checks": sum(r["verdict"] == "price_check" for r in got),
"min_usd": min(prices) if prices else None,
"median_usd": round(median(prices), 2) if prices else None,
"max_usd": max(prices) if prices else None})
return out
if __name__ == "__main__":
started = time.strftime("%Y-%m-%d %H:%M UTC", time.gmtime())
all_rows, report = [], {"fetched_at": started, "location": "United States", "products": []}
for spec in PRODUCTS:
rows, usage = compare(spec)
all_rows += rows
report["products"].append({"product": spec["name"], "table": price_table(rows), "llm_usage": usage})
with open("listings.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["product", "source", "id", "title", "seller", "extra", "price",
"currency", "verdict", "reason"])
w.writeheader()
w.writerows(all_rows)
print(json.dumps(report, indent=1))
Each product’s printed table has one row per source, like this one from my first final run (fetched_at 2026-10-03 13:42 UTC):
{"source": "amazon", "listings": 104, "matched": 6, "variants_flagged": 14, "price_checks": 1,
"min_usd": 79.99, "median_usd": 88.99, "max_usd": 118.99}
How the matching works
The rules run in a fixed order, and the order matters:
- Accessory wording first. “Case for Anker…”, “Lid replacement for Stanley…” and silicone boots mention the product by name, so they’d pass every later check. On Amazon’s Stanley search, 39 of 106 listings in one run stopped here.
- Brand and model family. The brand can sit in the title, the seller, the TikTok brand field or the Amazon URL slug.
- Bundles, condition, sibling models. “Pack of 2”, promotional bulk lots, “renewed”, “for Business”, “ProTour” and the like become
variantwith a reason. These are flagged, not silently dropped, solistings.csvshows what was excluded. - Model number or attributes. Anker’s A1259 (the number on Anker’s own product URL in the Google results) counts as a match on its own. Otherwise every attribute in the spec must be present and equal: capacity and wattage, the version after “MX Master”, cup size and the H2.0 generation. A conflicting value is a variant. A missing value makes the listing
ambiguous. - LLM for the ambiguous rest. Only those titles go to GPT-6 Luna, one call per product, as match, variant or unclear. Unclear listings are left out of prices.
- Price sanity. A match priced under half or over twice the cross-source median becomes
price_check. That caught a $999 listing on a non-US site, resale listings at $13 to $28.62 and at $200 and $250, and a $249.98 keyboard-and-mouse bundle whose title didn’t say “combo”. It also flagged a $139.99 Amazon colorway that may be a real price, which is why it’s a flag and not a delete.
The normalizer folds the spellings that broke naive matching in my dev runs (“10,000mAh”, “10K”, “H2.O”, “1.18 L”), and its DPI guard stops “8K DPI” on a mouse from becoming a battery capacity.
Results from two runs
I ran the program twice, at 13:42 and 13:46 UTC on 2026-10-03. “Matched” counts matches with a USD price. Ranges show run 1 then run 2 where they differed.
| Product | Source | Listings | Matched | Min USD | Median USD | Max USD |
|---|---|---|---|---|---|---|
| Anker Nano 10K 30W | Google cards | 29, 35 | 8, 9 | 45.10, 49.99 | 54.49, 53.99 | 72.17, 61.99 |
| Amazon | 123, 111 | 5 | 49.99 | 54.99 | 54.99 | |
| TikTok Shop | 25 | 1 | 49.99 | 49.99 | 49.99 | |
| Logitech MX Master 3S | Google cards | 16 | 9, 10 | 89.99 | 99, 104.50 | 129.99 |
| Amazon | 104, 107 | 6 | 79.99 | 88.99 | 118.99 | |
| TikTok Shop | 25 | 6 | 99.99 | 109.74 | 137.79 | |
| Stanley Quencher 40 oz | Google cards | 40, 16 | 27, 6 | 30, 40 | 50, 47.50 | 79.99, 62 |
| Amazon | 106, 100 | 29, 30 | 39.99 | 64.24, 64.28 | 109.99 | |
| TikTok Shop | 25 | 8 | 52.85 | 56.75 | 68.71 |
What I’d take from this:
- Amazon and TikTok Shop were stable, Google wasn’t. The same query returned 40 Stanley cards and then 16, with a different seller mix. Run Google two or three times and keep the range.
- Most “matches” are color variants. All five Anker Amazon matches were color ASINs of the same power bank at $49.99 or $54.99. Most of the 29 to 30 Stanley Amazon matches were colorways between $39.99 and $109.99. If you need one SKU, add the color to the spec.
- Flags are most of the output. The first run flagged 87 variants and 11 price checks against 99 priced matches. The Anker Amazon search alone produced 33 variants: 45 W Nanos, 3-in-1s, 2-packs and MagGo bundles.
- TikTok Shop’s short list is common. 10 of 16 TikTok calls across both runs came back with
has_more: falseand 5, 15 or 20 products. The code retries untilhas_moreis true, up to four tries, and otherwise keeps the short list. In these runs every product got a 25-product page, on the fourth try at worst.

Caption: The SERP Google Organic Live model page shows a Free base price and sync execution under an “API Free Week” banner, the dated pricing behind the cost section (captured 2026-10-03).
How accurate the matching was
I hand-checked every listing that run 1 priced as a match, 99 in total, from the title and seller alone. My rule: correct if the title names the reference product (same model and size, a single unit, any color or edition). 93 were correct, a precision of 94%.
| Product | Source | Priced matches | Correct |
|---|---|---|---|
| Anker Nano 10K 30W | Google / Amazon / TikTok | 8 / 5 / 1 | 8 / 5 / 1 |
| Logitech MX Master 3S | Google / Amazon / TikTok | 9 / 6 / 6 | 7 / 5 / 6 |
| Stanley Quencher 40 oz | Google / Amazon / TikTok | 27 / 29 / 8 | 24 / 29 / 8 |
The six I counted wrong: two Google cards for an “MX Master 3S Lite” (not a product name I could confirm with Logitech), one Amazon “Bluetooth, no receiver” package, and three Stanley resale titles: one with a flip-top lid, one for the Adventure line, and one plural “Tumblers” that may be more than one unit. Adding “lite”, “no receiver” and “adventure” to the variant patterns would have caught four of them. Two of the six are judgment calls: the no-receiver package is still an MX Master 3S, and “Adventure Quencher” was also used for this cup. Count both as correct and precision is 95 of 99. I tuned the rules on earlier dev runs, so these numbers are measured on familiar products. Expect lower on new ones until you add their variant names.
The LLM decided 16 listings in that run. It called 8 a match, and all 8 were right by my labels. It called the other 8 unclear, such as a bare “Anker Nano Power Bank” with no capacity, and those stayed out of the prices. I also looked for misses among rejected listings that mention the model family. The brand check dropped at least four: two Amazon Stanley listings with no brand in the title or URL slug, and two Google cards spelling the mouse brand “Logicool” (its Japanese name) and “Logi tech”. Precision here is the stronger number. I didn’t measure recall.
Why GPT-6 Luna, and what it cost
In our cheap LLM tier benchmark, GPT-6 Luna scored 51/51 on short tool tasks at about 1/20 of GPT-6.1 Sol’s cost per task. Sorting a few titles into three labels doesn’t need more.
The four Luna calls in my final runs used 286 to 440 prompt tokens and 297 to 739 completion tokens. Of those, 256 to 705 were reasoning tokens, which is why max_tokens is 2000. My first attempt at 400 ran out mid-answer and the JSON didn’t parse. GET /v1/tasks/<id>/cost reported $0.000193 to $0.000398 per call, or $0.00068 for run 1 and $0.00042 for run 2, matching the $0.10 and $0.50 per million token list price.
The data calls cost nothing at the listed price that day. GET /v1/models/<id> returned base_price: "0" for all three endpoints, and the three data runs I looked up in /v1/tasks/<id>/cost settled at $0.000000. Both lookups need the same Bearer key; the public model pages show the same prices without one. The DataForSEO model and catalog pages I captured also showed an “API Free Week” banner, so check the price again before a scheduled job. The cost field inside DataForSEO’s own envelope (0.0033 for Amazon, 0.004 for Google) is the vendor’s internal accounting, not your SandBase bill.

Caption: The DataForSEO catalog listed 74 endpoints with an Available, Free status and the API Free Week banner, so the Free price is a dated observation (captured 2026-10-03).
Limits and next steps
Scope: three branded products, US only, two final runs plus dev runs on one day, first page of results per source. I didn’t test other countries, unbranded products, or Google cards for queries that show no product block. Search results are ranked and personalized upstream and change by the minute, so a run describes those listings at that moment, not the market.
Where to take it next:
- Add a seller filter for Google cards. If you only want new retail listings, drop or tag marketplace sellers before the median. Run 1 shows why.
- Pin the SKU. Add color or part number to the spec when you need one exact item, not the model family.
- Schedule it and keep history. Store
listings.csvper run and compare ranges, not single numbers. - Go deeper per source. The DataForSEO research agent tutorial covers more SERP endpoints, and the TikTok Shop product research walkthrough covers detail pages and reviews.
To run it yourself, open the Amazon products API reference and the Chat Completions guide, then get a SandBase API key.
FAQ
Why not use the Google Shopping endpoint directly?
On 2026-10-03 both Google Shopping routes on SandBase returned DataForSEO status 40402 “Invalid Path” inside a completed run. DataForSEO documents Google Shopping products as a task-based flow, which SandBase doesn’t list. The product cards on Google’s regular results page were the working alternative.
How does the agent know two listings are the same product?
Mostly in code: brand and model family, then accessory, bundle and sibling-model rejects, then model number, capacity, wattage or size against the spec. Only titles missing an attribute go to the LLM. In a hand check, 93 of 99 priced matches were correct.
Are these prices for new items?
Not always. Google cards included eBay, Mercari, Poshmark and Whatnot sellers, and in one run they made up 24 of 27 Stanley matches. The code flags titles that say “renewed” or “used” and odd prices, but it can’t see condition that a title doesn’t state.
How much does a run cost?
On 2026-10-03 the data endpoints were listed as Free and settled at $0. The LLM step cost $0.00042 to $0.00068 per run of three products. Prices can change after a promotion ends, so check the model pages first.
Can I add my own products?
Yes. Add a PRODUCTS entry with the query, brand, model-family pattern, attributes and sibling models to reject. Run it once, read listings.csv, and add the look-alikes you see to other_models before trusting the medians.