Product Research Agent for TikTok Shop and Xiaohongshu
Build a product research agent that checks TikTok Shop, Xiaohongshu and Douyin for one product, computes stats in Python and has GPT-6.1 Sol write the brief.

The first time I ran this brief for “portable blender”, TikTok Shop gave me 10 listings, every one of them with sales, and the top three held 96% of units sold. That looks like a market locked up by a few sellers. Retrying the same request returned 30 different listings with a page token, and across two full pages (60 listings) the top-three share was about 65%. The first answer was a short fallback list, not the market.
This tutorial builds a product research agent for TikTok Shop and Xiaohongshu, with Douyin as a third signal. You give it one product keyword. It returns a one-page evidence brief: TikTok Shop listings (price, sold count, rating), Xiaohongshu notes and shop listings, Douyin videos, one CSV per platform, and a short summary written by an LLM from numbers your code already computed. It’s for sellers and agent builders who want evidence before stocking a product. Most of the work turned out to be spotting bad samples like that first one.
Key takeaway
- One SandBase API key covers the LLM step and all five data endpoints: TikTok Shop search and product detail, Xiaohongshu note and product search, and Douyin video search. All five data endpoints were listed as Free on 2026-10-03.
- Python does every count and median.
openai/gpt-6.1-solonly writes the narrative, and a check flags any number in its text that isn’t in the stats. Six briefs used 284 to 290 input tokens and 269 to 485 output tokens each, about $0.003 to $0.005 per brief.- TikTok Shop search returned a short fallback list (5 or 10 items,
has_more: false) on 10 of 21 calls in my final batch. The code retries it and never mixes it into a full page. If every retry comes back short, the brief keeps the short list and the LLM is told to call any sample under 30 thin.- Xiaohongshu medians moved a lot between runs minutes apart (160.5 to 354 for one keyword). Treat one run as a sample, not a measurement.
What the brief contains
This tutorial queries the TikTok Shop US storefront, so that side uses an English keyword. Xiaohongshu and Douyin are searched with the Chinese term. I tested two pairs: “portable blender” / 便携榨汁杯 (portable juicer cup) and “sunscreen” / 防晒霜 (sunscreen). Compare shapes (concentration, price band, engagement level), not raw totals across platforms.
| Signal | Endpoint | What the code computes |
|---|---|---|
| Listings, price, sold count, rating | tiktok/shop-web/search-products-list | listings with sales, median price, total sold, top-3 share of units sold, median rating of reviewed listings |
| Star mix of the best seller | tiktok/shop-web/product-detail-v2 | share of 1 and 2 star reviews |
| Notes about the product | xiaohongshu/app-v2/search-notes | note count, median likes + saves + comments, median saves, notes from the last 90 days, video notes |
| Xiaohongshu shop listings | xiaohongshu/app-v2/search-products | goods cards, median listed price |
| Short videos | douyin/search/video-search-v2 | video count, median likes, median likes + saves + comments + shares, videos from the last 90 days |
| Narrative | openai/gpt-6.1-sol via /v1/chat/completions | 120 to 180 words, using only the computed numbers |
One run made 11 to 13 HTTP calls and took about 55 seconds of wall time in my final runs.
Data boundary
This uses public, read-only data. It needs a SandBase API key. SandBase isn’t an official partner of TikTok, Xiaohongshu or Douyin. There are no private accounts, DMs, seller back-office analytics, orders or account actions. The CSVs keep note and video ids and counts, not authors or post text.
| Need | Use |
|---|---|
| Public listings, notes and videos for pre-launch research | SandBase public-data endpoints (this tutorial) |
| Your own shop’s orders, inventory and ads | TikTok Shop Partner Center |
| Your own Xiaohongshu account or store | Xiaohongshu open platform |
| Your own Douyin account data | Douyin open platform |
Which endpoints held up
All candidates are enabled in the SandBase registry and completed live on 2026-10-03 (UTC):
| Endpoint | What I saw | Used |
|---|---|---|
tiktok/shop-web/search-products-list | 30 listings per full page with has_more and a page_token. Often a short fallback list instead | Yes, with retries |
tiktok/shop-web/product-detail (v1) | Completed, but global_data.product_info held only empty seller_id fields. No rating data | No |
tiktok/shop-web/product-detail-v2 | Star histogram and category path. Intermittent HTTP 503 with upstream error 400 | Yes, optional |
xiaohongshu/app-v2/search-notes | 20 notes per page with counts and a timestamp | Yes |
xiaohongshu/app-v2/search-products | 20 cards per call, 18 or 19 of them goods cards with price_info | Yes |
douyin/search/video-search-v2 | 7 or 8 cards per page, statistics per video | Yes |
About the product-detail failures: earlier tests had reported a “region mismatch” on this route. I never saw that message. What I saw was HTTP 503 with upstream error 400 on the same US product id that worked in other runs. In about 20 practice runs earlier that day, the step was skipped after three failed attempts 4 times. In the final 6 runs, all 6 succeeded. So the code treats it as optional and the brief simply leaves out the star fields when it fails. A different region problem is easy to confuse with this. Sending region: "GB" for the same product (run c0788442-4d80-49e6-92fc-a1f69ac2d21c) came back completed with region_supported: false and a supported_regions list of ID, JP, MX, MY, PH, SG, TH, US and VN, even though the reference lists GB. The code stays on US. If you change regions, check region_supported first.
The search reference documents search_word as required, plus offset, page_token and region (default US):
Caption: The search-products-list reference shows the POST route and the four body fields the agent sends; its example leaves outputs[0].data empty, so the payload fields below come from live calls (captured 2026-10-03).
Tested on 2026-10-03 (UTC)
Inputs: the generic keywords above, region US. Each request is POST https://api.sandbase.ai/v1/api/<vendor>/<path> with Authorization: Bearer $SANDBASE_API_KEY and a JSON body of that endpoint’s fields only. The program reads only outputs[0].data. Every field name below comes from these calls. They’re observed-only, not documented guarantees, which is why the code reads each one with .get().
curl -s https://api.sandbase.ai/v1/api/tiktok/shop-web/search-products-list \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"search_word": "portable blender", "region": "US", "offset": 0, "page_token": ""}'
The first call in one run (788a0b0d-6927-4c25-9bac-0046471254fb) came back with the short list. The retry (02a0296a-504e-4aeb-916d-93ba80a96a49) was a full page. Trimmed, with one brand-owned card:
{"code": 0, "message": "success", "data": {
"has_more": true,
"load_more_params": {"api_source": 2, "offset": 30, "page_token": "20261003043528F72A01878A90B7109715"},
"products": [{
"product_id": "1732254901028033219",
"brand_info": {"brand_name": "ZHENMI"},
"sold_info": {"sold_count": 24647},
"rate_info": {"review_count": "2394", "score": 4.1},
"product_price_info": {"currency_name": "USD", "sale_price_decimal": "32.99", "origin_price_decimal": "99.99"}
}]}}
The short version had products with 10 items, has_more: false, and load_more_params with offset: 0 and an empty page_token. It was the same 10 product ids every time it appeared for this keyword (the sunscreen one was always the same 5), whatever offset or token I sent. That’s the signature the code checks.
product-detail-v2 with {"product_id": "1731581114818794403", "region": "US"} (run f5d014ea-b8bb-494a-8c64-c505b58eb6d9) returned a components_map list. The product_info component carried:
{"reviews_info": {"review_ratings": {"overall_score": 4.3, "review_count": "12280",
"rating_result": {"1": "1355", "2": "348", "3": "619", "4": "1086", "5": "8872"}}}}
The bread_crumbs component read TikTok Shop, Household Appliances, Kitchen Appliances, Juicers & Blenders. The full body was about 487 KB, so the code keeps only the histogram.
search-notes with {"keyword": "便携榨汁杯", "page": 1, "search_id": "", "search_session_id": ""} (run ace7c7d1-cf41-4638-8088-fca7eea678ba) returned search_id and search_session_id next to data, and 20 data.items[].note objects. One, with id, text and timestamp removed:
{"type": "video", "liked_count": 2885, "collected_count": 5744, "comments_count": 20,
"shared_count": 600, "timestamp": "<unix seconds, redacted>"}
Sending those two ids back with page: 2 gave 20 new notes in a probe. Without them, page 2 overlapped page 1. Even with them, two pages sometimes held only 30 to 35 unique notes, so the code deduplicates by id.
Caption: The search-notes reference documents page plus the search_id and search_session_id values from the first response, which is how the agent requests page 2 (captured 2026-10-03).
search-products with {"keyword": "便携榨汁杯"} (run 0922a59c-2179-48af-a249-4395c022e7e3) put cards under data.module.data. 19 of 20 had card_name: "cosmos_search_goods_card". A brand card:
{"title": "小米·米家随行便携榨汁杯2 家用小型多功能水果奶昔电动搅拌果汁机",
"price_info": {"price": 89.9, "origin_price": 89.9, "foreign_price": "≈$13.57", "symbol": "$"}}
The price values look like CNY. foreign_price is a converted figure whose currency changed between runs (GBP, then USD), so the code ignores it. Sold counts only appeared as text tags like “12k+ sold”, so I left them out.
video-search-v2 with {"keyword": "便携榨汁杯", "cursor": 0} (run 65d1a688-2f06-4401-8744-619eb8894eda) returned the payload directly in outputs[0].data, without the code/data wrapper the other four use. business_config.has_more was 1 and business_config.next_page held cursor: 8 and a search_id. The next page also needs business_config.backtrace. Seven of the eight cards were videos (one was a related-search card). Trimmed statistics for one:
{"digg_count": 233, "comment_count": 61, "share_count": 122, "collect_count": 54, "play_count": 0}
play_count was 0 on every card I looked at, so the brief doesn’t use views.
The complete program
One file, about 200 lines, standard library plus requests. Run it as python3 product_brief.py "portable blender" "便携榨汁杯".
#!/usr/bin/env python3
"""Cross-platform product research brief: TikTok Shop + Xiaohongshu + Douyin for one product.
Usage: python3 product_brief.py "portable blender" "便携榨汁杯"
Writes tiktok_shop.csv, xiaohongshu_notes.csv, douyin_videos.csv and brief.md.
"""
import csv
import json
import os
import re
import sys
import time
from statistics import median
import requests
BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
NOW = time.time()
def call(model: str, params: dict, retries: int = 2) -> dict:
"""POST /v1/api/<model>; return outputs[0].data or raise with status/error."""
for attempt in range(retries + 1):
try:
resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
body = resp.json()
except requests.RequestException as err: # dropped connection or truncated body
if attempt == retries:
raise RuntimeError(f"{model}: {err}") from err
time.sleep(3 * (attempt + 1))
continue
if resp.status_code >= 500 and attempt < retries:
time.sleep(3 * (attempt + 1)) # upstream hiccup: back off, then retry
continue
outputs = body.get("outputs") or []
if resp.status_code != 200 or body.get("status") != "completed" or not outputs:
raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
return outputs[0].get("data") or {}
raise RuntimeError(f"{model}: retries exhausted")
def num(x) -> float:
"""Counts arrive as ints or numeric strings; anything else counts as 0."""
try:
return float(x)
except (TypeError, ValueError):
return 0.0
def med(values):
values = [v for v in values if v is not None]
return round(median(values), 2) if values else None
def tiktok_shop(keyword: str, pages: int = 2, region: str = "US") -> list[dict]:
rows, offset, token = [], 0, ""
for _ in range(pages):
for _attempt in range(4): # a short fallback set (has_more false) comes back often; retry it
data = call("tiktok/shop-web/search-products-list",
{"search_word": keyword, "region": region, "offset": offset, "page_token": token})
page = data.get("data") or {}
if page.get("has_more"):
break
time.sleep(2)
if rows and not page.get("has_more"):
break # don't mix the fallback set into a full first page
for p in page.get("products") or []:
price = p.get("product_price_info") or {}
rate = p.get("rate_info") or {}
rows.append({"product_id": p.get("product_id"), "title": (p.get("title") or "")[:80],
"brand": (p.get("brand_info") or {}).get("brand_name") or "",
"price": num(price.get("sale_price_decimal")), "currency": price.get("currency_name"),
"sold": int(num((p.get("sold_info") or {}).get("sold_count"))),
"rating": num(rate.get("score")), "reviews": int(num(rate.get("review_count")))})
more = page.get("load_more_params") or {}
if not page.get("has_more") or not more.get("page_token"):
break
offset, token = more.get("offset", offset + 30), more["page_token"]
return rows
def tiktok_star_mix(product_id: str, region: str = "US") -> dict:
"""Star histogram of one listing from product-detail-v2; returns {} if the call fails."""
try:
data = call("tiktok/shop-web/product-detail-v2", {"product_id": product_id, "region": region})
except RuntimeError as err:
print(f" product-detail-v2 skipped: {err}")
return {}
comps = {c.get("component_name"): c.get("component_data") or {}
for c in (data.get("data") or {}).get("components_map") or [] if isinstance(c, dict)}
hist = ((comps.get("product_info") or {}).get("reviews_info") or {}).get("review_ratings", {}).get("rating_result") or {}
total = sum(num(v) for v in hist.values())
low = num(hist.get("1")) + num(hist.get("2"))
return {"top_listing_star_reviews": int(total),
"top_listing_low_star_share": round(low / total, 3) if total else None}
def xhs_notes(keyword: str, pages: int = 2) -> list[dict]:
rows, ids = [], {"search_id": "", "search_session_id": ""}
for page in range(1, pages + 1):
data = call("xiaohongshu/app-v2/search-notes", {"keyword": keyword, "page": page, **ids})
ids = {k: data.get(k) or "" for k in ids} # paging tokens sit next to `data`
for item in (data.get("data") or {}).get("items") or []:
n = item.get("note") or {}
if n.get("id"):
rows.append({"note_id": n["id"], "type": n.get("type"), "likes": int(num(n.get("liked_count"))),
"saves": int(num(n.get("collected_count"))),
"comments": int(num(n.get("comments_count"))),
"age_days": round((NOW - num(n.get("timestamp"))) / 86400)})
return list({r["note_id"]: r for r in rows}.values()) # drop duplicates across pages
def xhs_goods_prices(keyword: str) -> list[float]:
data = call("xiaohongshu/app-v2/search-products", {"keyword": keyword})
cards = ((data.get("data") or {}).get("module") or {}).get("data") or []
return [num((c.get("content") or {}).get("price_info", {}).get("price"))
for c in cards if c.get("card_name") == "cosmos_search_goods_card"]
def douyin_videos(keyword: str, pages: int = 3) -> list[dict]:
rows, params = [], {"keyword": keyword, "cursor": 0}
for _ in range(pages):
data = call("douyin/search/video-search-v2", params)
for card in data.get("business_data") or []:
a = (card.get("data") or {}).get("aweme_info") or {}
s = a.get("statistics") or {}
if a.get("aweme_id"):
rows.append({"aweme_id": a["aweme_id"], "likes": int(num(s.get("digg_count"))),
"comments": int(num(s.get("comment_count"))), "shares": int(num(s.get("share_count"))),
"saves": int(num(s.get("collect_count"))),
"age_days": round((NOW - num(a.get("create_time"))) / 86400)})
cfg = data.get("business_config") or {}
nxt = cfg.get("next_page") or {}
if not cfg.get("has_more") or not nxt:
break
params = {"keyword": keyword, "cursor": nxt.get("cursor"), "search_id": nxt.get("search_id", ""),
"backtrace": cfg.get("backtrace", "")}
return list({r["aweme_id"]: r for r in rows}.values())
def summarize(tt, star, notes, prices, videos) -> dict:
sold = sorted((r["sold"] for r in tt), reverse=True)
rated = [r["rating"] for r in tt if r["reviews"] > 0]
eng = lambda r: r["likes"] + r["saves"] + r["comments"] # noqa: E731
return {
"tiktok_shop": {"listings": len(tt), "with_sales": sum(1 for s in sold if s > 0),
"median_price_usd": med([r["price"] for r in tt]), "total_sold": sum(sold),
"top3_sold_share": round(sum(sold[:3]) / sum(sold), 3) if sum(sold) else None,
"median_rating_reviewed": med(rated), **star},
"xiaohongshu": {"notes": len(notes), "median_engagement": med([eng(n) for n in notes]),
"median_saves": med([n["saves"] for n in notes]),
"posted_last_90d": sum(1 for n in notes if n["age_days"] <= 90),
"video_notes": sum(1 for n in notes if n["type"] == "video"),
"goods_cards": len(prices), "median_goods_price_cny": med(prices)},
"douyin": {"videos": len(videos), "median_likes": med([v["likes"] for v in videos]),
"median_engagement": med([eng(v) + v["shares"] for v in videos]),
"posted_last_90d": sum(1 for v in videos if v["age_days"] <= 90)},
}
def narrate(stats: dict, en_kw: str, zh_kw: str) -> tuple[str, dict]:
prompt = (f"Product: '{en_kw}' (TikTok Shop US) / '{zh_kw}' (Xiaohongshu, Douyin). "
"Write a 120-180 word stocking brief: demand signal, competition, price band, risks, and one "
"next check. Sample sizes are listings, notes, goods_cards and videos; call one under 30 thin. "
"Use only numbers that appear in this JSON; do not compute new ones.\n"
+ json.dumps(stats, ensure_ascii=False))
resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=180,
json={"model": "openai/gpt-6.1-sol", "max_completion_tokens": 800,
"messages": [{"role": "user", "content": prompt}]})
resp.raise_for_status()
body = resp.json()
return body["choices"][0]["message"]["content"] or "", body.get("usage") or {}
def unknown_numbers(text: str, stats: dict) -> list[str]:
"""Numbers in the narrative that are not in the stats JSON (percent forms of shares allowed)."""
known = set()
for v in re.findall(r"\d+(?:\.\d+)?", json.dumps(stats)):
known |= {v, v.rstrip("0").rstrip(".") if "." in v else v}
if float(v) < 1: # shares may be quoted as percentages
known |= {f"{float(v) * 100:.1f}".rstrip("0").rstrip("."), str(round(float(v) * 100))}
found = [n.replace(",", "") for n in re.findall(r"\d[\d,]*(?:\.\d+)?", text)]
return sorted({n for n in found if n not in known})
def write_csv(path: str, rows: list[dict]):
with open(path, "w", newline="", encoding="utf-8-sig") as f:
if rows:
w = csv.DictWriter(f, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)
if __name__ == "__main__":
en_kw, zh_kw = sys.argv[1], sys.argv[2]
tt = tiktok_shop(en_kw)
top = max(tt, key=lambda r: r["sold"]) if tt else None
star = tiktok_star_mix(top["product_id"]) if top else {}
notes, prices, videos = xhs_notes(zh_kw), xhs_goods_prices(zh_kw), douyin_videos(zh_kw)
write_csv("tiktok_shop.csv", tt)
write_csv("xiaohongshu_notes.csv", notes)
write_csv("douyin_videos.csv", videos)
stats = summarize(tt, star, notes, prices, videos)
text, usage = narrate(stats, en_kw, zh_kw)
flags = unknown_numbers(text, stats)
with open("brief.md", "w", encoding="utf-8") as f:
f.write(f"# Product brief: {en_kw} / {zh_kw}\n\n```json\n{json.dumps(stats, ensure_ascii=False, indent=2)}\n```\n\n")
f.write(text + "\n\n" + (f"Check these numbers: {', '.join(flags)}\n" if flags else "Numbers check: OK\n"))
print(json.dumps(stats, ensure_ascii=False, indent=2))
print(text)
print("unverified numbers:", flags or "none", "| usage:", usage)
A few design choices worth knowing before you change anything:
- Medians, not means. One sunscreen listing marked “COMING SOON” showed a price of 3000 USD and 0 sold. A mean price would have been pulled up by it. The median barely moves.
- The fallback check. A page with
has_morefalse gets up to four tries. If page 2 still comes back short after a full page 1, the loop stops rather than adding the 5 or 10 unrelated fallback listings. If page 1 never recovers, the brief is built on the short list, which is why the prompt asks the model to flag thin samples. The trade-off: a genuine short last page 2 also gets dropped. For a broad keyword I never saw one. - Engagement means different things. Xiaohongshu’s sum is likes + saves + comments. Douyin’s adds shares. Don’t compare the two medians as if they were the same metric.
- The numbers check.
unknown_numberscompares every number in the LLM text with the stats JSON, allowing shares written as percentages. It flagged nothing in my six final runs. It can’t tell whether a correct number sits in the wrong sentence.
Why GPT-6.1 Sol writes the narrative
In our Claude Sonnet 5.5 vs GPT-6.1 Sol agent test, GPT-6.1 Sol answered 30 of 30 short tool tasks correctly at a median of about $0.0066 per task. That test also showed where models slip: averaging and counting over raw tool output. So here the model never sees raw listings. It gets a JSON block of about a dozen numbers per platform and writes prose.
Caption: The GPT-6.1 Sol model page lists $2.00 per million input tokens and $10.00 per million output tokens, the prices used for the cost math below (captured 2026-10-03).
The six final briefs used 284 to 290 prompt tokens and 269 to 485 completion tokens. Between 0 and 213 of those completion tokens were reasoning tokens. At $2/M input and $10/M output (the model catalog also lists a higher tier above 272K prompt tokens, which this never gets near), that’s $0.0033 to $0.0054 per brief, or $0.027 for all six. The data calls cost nothing at the listed price. The Douyin model page below shows Free, and the other four data endpoints returned base_price: "0" from GET /v1/models/<model> the same day. That’s the listing on 2026-10-03, not a guarantee.
Caption: The douyin/search/video-search-v2 model page shows a Free base price, sync execution and 8 input fields, including the cursor, search_id and backtrace fields the agent pages with (captured 2026-10-03).
What three runs per keyword looked like
I ran the program three times per keyword pair between 04:33 and 04:38 UTC on 2026-10-03. Ranges across the three runs:
| Metric | portable blender | sunscreen |
|---|---|---|
| TikTok listings kept | 60, 60, 60 | 60, 60, 30 |
| TikTok median price (USD) | 16.09 to 16.86 | 20.91 to 23.20 |
| TikTok top-3 share of units sold | 0.636 to 0.654 | 0.421 to 0.615 |
| TikTok median rating (reviewed listings) | 4.45 to 4.5 | 4.7 |
| Best seller’s 1 and 2 star share | 0.139 | 0.04 to 0.138 |
| Xiaohongshu unique notes | 30 to 40 | 32 to 40 |
| Xiaohongshu median engagement | 160.5 to 354 | 52 to 79.5 |
| Xiaohongshu median goods price (CNY) | 57.99 to 59.9 | 54 to 103.25 |
| Douyin videos | 16 to 21 | 20 |
| Douyin median likes | 101 to 152 | 8,887 to 8,899 |
TikTok Shop was steady once the sample was full, so one good run gives usable price and concentration numbers. The sunscreen best seller changed in one run (a different listing ranked first), which moved the low-star share from 0.138 to 0.04, so name the listing next to that number. Xiaohongshu was the noisy one: median engagement for the portable blender keyword went 263.5, 354, then 160.5 across runs two minutes apart. Practice runs ranged wider still, 61 to 354. Douyin’s top results changed less. They were also old: only 3 to 5 of about 20 videos were from the last 90 days, because default search ranks by relevance.
Here’s an excerpt from one brief the program wrote (the run with the 02a0296a... search above, LLM run 3eb63d10-25c2-412a-90b2-a6de9dab4770, 284 input and 485 output tokens):
Competition: TikTok sales are concentrated: top-listing concentration, measured by top3_sold_share, is 0.654. Enter cautiously rather than assuming demand is evenly accessible. Price band: The data establish median price anchors, not a defensible range: USD 16.86 on TikTok and CNY 59.9 on Xiaohongshu. Xiaohongshu’s 19 goods_cards are a thin sample.
Its numbers check came back OK. The prose is cautious, maybe too cautious, but it invents nothing.
Limits and next steps
Scope of testing: two keyword pairs, region US for TikTok Shop, 6 final runs plus about 20 practice runs on one day. I didn’t test other TikTok regions, long-tail keywords, or volumes beyond a few dozen calls. Search ranking on all three platforms is personalized and changes over time, so a brief describes what search showed at that moment, not the market.
Ways to extend it:
- Run it two or three times and keep the ranges. That’s cheap, and it’s the only fix I found for Xiaohongshu noise.
- Read reviews when the low-star share is high. The TikTok Shop product research tutorial pages review text with
product-reviews-v2. - Go deeper on one platform. See the Xiaohongshu product research API walkthrough or Douyin creator video research if you want sorting, time filters or creator follow-ups.
To try it, get a SandBase API key and use the Chat Completions guide for the LLM step.
FAQ
Can one keyword cover TikTok Shop, Xiaohongshu and Douyin?
Not well. TikTok Shop US listings are in English, so the agent takes an English keyword there and a Chinese one for Xiaohongshu and Douyin. Compare concentration and price band rather than totals.
Why did TikTok Shop return only 5 or 10 products?
That was a short fallback list: has_more false, an empty page_token, and the same product ids whatever offset I sent. It showed up on 10 of 21 search calls in my final batch. The program tries each page up to four times. In 1 of 6 final runs, page 2 never recovered and the brief kept 30 listings.
Does Xiaohongshu search return a total note count?
Not in the responses I saw. The brief counts the notes it fetched (two pages, up to 40, deduplicated). That’s a sample size, not the size of the topic.
How much does one brief cost?
At the prices listed on 2026-10-03, the five data endpoints were Free. The GPT-6.1 Sol call used up to 290 input and 485 output tokens, which is at most about $0.0054. A daily brief on 20 products would cost roughly $0.11 in LLM tokens at that rate. Check the model page before a large run.
Why not let the LLM read the raw listings?
Arithmetic is where models slip, and one TikTok detail payload was about 487 KB. Python computes the numbers, the model writes around them, and the check catches any figure that isn’t in the stats.