Blog/Tutorials/

Product Review Analysis Agent for TikTok Shop and Xiaohongshu

Build a product review analysis agent for TikTok Shop and Xiaohongshu, then check its labels against 40 hand-labeled reviews: GPT-6 Luna vs GPT-6.1 Sol.

Cover for a product review analysis agent tutorial covering TikTok Shop and Xiaohongshu reviews

The first time I pulled 300 Xiaohongshu reviews for three magnetic phone cases, every single one had 5 stars. The review overview for the biggest listing said otherwise: 74 of its 4,688 ratings were 1 to 3 stars. The default order was showing the best reviews first. Switching to the “Latest” sort fixed the star mix, and then 179 of the next 300 reviews turned out to be the platform’s automatic text for buyers who left only a star rating. Neither throws an error. Both quietly skew any theme analysis built on top.

This tutorial builds a product review analysis agent for TikTok Shop and Xiaohongshu. You give it one product keyword per platform. It finds the three most-reviewed public listings on each, pulls up to 100 written reviews per listing, has an LLM tag every review with one theme and a sentiment, then counts everything in Python. Output: a CSV with one row per review and a one-page brief. It’s for sellers and agent builders who want to know what buyers complain about. The measured part is the classifier: I hand-labeled 40 reviews and compared openai/gpt-6-luna with openai/gpt-6.1-sol on agreement, stability and cost.

Key takeaway

  • Against 40 hand-labeled reviews, GPT-6.1 Sol matched my theme on 35 of 40 in both runs (Cohen’s kappa 0.855). GPT-6 Luna matched 32 and 35 of 40. On the 29 reviews I found unambiguous, Sol got 28 both times; Luna got 25 and 28.
  • On all 578 reviews, Luna cost $0.012 per full run and Sol about $0.147, roughly 12 times more. Luna wrote about 1.8 times as many output tokens, so the gap is smaller than the 20 times list-price ratio.
  • Sol was steadier: two runs agreed on 96.9% of themes, versus 88.6% for Luna. Luna also skipped 1 or 2 of 578 reviews per run and, in one run, labeled two complaints positive.
  • The aggregate answer barely moved with the model. On TikTok Shop, durability, magnetic and fit were the top three complaint themes in all five runs. Xiaohongshu had only 4 to 8 complaints per theme, too few to rank.

What the agent does

The category is magnetic phone cases. On TikTok Shop US the keyword is “magsafe iphone case”. On Xiaohongshu it’s 磁吸手机壳 (“magnetic phone case”). I started with portable blenders, as in the product research agent tutorial, but the Xiaohongshu blender listings I checked had at most 26 reviews. Phone cases had hundreds to thousands per listing on both platforms.

StepEndpointWhat the code keeps
Find TikTok Shop listingstiktok/shop-web/search-products-listproduct id, title, review count; top 3 by reviews
TikTok Shop reviewstiktok/shop-web/product-reviews-v2review id, stars, text; star-only reviews skipped
Find Xiaohongshu listingsxiaohongshu/app-v2/search-productsfirst 8 goods cards
Rank them by review countxiaohongshu/app-v2/product-review-overviewtotal; top 3
Xiaohongshu reviewsxiaohongshu/app-v2/product-reviewsreview id, stars, text; “Latest” sort; placeholder text skipped
Theme + sentimentopenai/gpt-6-luna (default) or openai/gpt-6.1-sol via /v1/chat/completionsone of 9 themes, positive / negative / mixed

The theme list is fixed for phone cases: fit, protection, magnetic, look, feel, durability, value, shipping, other. The labeling rule matters as much as the list. If a review complains, its theme is the main complaint; otherwise it’s the aspect discussed most; generic praise is “other”. That rule makes the complaint ranking meaningful, and it’s the same rule I used when hand-labeling.

Data boundary

This reads public, read-only data. It needs a SandBase API key. SandBase isn’t an official partner of TikTok or Xiaohongshu. There are no private accounts, DMs, seller back-office analytics, orders or account actions. Review text and reviewer names are third-party personal content, so this article shows only counts, shares and generic paraphrases, and the CSV the program writes should stay on your machine.

NeedUse
Public reviews of any listing, for research or competitor analysisSandBase public-data endpoints (this tutorial)
Reviews and orders of your own TikTok Shop storeTikTok Shop Partner Center
Your own Xiaohongshu store dataXiaohongshu open platform

Two traps in the review endpoints

Xiaohongshu’s default order is curated. With the default sort_strategy_type of 0, three listings gave 300 reviews and all 300 had itemScore 5. The overview endpoint lists two sort options for these listings, Default (type 0) and Latest (type 1). With type 1, the first 36 reviews of the biggest listing had a 2-star, a 3-star and seven 4-star ratings. Types 2 and 3 returned no reviews. The agent sends sort_strategy_type: 1.

SandBase endpoint reference for xiaohongshu/app-v2/product-reviews showing POST /v1/api/xiaohongshu/app-v2/product-reviews with from_page, page, share_pics_only, sku_id and sort_strategy_type

Caption: The product-reviews reference documents sort_strategy_type with default 0 but doesn’t list its values, which is why the “Latest” value 1 below comes from the overview response and a live test (captured 2026-10-03).

Many Xiaohongshu “reviews” have no words from the buyer. In the “Latest” order, 179 of 300 reviews were system text meaning “this user gave a 5-star review”, “this user gave a positive review” or “this user thinks the product is average”. An LLM would happily file those under “other, positive” and dilute every share. The code drops anything matching ^该用户(给出|认为) (“this user gave / thinks”). It’s a pattern from what I observed, not a documented flag, so check it if your counts look odd. TikTok Shop has the simpler version of this: star-only reviews come back with an empty review_text, and the code skips them.

TikTok Shop’s default order (sort_rule 2) didn’t look curated: 73 of 300 reviews were 1 or 2 stars. The reference doesn’t document what the other sort_rule values mean, so I left the default.

SandBase endpoint reference for tiktok/shop-web/product-reviews-v2 showing filter_type, filter_value, page_start, product_id, region and sort_rule

Caption: The product-reviews-v2 reference lists page_start, product_id, region and sort_rule (default 2), the fields the agent uses for paging TikTok Shop reviews (captured 2026-10-03).

Tested on 2026-10-03 (UTC)

Inputs: the two generic keywords above, TikTok Shop region US, and public listing ids returned by search. Each data call is POST https://api.sandbase.ai/v1/api/<vendor>/<path> with Authorization: Bearer $SANDBASE_API_KEY and a JSON body of that endpoint’s fields only. The documented payload is outputs[0].data. Some endpoints have been reported to return a different shape, so the reader in the code prefers outputs[0].data, falls back to outputs[0] and then a top-level output, and raises on anything not completed. All five endpoints here returned outputs[0].data in every call I logged. Field names below are observed-only, not documented guarantees, which is why the code reads them with .get().

curl -s https://api.sandbase.ai/v1/api/xiaohongshu/app-v2/product-review-overview \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"sku_id": "68cc0e65d1f7b800015dacdc"}'

Run d3568c3c-2df6-4833-9e75-afc0118319ca returned this inside outputs[0].data.data (other keys omitted):

{"total": 4688, "avgScore": "4.5",
 "scoreToCnt": {"1": 23, "2": 12, "3": 39, "4": 352, "5": 4262},
 "sortStrategyList": [{"name": "Default", "type": 0}, {"name": "Latest", "type": 1}]}

The reviews call for the same listing, {"sku_id": "68cc0e65d1f7b800015dacdc", "page": 0, "sort_strategy_type": 1} (run d272e435-016b-4556-bd96-f074720f201e), returned 12 reviews and hasMore: true. The first review’s reviewInfo, with text, ids, images, time and user fields omitted:

{"itemScore": 5, "logisticsScore": 5, "serviceScore": 5}

TikTok Shop, {"product_id": "1730831546437308868", "region": "US", "page_start": 1} (run eaee7cd8-c911-48d3-80b2-9643d08c140d), returned 20 product_reviews and has_more: true. The first one, with text, reviewer, ids, SKU and image fields omitted:

{"product_id": "1730831546437308868", "review_rating": 5, "is_verified_purchase": true,
 "is_incentivized_review": false, "review_country": "US"}

The complete program

One file, 194 lines, standard library plus requests. Run it as python3 review_agent.py "magsafe iphone case" "磁吸手机壳", optionally with openai/gpt-6.1-sol as a third argument. This is the exact file I ran.

#!/usr/bin/env python3
"""Product review analysis agent: TikTok Shop + Xiaohongshu reviews -> themes, sentiment, CSV, brief.

Usage: python3 review_agent.py "magsafe iphone case" "磁吸手机壳" [openai/gpt-6-luna]
Writes reviews.csv (one row per review) and brief.md (aggregates per platform and theme).
"""
import csv
import json
import os
import re
import sys
import time
from collections import Counter

import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
THEMES = ["fit", "protection", "magnetic", "look", "feel", "durability", "value", "shipping", "other"]
SENTIMENTS = ["positive", "negative", "mixed"]
PRICE = {"openai/gpt-6-luna": (0.1, 0.5), "openai/gpt-6.1-sol": (2.0, 10.0)}  # USD per 1M tokens, 2026-10-03


def call(model: str, params: dict, retries: int = 2) -> dict:
    """POST /v1/api/<model>. Prefer outputs[0].data, fall back to outputs[0], then output."""
    for attempt in range(retries + 1):
        try:
            resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
            body = resp.json()
        except requests.RequestException as err:
            if attempt == retries:
                raise RuntimeError(f"{model}: {err}") from err
            time.sleep(3 * (attempt + 1))
            continue
        if resp.status_code >= 500 and attempt < retries:
            time.sleep(3 * (attempt + 1))
            continue
        if resp.status_code != 200 or body.get("status") != "completed":
            raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
        first = (body.get("outputs") or [None])[0]
        if isinstance(first, dict):
            return first["data"] if "data" in first else first
        if body.get("output") is not None:
            return body["output"]
        raise RuntimeError(f"{model}: completed without a payload")
    raise RuntimeError(f"{model}: retries exhausted")


def tiktok_listings(keyword: str, n: int = 3) -> list[dict]:
    for _attempt in range(4):  # a short fallback list (has_more false) comes back often; retry it
        page = call("tiktok/shop-web/search-products-list", {"search_word": keyword, "region": "US"}).get("data") or {}
        if page.get("has_more"):
            break
        time.sleep(2)
    rows = [{"platform": "tiktok", "listing": p["product_id"], "title": (p.get("title") or "")[:60],
             "reviews_total": int((p.get("rate_info") or {}).get("review_count") or 0)}
            for p in page.get("products") or [] if p.get("product_id")]
    return sorted(rows, key=lambda r: -r["reviews_total"])[:n]


def tiktok_reviews(product_id: str, limit: int = 100, max_pages: int = 10) -> list[dict]:
    out = []
    for page in range(1, max_pages + 1):
        data = call("tiktok/shop-web/product-reviews-v2",
                    {"product_id": product_id, "region": "US", "page_start": page}).get("data") or {}
        for r in data.get("product_reviews") or []:
            if (r.get("review_text") or "").strip():  # star-only reviews have nothing to classify
                out.append({"review_id": r.get("review_id"), "stars": r.get("review_rating"),
                            "text": r["review_text"].strip()})
        if len(out) >= limit or not data.get("has_more"):
            break
    return out[:limit]


def xhs_listings(keyword: str, n: int = 3, candidates: int = 8) -> list[dict]:
    data = call("xiaohongshu/app-v2/search-products", {"keyword": keyword}).get("data") or {}
    cards = [m.get("content") or {} for m in (data.get("module") or {}).get("data") or []
             if m.get("card_name") == "cosmos_search_goods_card"][:candidates]
    rows = []
    for c in cards:  # search cards carry no review count, so ask the overview endpoint
        ov = call("xiaohongshu/app-v2/product-review-overview", {"sku_id": c["id"]}).get("data") or {}
        rows.append({"platform": "xiaohongshu", "listing": c["id"], "title": (c.get("title") or "")[:30],
                     "reviews_total": int(ov.get("total") or 0)})
    return sorted(rows, key=lambda r: -r["reviews_total"])[:n]


XHS_PLACEHOLDER = re.compile(r"^该用户(给出|认为)")  # auto text for star-only reviews ("this user gave 5 stars")


def xhs_reviews(sku_id: str, limit: int = 100, max_pages: int = 25) -> list[dict]:
    out = []
    for page in range(max_pages):  # pages start at 0; sort 1 = "Latest" (default order was all 5-star)
        data = call("xiaohongshu/app-v2/product-reviews",
                    {"sku_id": sku_id, "page": page, "sort_strategy_type": 1}).get("data") or {}
        for r in data.get("reviews") or []:
            info = r.get("reviewInfo") or {}
            text = (info.get("content") or "").strip()
            if text and not XHS_PLACEHOLDER.match(text):
                out.append({"review_id": info.get("reviewId"), "stars": info.get("itemScore"),
                            "text": text})
        if len(out) >= limit or not data.get("hasMore"):
            break
    return out[:limit]


PROMPT = """You label product reviews of phone cases. For each numbered review return one primary theme and a sentiment.
Themes: fit (fits the phone model, cutouts, buttons), protection (drops, camera or screen coverage),
magnetic (MagSafe or magnet strength, wireless charging, built-in stand or ring), look (design, color, clarity,
matches photos), feel (grip, texture, thickness, material feel), durability (yellowing, cracking, peeling, wear),
value (price, worth the money), shipping (delivery, packaging, seller service, wrong item), other (generic praise
or anything else). If the review complains about something, the theme is the main complaint; otherwise it is the
aspect discussed most. Sentiment: positive, negative or mixed. Reviews may be English or Chinese.
Reply with JSON only: {"labels": [{"i": 0, "theme": "fit", "sentiment": "positive"}]}"""


def classify(reviews: list[dict], model: str, batch: int = 20) -> dict:
    """Adds theme/sentiment to each review in place; returns token usage and cost."""
    usage = Counter()
    for start in range(0, len(reviews), batch):
        chunk = reviews[start:start + batch]
        numbered = "\n".join(f"{i}. {r['text'][:400]}" for i, r in enumerate(chunk))
        resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=180,
                             json={"model": model, "max_tokens": 4000,
                                   "messages": [{"role": "system", "content": PROMPT},
                                                {"role": "user", "content": numbered}]})
        resp.raise_for_status()
        body = resp.json()
        for k in ("prompt_tokens", "completion_tokens"):
            usage[k] += (body.get("usage") or {}).get(k, 0)
        usage["calls"] += 1
        match = re.search(r"\{.*\}", body["choices"][0]["message"]["content"] or "", re.S)
        labels = {x.get("i"): x for x in (json.loads(match.group(0)).get("labels") or [])} if match else {}
        for i, r in enumerate(chunk):
            lab = labels.get(i) or {}
            r["theme"] = lab.get("theme") if lab.get("theme") in THEMES else "other"
            r["sentiment"] = lab.get("sentiment") if lab.get("sentiment") in SENTIMENTS else "mixed"
            r["unlabeled"] = not lab  # the model skipped this one; counted, never hidden
    pin, pout = PRICE[model]
    usage["usd"] = round((usage["prompt_tokens"] * pin + usage["completion_tokens"] * pout) / 1e6, 5)
    return dict(usage)


def aggregate(rows: list[dict]) -> dict:
    out = {}
    for platform in sorted({r["platform"] for r in rows}):
        rs = [r for r in rows if r["platform"] == platform]
        neg = [r for r in rs if r["sentiment"] in ("negative", "mixed")]
        themes = Counter(r["theme"] for r in rs)
        neg_themes = Counter(r["theme"] for r in neg)
        out[platform] = {
            "reviews": len(rs), "low_star_share": round(sum(1 for r in rs if (r["stars"] or 5) <= 2) / len(rs), 3),
            "negative_or_mixed_share": round(len(neg) / len(rs), 3),
            "unlabeled": sum(1 for r in rs if r["unlabeled"]),
            "theme_share": {t: round(themes[t] / len(rs), 3) for t in THEMES if themes[t]},
            "top_complaints": [(t, c) for t, c in neg_themes.most_common(3)],
        }
    return out


def write_brief(path: str, listings: list[dict], stats: dict, usage: dict, model: str):
    lines = [f"# Review brief ({time.strftime('%Y-%m-%d %H:%M UTC', time.gmtime())}, {model})", "",
             "| Platform | Listing | Reviews on listing | Sampled |", "|---|---|---|---|"]
    lines += [f"| {l['platform']} | {l['listing']} | {l['reviews_total']} | {l['sampled']} |" for l in listings]
    for platform, s in stats.items():
        lines += ["", f"## {platform}: {s['reviews']} reviews, {s['negative_or_mixed_share']:.0%} negative or mixed, "
                      f"{s['low_star_share']:.0%} at 1-2 stars", "", "| Theme | Share of reviews |", "|---|---|"]
        lines += [f"| {t} | {v:.0%} |" for t, v in sorted(s["theme_share"].items(), key=lambda kv: -kv[1])]
        lines += ["", "Top complaint themes: " + ", ".join(f"{t} ({c})" for t, c in s["top_complaints"])]
    lines += ["", f"LLM usage: {usage['calls']} calls, {usage['prompt_tokens']} in / "
                  f"{usage['completion_tokens']} out tokens, about ${usage['usd']}"]
    with open(path, "w", encoding="utf-8") as f:
        f.write("\n".join(lines) + "\n")


if __name__ == "__main__":
    en_kw, zh_kw = sys.argv[1], sys.argv[2]
    model = sys.argv[3] if len(sys.argv) > 3 else "openai/gpt-6-luna"
    listings, rows = tiktok_listings(en_kw) + xhs_listings(zh_kw), []
    for l in listings:
        fetch = tiktok_reviews if l["platform"] == "tiktok" else xhs_reviews
        got = fetch(l["listing"])
        l["sampled"] = len(got)
        rows += [{"platform": l["platform"], "listing": l["listing"], **r} for r in got]
        print(f"{l['platform']:12} {l['listing']} total={l['reviews_total']} sampled={len(got)}")
    usage = classify(rows, model)
    stats = aggregate(rows)
    with open("reviews.csv", "w", newline="", encoding="utf-8-sig") as f:
        w = csv.DictWriter(f, fieldnames=["platform", "listing", "review_id", "stars", "theme", "sentiment",
                                          "unlabeled", "text"])
        w.writeheader()
        w.writerows(rows)
    write_brief("brief.md", listings, stats, usage, model)
    print(json.dumps(stats, ensure_ascii=False, indent=2))
    print("usage:", usage)

Design choices worth knowing before you change it:

  • The LLM labels; Python counts. The model never sees totals or shares. In our GPT-6.1 Sol agent test, arithmetic over raw tool output was where models slipped, so every number in the brief comes from aggregate.
  • Skipped labels are visible. If the model drops a review from its JSON, the row becomes “other, mixed” with unlabeled set to true, and the brief reports the count. Mixed counts toward “negative or mixed”, so a skipped review nudges that share up slightly.
  • Batches of 20, text cut at 400 characters. 578 reviews took 29 calls. Median review length was 63 characters on TikTok Shop and 21 on Xiaohongshu.

My verbatim end-to-end run (Luna, 2026-10-03 14:11 UTC) took 528 seconds, almost all of it paging reviews. It kept 300 TikTok Shop and 278 Xiaohongshu reviews. Compared with my evaluation fetch about 20 minutes earlier, one of its three Xiaohongshu listings was different and the other two came back under different SKU ids, because search results change between calls.

Does the cheap model label well enough?

I fetched the reviews once (13:50 UTC, 578 reviews), then drew 40 for hand labeling: 20 per platform, of which 8 were 1 to 3 star reviews and 12 were random, seed fixed. That over-samples complaints on purpose, since complaints are what the brief is for. I labeled all 40 myself before any model saw them, with the same written rule as the prompt, and marked 11 where I hesitated between two themes. My labels: 20 positive, 14 negative, 6 mixed; look 9, durability 6, protection 6, magnetic 5, feel 5, fit 4, other 4, shipping 1. One annotator is a limit: someone else would draw some theme lines differently.

Then each model labeled all 578 reviews twice with the program’s classify function.

Against my 40 labelsLuna run 1Luna run 2Sol run 1Sol run 2
Theme matches32/40 (κ 0.770)35/40 (κ 0.854)35/40 (κ 0.855)35/40 (κ 0.855)
… on the 29 clear reviews25/2928/2928/2928/29
… on the 11 ambiguous ones7/117/117/117/11
Sentiment matches (3 classes)36/40 (κ 0.830)36/40 (κ 0.832)38/40 (κ 0.916)39/40 (κ 0.958)
Complaint vs not (negative or mixed)38/4040/4040/4040/40
Theme and sentiment both match31/4032/4034/4035/40
On all 578 reviewsGPT-6 LunaGPT-6.1 Sol
Same theme in both runs88.6%96.9%
Same sentiment in both runs94.3%99.5%
Reviews skipped by the model1 and 20 and 0
Tokens per run (in / out)19,547 / 19,416 to 20,36319,547 / 10,824 to 10,880
Billed per run$0.0117 and $0.0120$0.1479 and $0.1459
Wall time, 29 sequential calls261 to 273 s248 to 250 s

Luna and Sol (first runs) chose the same theme for 85.1% of the 578 reviews and the same sentiment for 95.3%.

All four runs called a magnet plate that came off with the phone “durability”; I’d called it “magnetic”. Both models leaned on “value” or “feel” for long reviews that complain about several things. Those are rule questions, not model failures. Luna’s run 1 had the two errors that matter: it called a complaint about a too-tight case and one about a camera lip that wasn’t raised enough “positive”, and filed them under value and shipping. Those are the mistakes that hide a complaint from the brief.

Billing reconciled with the estimate to within $0.002 per run. All 116 chat-completion ids resolved through GET /v1/tasks/<id>/cost, which needs the same Bearer key as the calls. One Sol batch, as returned:

{"id": "c999e1ed-74f8-4922-8eaf-4095fe564fcd", "status": "completed", "settled": true, "currency": "USD",
 "cost": "0.004890", "estimated_cost": "0.004890",
 "usage": {"prompt_tokens": 700, "completion_tokens": 349, "total_tokens": 1049, "cached_tokens": 0,
           "cache_creation_tokens": 0, "reasoning_tokens": 68}}

My read: Luna is good enough for the aggregate brief, and Sol is worth it when individual labels matter, for example when you route each complaint to a team. Our model tier benchmark found Luna 51/51 on tool tasks at about 1/20 of Sol’s cost per task. On this labeling job it was a little less exact and a little less stable, at about 1/12 of the cost.

What the brief said

The numbers below come from the four evaluation runs plus the end-to-end run, all on 2026-10-03.

TikTok Shop (300 reviews)Xiaohongshu (278 reviews)
1 to 2 star share24.3%1.4%
Negative or mixed (LLM)39.3% to 40.7%11.5% to 14.0%
Most common themelook, 32% to 36%look, about 25% to 32%
Top complaint themesdurability (37 to 40), magnetic (23 to 25), fit (15 to 17)fit, durability, feel at 4 to 8 each; order changed between runs

On both platforms the text found far more complaints than the stars did. On Xiaohongshu, almost every written review had 4 or 5 stars, yet one in eight complained, usually a loose or tight fit, fingerprints and lint on soft materials, or edges that felt sharp. On TikTok Shop, the durability complaints were mostly peeling, chipping or fading within days to weeks, and the magnetic ones were weak magnets dropping chargers or grips. Those are paraphrases of patterns, not quotes.

The Xiaohongshu side is thin once filtered: with 4 to 8 complaints per theme, one relabeled review reorders the top three. On TikTok Shop, the same top three came out in all five runs.

SandBase model page for xiaohongshu/app-v2/product-reviews showing Free base price, sync execution, api model type and 5 input fields, under an API Free Week banner

Caption: The Xiaohongshu product-reviews model page showed a Free base price and 5 input fields on 2026-10-03, under an “API Free Week” banner, so treat the price as dated (captured 2026-10-03).

The data calls cost nothing at the listed price that day. All five data endpoints returned base_price: "0" from GET /v1/models/<model> (same Bearer key; the public model pages show the same prices without one). The page above carried an “API Free Week” promotion banner, so check again before a large run. GPT-6 Luna was listed at $0.10 input and $0.50 output per million tokens, GPT-6.1 Sol at $2 and $10, both below a 272K-token prompt tier this job never reaches.

Limits

Scope: one product category, three listings per platform, 578 reviews fetched once for the evaluation, 40 hand labels by one annotator, two runs per model, one day. The 40-review sample over-weights low-star reviews, so its agreement rates describe complaint-heavy text better than the average review. I didn’t test other categories or TikTok Shop regions. The placeholder pattern and the “Latest” sort value come from observation, not documentation. The raw reviews and labels are kept internally because they’re third-party content; the method, prompt, taxonomy, sampling rule and aggregates are all in this article.

Next steps

  • Change the taxonomy for your category. Keep it under ten themes and keep the “main complaint wins” rule. Hand-label 40 before trusting it.
  • Route by disagreement. Run Luna on everything, then send only reviews where a second Luna run disagrees to Sol. On this data, the two Luna runs disagreed on the theme of about 11% of reviews.
  • Pair it with listing research. The TikTok Shop product research tutorial covers search, detail and review paging on one platform.

To try it, read the product-reviews-v2 reference and the Chat Completions guide, then get a SandBase API key.

FAQ

Why not just use star ratings?

Stars missed most complaints here. On Xiaohongshu, 1.4% of written reviews had 1 or 2 stars, but the classifier flagged 11.5% to 14% as negative or mixed.

Why did Xiaohongshu return only 5-star reviews?

The default sort (sort_strategy_type 0) returned 300 of 300 five-star reviews for three listings. The overview response lists a “Latest” option with type 1, which returned a normal star mix. The reference documents the field but not its values.

How many reviews does one run read?

Up to 100 written reviews for each of three listings per platform. TikTok Shop pages held 20 reviews and Xiaohongshu pages 12, and many are skipped for having no text, so Xiaohongshu needs the most pages. My end-to-end run took 528 seconds, nearly all of it paging; I didn’t log the exact call count.

Is GPT-6 Luna accurate enough?

For aggregate shares and complaint rankings, yes on this data: the TikTok Shop top three never changed. For per-review routing, GPT-6.1 Sol was more exact (35/40 themes both runs) and far more stable (96.9% vs 88.6% run to run).

Can I publish the reviews the agent collects?

Treat them as third-party personal content. Publish counts, shares and paraphrased patterns, not review text or reviewer names.