商品评论分析 Agent:TikTok Shop 和小红书评论自动归类
商品评论分析 Agent 实战:抓 TikTok Shop 和小红书磁吸手机壳评论,LLM 打主题和情感,Python 统计。用 40 条人工标注对比 GPT-6 Luna 和 GPT-6.1 Sol 的准确率与成本。

第一次给三款磁吸手机壳拉小红书评论,300 条全是 5 星。可同一个商品的评价概览明明写着:4,688 个评分里有 74 个是 1~3 星。默认排序把好评排在了前面。换成「最新」排序,星级分布正常了,结果又发现 300 条里有 179 条是平台自动生成的占位文字,买家其实只打了星、一个字没写。这两个坑都不会报错,但都会悄悄把后面的主题统计带歪。
这篇教程做的是一个商品评论分析 Agent:每个平台给一个商品词,它找出评论最多的三个公开商品,每个商品最多拉 100 条有文字的评论,让 LLM 给每条打一个主题和一个情感,然后全部在 Python 里计数,输出一份逐条评论的 CSV 和一页简报。适合想知道买家到底在抱怨什么的卖家和 Agent 开发者。文章的实测部分是分类器本身:我人工标了 40 条评论,对比 openai/gpt-6-luna 和 openai/gpt-6.1-sol 的一致率、稳定性和成本。
先说结论
- 对照 40 条人工标注,GPT-6.1 Sol 两次运行都有 35 条主题和我一致(Cohen’s kappa 0.855);GPT-6 Luna 分别是 32 条和 35 条。在我认为没有歧义的 29 条上,Sol 两次都对 28 条,Luna 是 25 条和 28 条。
- 全量 578 条评论跑一遍,Luna 约 $0.012,Sol 约 $0.147,差 12 倍左右。Luna 的输出 token 是 Sol 的 1.8 倍左右,所以差距没有标价上的 20 倍那么大。
- Sol 更稳:两次运行主题一致 96.9%,Luna 是 88.6%。Luna 每次还会漏标 1~2 条,其中一次把两条抱怨标成了正面。
- 聚合结论几乎不受模型影响。TikTok Shop 的前三大抱怨主题在五次运行里都是耐用性、磁吸、贴合度。小红书每个主题只有 4~8 条抱怨,样本太少,排不出可靠的名次。
这个 Agent 做什么
品类选的是磁吸手机壳。TikTok Shop 美区用英文词 “magsafe iphone case”,小红书用 磁吸手机壳。一开始我想沿用跨平台选品 Agent 里的便携榨汁杯,但查到的小红书榨汁杯商品最多只有 26 条评论;手机壳在两个平台上都是每个商品几百到几千条。
| 步骤 | 端点 | 代码保留什么 |
|---|---|---|
| 找 TikTok Shop 商品 | tiktok/shop-web/search-products-list | 商品 id、标题、评论数,取评论最多的 3 个 |
| TikTok Shop 评论 | tiktok/shop-web/product-reviews-v2 | 评论 id、星级、正文;只打星的跳过 |
| 找小红书商品 | xiaohongshu/app-v2/search-products | 前 8 张商品卡 |
| 按评论数排序 | xiaohongshu/app-v2/product-review-overview | total,取前 3 |
| 小红书评论 | xiaohongshu/app-v2/product-reviews | 评论 id、星级、正文;「最新」排序,占位文字跳过 |
| 主题 + 情感 | openai/gpt-6-luna(默认)或 openai/gpt-6.1-sol,走 /v1/chat/completions | 9 个主题之一,正面 / 负面 / 混合 |
主题表是为手机壳定的:fit(贴合)、protection(防护)、magnetic(磁吸)、look(外观)、feel(手感)、durability(耐用)、value(性价比)、shipping(物流与售后)、other(其他)。比主题表更重要的是打标规则:评论里有抱怨,主题就取主要抱怨;没有抱怨,就取谈得最多的方面;泛泛的夸奖算 other。这条规则让「抱怨排行」有意义,我人工标注时用的也是同一条。
数据边界
只读公开数据,需要 SandBase API Key。SandBase 不是 TikTok 或小红书的官方合作方,不涉及私人账号、私信、商家后台数据、订单或任何账号操作。评论正文和评论者昵称属于第三方个人内容,所以本文只给计数、占比和泛化转述;程序写出的 CSV 请留在本地。
| 需求 | 用什么 |
|---|---|
| 任意商品的公开评论,做调研或竞品分析 | SandBase 公开数据端点(本文) |
| 自己 TikTok Shop 店铺的评论和订单 | TikTok Shop Partner Center |
| 自己小红书店铺的数据 | 小红书开放平台 |
评论接口的两个坑
小红书默认排序是筛过的。 sort_strategy_type 用默认值 0 时,三个商品拉到 300 条,itemScore 全是 5。评价概览接口给这些商品列了两个排序选项:Default(type 0)和 Latest(type 1)。换成 1 之后,最大那个商品的前 36 条里出现了 1 条 2 星、1 条 3 星和 7 条 4 星;type 2 和 3 什么都没返回。所以 Agent 固定传 sort_strategy_type: 1。

截图:product-reviews 参考页写了 sort_strategy_type 默认为 0,但没列出可选值,所以「最新」= 1 是从概览返回和实测里得出的(2026-10-03 截取)。
很多小红书「评论」根本没有买家写的字。 「最新」排序下,300 条里有 179 条是系统文字,意思分别是「该用户给了 5 星好评」「该用户给了好评」「该用户认为商品一般」。LLM 会老老实实把它们归进「其他、正面」,把所有占比都稀释掉。代码用 ^该用户(给出|认为) 把它们剔除。这是我观察出来的规律,不是文档里的字段,数字看着不对劲时先查它。TikTok Shop 也有类似情况,只是简单些:只打星的评论 review_text 为空,代码直接跳过。
TikTok Shop 的默认排序(sort_rule 2)看起来没有筛过:300 条里 73 条是 1~2 星。参考页没说明 sort_rule 其他取值的含义,我就保留默认。

截图:product-reviews-v2 参考页列出了 page_start、product_id、region 和 sort_rule(默认 2),Agent 翻 TikTok Shop 评论用的就是这几个字段(2026-10-03 截取)。
实测记录:2026-10-03(UTC)
输入:上面两个通用商品词、TikTok Shop 美区,以及搜索返回的公开商品 id。每个数据调用都是 POST https://api.sandbase.ai/v1/api/<vendor>/<path>,带 Authorization: Bearer $SANDBASE_API_KEY,body 只放该端点自己的参数。文档里的数据位置是 outputs[0].data;有报告说部分端点返回的结构不一样,所以代码里的读取函数优先读 outputs[0].data,读不到再退到 outputs[0],再退到顶层 output,状态不是 completed 就直接抛错。本文五个数据端点在我记录的每一次调用里都是 outputs[0].data。下面的字段名都是实测观察到的,不是文档保证,所以代码一律用 .get() 读。
curl -s https://api.sandbase.ai/v1/api/xiaohongshu/app-v2/product-review-overview \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"sku_id": "68cc0e65d1f7b800015dacdc"}'
run d3568c3c-2df6-4833-9e75-afc0118319ca 在 outputs[0].data.data 里返回了下面这些(其余键省略):
{"total": 4688, "avgScore": "4.5",
"scoreToCnt": {"1": 23, "2": 12, "3": 39, "4": 352, "5": 4262},
"sortStrategyList": [{"name": "Default", "type": 0}, {"name": "Latest", "type": 1}]}
同一商品的评论请求 {"sku_id": "68cc0e65d1f7b800015dacdc", "page": 0, "sort_strategy_type": 1}(run d272e435-016b-4556-bd96-f074720f201e)返回 12 条评论和 hasMore: true。第一条的 reviewInfo,正文、id、图片、时间和用户字段已省略:
{"itemScore": 5, "logisticsScore": 5, "serviceScore": 5}
TikTok Shop 请求 {"product_id": "1730831546437308868", "region": "US", "page_start": 1}(run eaee7cd8-c911-48d3-80b2-9643d08c140d)返回 20 条 product_reviews 和 has_more: true。第一条,正文、评论者、id、SKU 和图片字段已省略:
{"product_id": "1730831546437308868", "review_rating": 5, "is_verified_purchase": true,
"is_incentivized_review": false, "review_country": "US"}
完整程序
单文件 194 行,标准库加 requests。运行 python3 review_agent.py "magsafe iphone case" "磁吸手机壳",第三个参数可以换成 openai/gpt-6.1-sol。下面就是我实际跑的那份文件。
#!/usr/bin/env python3
"""Product review analysis agent: TikTok Shop + Xiaohongshu reviews -> themes, sentiment, CSV, brief.
Usage: python3 review_agent.py "magsafe iphone case" "磁吸手机壳" [openai/gpt-6-luna]
Writes reviews.csv (one row per review) and brief.md (aggregates per platform and theme).
"""
import csv
import json
import os
import re
import sys
import time
from collections import Counter
import requests
BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
THEMES = ["fit", "protection", "magnetic", "look", "feel", "durability", "value", "shipping", "other"]
SENTIMENTS = ["positive", "negative", "mixed"]
PRICE = {"openai/gpt-6-luna": (0.1, 0.5), "openai/gpt-6.1-sol": (2.0, 10.0)} # USD per 1M tokens, 2026-10-03
def call(model: str, params: dict, retries: int = 2) -> dict:
"""POST /v1/api/<model>. Prefer outputs[0].data, fall back to outputs[0], then output."""
for attempt in range(retries + 1):
try:
resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
body = resp.json()
except requests.RequestException as err:
if attempt == retries:
raise RuntimeError(f"{model}: {err}") from err
time.sleep(3 * (attempt + 1))
continue
if resp.status_code >= 500 and attempt < retries:
time.sleep(3 * (attempt + 1))
continue
if resp.status_code != 200 or body.get("status") != "completed":
raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
first = (body.get("outputs") or [None])[0]
if isinstance(first, dict):
return first["data"] if "data" in first else first
if body.get("output") is not None:
return body["output"]
raise RuntimeError(f"{model}: completed without a payload")
raise RuntimeError(f"{model}: retries exhausted")
def tiktok_listings(keyword: str, n: int = 3) -> list[dict]:
for _attempt in range(4): # a short fallback list (has_more false) comes back often; retry it
page = call("tiktok/shop-web/search-products-list", {"search_word": keyword, "region": "US"}).get("data") or {}
if page.get("has_more"):
break
time.sleep(2)
rows = [{"platform": "tiktok", "listing": p["product_id"], "title": (p.get("title") or "")[:60],
"reviews_total": int((p.get("rate_info") or {}).get("review_count") or 0)}
for p in page.get("products") or [] if p.get("product_id")]
return sorted(rows, key=lambda r: -r["reviews_total"])[:n]
def tiktok_reviews(product_id: str, limit: int = 100, max_pages: int = 10) -> list[dict]:
out = []
for page in range(1, max_pages + 1):
data = call("tiktok/shop-web/product-reviews-v2",
{"product_id": product_id, "region": "US", "page_start": page}).get("data") or {}
for r in data.get("product_reviews") or []:
if (r.get("review_text") or "").strip(): # star-only reviews have nothing to classify
out.append({"review_id": r.get("review_id"), "stars": r.get("review_rating"),
"text": r["review_text"].strip()})
if len(out) >= limit or not data.get("has_more"):
break
return out[:limit]
def xhs_listings(keyword: str, n: int = 3, candidates: int = 8) -> list[dict]:
data = call("xiaohongshu/app-v2/search-products", {"keyword": keyword}).get("data") or {}
cards = [m.get("content") or {} for m in (data.get("module") or {}).get("data") or []
if m.get("card_name") == "cosmos_search_goods_card"][:candidates]
rows = []
for c in cards: # search cards carry no review count, so ask the overview endpoint
ov = call("xiaohongshu/app-v2/product-review-overview", {"sku_id": c["id"]}).get("data") or {}
rows.append({"platform": "xiaohongshu", "listing": c["id"], "title": (c.get("title") or "")[:30],
"reviews_total": int(ov.get("total") or 0)})
return sorted(rows, key=lambda r: -r["reviews_total"])[:n]
XHS_PLACEHOLDER = re.compile(r"^该用户(给出|认为)") # auto text for star-only reviews ("this user gave 5 stars")
def xhs_reviews(sku_id: str, limit: int = 100, max_pages: int = 25) -> list[dict]:
out = []
for page in range(max_pages): # pages start at 0; sort 1 = "Latest" (default order was all 5-star)
data = call("xiaohongshu/app-v2/product-reviews",
{"sku_id": sku_id, "page": page, "sort_strategy_type": 1}).get("data") or {}
for r in data.get("reviews") or []:
info = r.get("reviewInfo") or {}
text = (info.get("content") or "").strip()
if text and not XHS_PLACEHOLDER.match(text):
out.append({"review_id": info.get("reviewId"), "stars": info.get("itemScore"),
"text": text})
if len(out) >= limit or not data.get("hasMore"):
break
return out[:limit]
PROMPT = """You label product reviews of phone cases. For each numbered review return one primary theme and a sentiment.
Themes: fit (fits the phone model, cutouts, buttons), protection (drops, camera or screen coverage),
magnetic (MagSafe or magnet strength, wireless charging, built-in stand or ring), look (design, color, clarity,
matches photos), feel (grip, texture, thickness, material feel), durability (yellowing, cracking, peeling, wear),
value (price, worth the money), shipping (delivery, packaging, seller service, wrong item), other (generic praise
or anything else). If the review complains about something, the theme is the main complaint; otherwise it is the
aspect discussed most. Sentiment: positive, negative or mixed. Reviews may be English or Chinese.
Reply with JSON only: {"labels": [{"i": 0, "theme": "fit", "sentiment": "positive"}]}"""
def classify(reviews: list[dict], model: str, batch: int = 20) -> dict:
"""Adds theme/sentiment to each review in place; returns token usage and cost."""
usage = Counter()
for start in range(0, len(reviews), batch):
chunk = reviews[start:start + batch]
numbered = "\n".join(f"{i}. {r['text'][:400]}" for i, r in enumerate(chunk))
resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=180,
json={"model": model, "max_tokens": 4000,
"messages": [{"role": "system", "content": PROMPT},
{"role": "user", "content": numbered}]})
resp.raise_for_status()
body = resp.json()
for k in ("prompt_tokens", "completion_tokens"):
usage[k] += (body.get("usage") or {}).get(k, 0)
usage["calls"] += 1
match = re.search(r"\{.*\}", body["choices"][0]["message"]["content"] or "", re.S)
labels = {x.get("i"): x for x in (json.loads(match.group(0)).get("labels") or [])} if match else {}
for i, r in enumerate(chunk):
lab = labels.get(i) or {}
r["theme"] = lab.get("theme") if lab.get("theme") in THEMES else "other"
r["sentiment"] = lab.get("sentiment") if lab.get("sentiment") in SENTIMENTS else "mixed"
r["unlabeled"] = not lab # the model skipped this one; counted, never hidden
pin, pout = PRICE[model]
usage["usd"] = round((usage["prompt_tokens"] * pin + usage["completion_tokens"] * pout) / 1e6, 5)
return dict(usage)
def aggregate(rows: list[dict]) -> dict:
out = {}
for platform in sorted({r["platform"] for r in rows}):
rs = [r for r in rows if r["platform"] == platform]
neg = [r for r in rs if r["sentiment"] in ("negative", "mixed")]
themes = Counter(r["theme"] for r in rs)
neg_themes = Counter(r["theme"] for r in neg)
out[platform] = {
"reviews": len(rs), "low_star_share": round(sum(1 for r in rs if (r["stars"] or 5) <= 2) / len(rs), 3),
"negative_or_mixed_share": round(len(neg) / len(rs), 3),
"unlabeled": sum(1 for r in rs if r["unlabeled"]),
"theme_share": {t: round(themes[t] / len(rs), 3) for t in THEMES if themes[t]},
"top_complaints": [(t, c) for t, c in neg_themes.most_common(3)],
}
return out
def write_brief(path: str, listings: list[dict], stats: dict, usage: dict, model: str):
lines = [f"# Review brief ({time.strftime('%Y-%m-%d %H:%M UTC', time.gmtime())}, {model})", "",
"| Platform | Listing | Reviews on listing | Sampled |", "|---|---|---|---|"]
lines += [f"| {l['platform']} | {l['listing']} | {l['reviews_total']} | {l['sampled']} |" for l in listings]
for platform, s in stats.items():
lines += ["", f"## {platform}: {s['reviews']} reviews, {s['negative_or_mixed_share']:.0%} negative or mixed, "
f"{s['low_star_share']:.0%} at 1-2 stars", "", "| Theme | Share of reviews |", "|---|---|"]
lines += [f"| {t} | {v:.0%} |" for t, v in sorted(s["theme_share"].items(), key=lambda kv: -kv[1])]
lines += ["", "Top complaint themes: " + ", ".join(f"{t} ({c})" for t, c in s["top_complaints"])]
lines += ["", f"LLM usage: {usage['calls']} calls, {usage['prompt_tokens']} in / "
f"{usage['completion_tokens']} out tokens, about ${usage['usd']}"]
with open(path, "w", encoding="utf-8") as f:
f.write("\n".join(lines) + "\n")
if __name__ == "__main__":
en_kw, zh_kw = sys.argv[1], sys.argv[2]
model = sys.argv[3] if len(sys.argv) > 3 else "openai/gpt-6-luna"
listings, rows = tiktok_listings(en_kw) + xhs_listings(zh_kw), []
for l in listings:
fetch = tiktok_reviews if l["platform"] == "tiktok" else xhs_reviews
got = fetch(l["listing"])
l["sampled"] = len(got)
rows += [{"platform": l["platform"], "listing": l["listing"], **r} for r in got]
print(f"{l['platform']:12} {l['listing']} total={l['reviews_total']} sampled={len(got)}")
usage = classify(rows, model)
stats = aggregate(rows)
with open("reviews.csv", "w", newline="", encoding="utf-8-sig") as f:
w = csv.DictWriter(f, fieldnames=["platform", "listing", "review_id", "stars", "theme", "sentiment",
"unlabeled", "text"])
w.writeheader()
w.writerows(rows)
write_brief("brief.md", listings, stats, usage, model)
print(json.dumps(stats, ensure_ascii=False, indent=2))
print("usage:", usage)
改代码之前值得知道的几个取舍:
- LLM 只打标,计数交给 Python。 模型看不到任何总数或占比。我们在 GPT-6.1 Sol Agent 实测里发现,模型最容易翻车的就是对原始工具输出做算术,所以简报里每个数字都出自
aggregate。 - 漏标不藏。 模型返回的 JSON 漏了某条,这条就记为「other、mixed」,
unlabeled置为 true,简报里会报出数量。mixed 计入「负面或混合」,所以漏标会让这个占比略微偏高。 - 每批 20 条,正文截到 400 字符。 578 条用了 29 次调用。评论普遍很短:TikTok Shop 中位数 63 个字符,小红书 21 个。
我照原样端到端跑了一次(Luna,2026-10-03 14:11 UTC),耗时 528 秒,几乎全花在翻评论页上,最后保留 TikTok Shop 300 条、小红书 278 条。和大约 20 分钟前的评测抓取相比,小红书三个商品里有一个换了,另外两个返回的是不同的 SKU id,因为搜索结果每次都会变。
便宜模型打标够不够用?
评测用的评论只抓一次(13:50 UTC,共 578 条),从中抽 40 条人工标注:每个平台 20 条,其中 8 条是 1~3 星,12 条随机,随机种子固定。这样做是故意多抽差评,因为简报要看的正是抱怨。40 条全部由我在模型看到之前标完,规则和 prompt 里写的一样,其中 11 条我在两个主题之间犹豫过,单独标记为「有歧义」。我的标注结果:正面 20、负面 14、混合 6;外观 9、耐用 6、防护 6、磁吸 5、手感 5、贴合 4、其他 4、物流 1。只有一个标注者是个局限,换个人来,有些主题边界会划得不一样。
然后两个模型各用程序里的 classify 把 578 条全部标两遍。
| 对照我的 40 条标注 | Luna 第 1 次 | Luna 第 2 次 | Sol 第 1 次 | Sol 第 2 次 |
|---|---|---|---|---|
| 主题一致 | 32/40(κ 0.770) | 35/40(κ 0.854) | 35/40(κ 0.855) | 35/40(κ 0.855) |
| 其中 29 条无歧义 | 25/29 | 28/29 | 28/29 | 28/29 |
| 其中 11 条有歧义 | 7/11 | 7/11 | 7/11 | 7/11 |
| 情感一致(三分类) | 36/40(κ 0.830) | 36/40(κ 0.832) | 38/40(κ 0.916) | 39/40(κ 0.958) |
| 是否抱怨(负面或混合) | 38/40 | 40/40 | 40/40 | 40/40 |
| 主题和情感都一致 | 31/40 | 32/40 | 34/40 | 35/40 |
| 全部 578 条 | GPT-6 Luna | GPT-6.1 Sol |
|---|---|---|
| 两次运行主题相同 | 88.6% | 96.9% |
| 两次运行情感相同 | 94.3% | 99.5% |
| 模型漏标条数 | 1 和 2 | 0 和 0 |
| 每次 token(输入 / 输出) | 19,547 / 19,416~20,363 | 19,547 / 10,824~10,880 |
| 每次实际计费 | $0.0117 和 $0.0120 | $0.1479 和 $0.1459 |
| 耗时(29 次顺序调用) | 261~273 秒 | 248~250 秒 |
Luna 和 Sol(各取第一次)在 578 条里主题相同 85.1%,情感相同 95.3%。
分歧本身比总分更有看头。四次运行都把「磁吸片跟着手机一起掉下来」归成了耐用性,我归的是磁吸;遇到一条评论抱怨好几件事的长评,两个模型都偏向性价比或手感。这些是规则怎么定的问题,算不上模型出错。真正要命的是 Luna 第一次运行里的两处:一条嫌壳太紧、一条嫌镜头边不够高,都被标成了正面,主题还分别落到了性价比和物流。这类错误会把抱怨从简报里藏掉。
计费和估算每次相差不到 $0.002。116 个 chat completion id 全部能通过 GET /v1/tasks/<id>/cost 查到(要用和调用时同一把 Bearer Key)。Sol 的一批,原样返回如下:
{"id": "c999e1ed-74f8-4922-8eaf-4095fe564fcd", "status": "completed", "settled": true, "currency": "USD",
"cost": "0.004890", "estimated_cost": "0.004890",
"usage": {"prompt_tokens": 700, "completion_tokens": 349, "total_tokens": 1049, "cached_tokens": 0,
"cache_creation_tokens": 0, "reasoning_tokens": 68}}
我的判断:做聚合简报,Luna 够用;如果每条标签都要落到具体动作上,比如把每条抱怨分派给不同团队,Sol 值这个钱。我们的模型分档实测里,Luna 在工具调用题上 51/51 全对,单任务成本约为 Sol 的 1/20;换到打标这件事上,它略欠精确、略欠稳定,成本约为 1/12。
简报里说了什么
下表来自 2026-10-03 的四次评测运行加一次端到端运行。
| TikTok Shop(300 条) | 小红书(278 条) | |
|---|---|---|
| 1~2 星占比 | 24.3% | 1.4% |
| 负面或混合(LLM) | 39.3%~40.7% | 11.5%~14.0% |
| 最常见主题 | 外观,32%~36% | 外观,约 25%~32% |
| 前三大抱怨主题 | 耐用(37~40)、磁吸(23~25)、贴合(15~17) | 贴合、耐用、手感各 4~8 条,名次每次都变 |
两个平台上,文字里的抱怨都远多于星级反映出来的。小红书几乎每条有文字的评论都是 4~5 星,但每八条里就有一条在抱怨,常见的是太松或太紧、软材质沾指纹沾毛、边框割手。TikTok Shop 的耐用性抱怨多是几天到几周内掉漆、崩角、褪色,磁吸类则是磁力弱、充电器和支架吸不住。以上是对规律的转述,不是原文引用。
小红书这边过滤完就很薄了:每个主题只有 4~8 条抱怨,改一条标注,前三名就会换位。TikTok Shop 的前三名在五次运行里一次没变。

截图:2026-10-03 小红书 product-reviews 模型页显示基础价格 Free、5 个输入字段,页顶挂着「API Free Week」横幅,价格请按当天看待(2026-10-03 截取)。
当天按标价,数据调用不花钱:五个数据端点用 GET /v1/models/<model> 查到的都是 base_price: "0"(同样要带 Bearer Key;公开的模型页不登录也能看到价格)。上图页面挂着「API Free Week」活动横幅,大批量跑之前请再确认一次。GPT-6 Luna 标价每百万 token 输入 $0.10、输出 $0.50,GPT-6.1 Sol 是 $2 和 $10,这是 272K prompt token 以下的档位,这个任务远远碰不到上一档。
局限
测试范围:一个品类,每个平台三个商品,评测用的 578 条评论只抓了一次,40 条人工标注只有一个标注者,每个模型跑两次,只测了一天。40 条样本刻意多抽了差评,所以一致率更能代表「抱怨多的评论」,不代表平均水平。其他品类和 TikTok Shop 其他地区都没测。占位文字的规律和「最新」排序的取值都来自观察,不是文档。原始评论和标注属于第三方内容,只在内部保留;方法、prompt、主题表、抽样规则和聚合结果都已写在本文里。
下一步
- 按你的品类改主题表。 控制在十个以内,保留「有抱怨就取主要抱怨」这条规则。正式用之前先人工标 40 条。
- 按分歧分流。 全量用 Luna 跑,只把两次 Luna 结果不一致的评论交给 Sol。在这批数据上,两次 Luna 的主题分歧大约占 11%。
- 和选品调研配合。 TikTok Shop 选品调研教程讲了单平台的搜索、详情和评论翻页。
想动手,先看 product-reviews-v2 参考文档和 Chat Completions 指南,然后获取 SandBase API Key。
常见问题
直接看星级不行吗?
这次数据里,星级漏掉了大部分抱怨。小红书有文字的评论只有 1.4% 是 1~2 星,分类器却标出 11.5%~14% 是负面或混合。
为什么小红书只返回 5 星评论?
默认排序(sort_strategy_type 0)下,三个商品 300 条评论全是 5 星。概览接口列出了 type 1 的「Latest」选项,用它就能拿到正常的星级分布。参考文档写了这个字段,但没写取值。
跑一次要读多少评论?
每个平台三个商品,每个商品最多 100 条有文字的评论。TikTok Shop 每页 20 条,小红书每页 12 条,大量没有文字的会被跳过,所以小红书要翻的页最多。我那次端到端运行耗时 528 秒,几乎都在翻页,具体调用次数没有记录。
GPT-6 Luna 准确度够吗?
看聚合占比和抱怨排行,这批数据上够用:TikTok Shop 的前三名从没变过。要逐条分派的话,GPT-6.1 Sol 更准(两次都是 35/40),也稳得多(两次一致率 96.9% 对 88.6%)。
Agent 抓到的评论能直接发出来吗?
把它们当作第三方个人内容。可以发计数、占比和转述后的规律,不要发评论原文或评论者昵称。