舆情监测 Agent:微博、知乎、B站,便宜模型够不够用
用 Python 写一个舆情监测 Agent:微博、知乎、B站三平台抓同一话题,大模型打立场和主题标签,代码做统计。再用 50 条人工标注对比 GPT-6 Luna 和 GPT-6.1 Sol 的准确率、成本与耗时。

同一批 78 条华为 Mate 90 相关的帖子、回答和评论,我分别交给两个模型打标签。GPT-6.1 Sol 说 B站在话题内的评论里 45% 是正面的;GPT-6 Luna 说只有 17%,还把 B站写成了「主要风险区」。内容、打标签的提示词、代码完全一样,两份简报对 B站的判断却差得很远。微博加知乎的 38 条,两个模型有 33 条立场一致;B站的 40 条,只有 26 条一致。
「便宜模型够不够用」,答案就藏在这个差距里。本文搭一个舆情监测 Agent(微博、知乎、B站):针对一个消费科技话题,先看它在三个平台热榜上的位置,再拉公开的微博帖子、知乎回答和 B站评论,用大模型给每条内容打立场和主题,在 Python 里汇总,最后生成一页简报。再用我手工标注的 50 条对比 GPT-6 Luna 和 GPT-6.1 Sol。适合想每天看新品口碑、又不想每条标签都付旗舰价的产品、公关和研究团队。
先说结论
- 一把 SandBase API Key 覆盖 7 个公开数据端点(三个热榜、微博实时搜索、知乎问题回答、B站搜索和评论)以及两个大模型。2026-10-03 这些数据端点在目录里标为 Free,当时页面挂着「API Free Week」横幅。
- 50 条人工标注上,GPT-6 Luna 三次运行的立场准确率是 70% 到 80%(Cohen’s kappa 0.56 到 0.70);GPT-6.1 Sol 是 78% 到 84%(kappa 0.67 到 0.76)。主题准确率分别是 80% 到 82% 和 86% 到 90%。
- 一份简报 Luna 约 0.003 美元,Sol 约 0.03 到 0.04 美元,约 1/11 到 1/12。耗时差不多,每份简报大模型时间约 40 到 56 秒。
- 微博和知乎用 Luna 足够。B站评论里反讽和玩梗多,Luna 把夸奖读成批评的次数多到足以改变结论。要么跑两遍,要么 B站交给 Sol。
这个 Agent 做什么
话题选的是华为 Mate 90 系列,测试前几天刚发布。2026-10-03 当天,知乎上一个关于该机芯片解析视频的问题排在热榜第 12 到 13 位,B站热搜里对应关键词排在第 21 到 25 位。选它是因为它属于消费电子,多个平台都有热度,且不涉及时政或敏感议题。代码里搜索词用 华为Mate90,热榜标题匹配用别名 Mate 90 和 Mate90。
| 步骤 | 端点 | 代码保留什么 |
|---|---|---|
| 看热度 | weibo/web-v2/hot-search、zhihu/web/hot-list、bilibili/web/hot-search | 标题命中别名的条目位置;知乎问题 id |
| 微博帖子 | weibo/web-v2/realtime-search | 2 页,帖子 id、正文、点赞数 |
| 知乎回答 | zhihu/web/question-answers | 热榜问题下 20 条回答的 id、正文、赞同数 |
| B站视频 | bilibili/web/general-search | 播放量最高的 2 个相关视频 |
| B站评论 | bilibili/web/video-comments | 每个视频第 1 页评论的 id、正文、点赞数 |
| 打标签、写简报 | openai/gpt-6-luna 或 openai/gpt-6.1-sol,走 /v1/chat/completions | 每条的立场和主题;120 到 180 词简报 |
立场四选一:positive、negative、neutral(陈述事实、提问或褒贬参半)、off_topic。主题八选一:chip_performance、camera、price_value、design_build、battery_charging、software_system、availability_sales、other。所有数字都由 Python 算:话题内条数、正负面占比、前三主题、负面内容拿到的点赞占比。大模型只根据这份 JSON 写简报,代码再检查简报里有没有 JSON 之外的数字。
它和我们之前的微博加抖音舆情监测 Agent 不同:那篇按关键词长期追踪,这篇是单话题三平台简报,外加模型选型实测。想在单个平台挖深,可以看知乎问答挖掘和 B站视频评论分析。
数据边界
只读公开数据,需要 SandBase API Key。SandBase 不是微博、知乎或 B站的官方合作方。不碰私有账号、私信、账号后台数据,也不做任何账号操作。帖子和评论是个人写的,所以本文只放汇总结果和泛化转述,不贴原文、不出现用户名。程序生成的 labels.csv 只有 id、点赞数和标签;items.json 里有模型看过的原文,请只留在内部。
| 需求 | 用什么 |
|---|---|
| 看某个产品或发布会的公开讨论 | SandBase 公开数据端点(本文) |
| 自己微博账号的帖子和粉丝数据 | 微博开放平台 |
| 自己 B站频道的数据 | B站开放平台 |
实测记录:2026-10-03(UTC)
每个请求都是 POST https://api.sandbase.ai/v1/api/<vendor>/<path>,带 Authorization: Bearer $SANDBASE_API_KEY,body 只放该端点自己的字段。参考文档把数据放在 outputs[0].data,但有些端点被观察到放在别处,所以读取函数优先取 outputs[0].data,取不到再退到 outputs[0],最后才看顶层 output;状态不是 completed 就直接报错。当天我记录的每一次调用,这 7 个端点都返回在 outputs[0].data。下面提到的字段名都来自这些调用,属于实测观察,不是文档保证,所以代码一律用 .get() 读。
curl -s https://api.sandbase.ai/v1/api/bilibili/web/hot-search \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"limit": 30}'
运行 a8eec5b3-ae9f-46a7-8500-94c6e8415aaa(12:44 UTC)返回 {"id", "status": "completed", "model", "outputs"},outputs[0].data 下面的列表在 data.trending.list。命中的那一条,完整如下:
{"goto": "", "heat_score": 438461, "icon": "", "keyword": "华为Mate90系列韬定律解析", "show_name": "华为Mate90系列韬定律解析", "uri": ""}
zhihu/web/hot-list 传 {"limit": "50"}(运行 eb994e55-fe36-42a4-959e-d80e80796f4e),条目在 data 下,每条有个 target。命中问题的 target 一共 12 个键,这里只展示 4 个(author、excerpt 等其余键省略):
{"id": 2089412797440455404, "answer_count": 178, "follower_count": 447, "created": 1790934830}
Agent 把这个 id 传给 question-answers:{"question_id": "2089412797440455404", "limit": 20}。运行 30eb337a-c6e2-41e6-92a8-5de83df30245 返回 20 条,每条 target 里有 content、voteup_count、comment_count,另有一个 paging 对象,is_end: false,带一个 next 链接。我从链接里解析出 cursor、offset、session_id 再请求一次(运行 9ba6fdcf-3b01-4188-9379-22c3caa50589),状态 completed,但 0 条回答。所以程序只读第 1 页。
weibo/web-v2/realtime-search 传 {"query": "华为Mate90", "page": 1},结果在 data.parsed_data.results,我的调用里每页 9 到 10 条,每条有 weibo_id、content,以及包含 comment_count、like_count、repost_count 的 interaction。抓到的微博全都是 0 赞(发出才几分钟),所以「负面点赞占比」这一项实际上只反映知乎和 B站。第 2、3 页能拿到新的帖子 id。
最不稳的是 bilibili/web/video-comments,参数 {"bv_id": "<视频 id>", "pn": 1}。对播放量最高的那个视频,有一次调用(运行 93cbc90a-332c-4b97-8206-3b3d4889e926)只回了 3 条评论,data.page 却写着:
{"acount": 26460, "count": 26460, "num": 1, "size": 20}
之后再请求 pn: 1 回了 20 条,pn: 2 回 0 条且 page 各字段全是 0,pn: 3 又回 20 条。所以程序对第 1 页最多重试三次,直到拿到至少 10 条,不再往后翻。
![SandBase weibo/web-v2/realtime-search 端点参考页,显示 POST 路由、query 和 page 字段以及 outputs[0].data 响应信封](https://static.sandbase.ai/blog/screenshots/public-opinion-agent-weibo-zhihu-bilibili-cheap-model/weibo-realtime-ref-d89e83d3363e.webp)
截图:微博实时搜索参考页写明了 POST 路由、必填的 query 和可选的 page,响应信封里 outputs[0].data 的示例是空的,所以上文的帖子字段都来自实测调用(2026-10-03 截取)。

截图:知乎问题回答参考页列出了用于翻页的 cursor、offset 和 session_id;我实测第二页返回为空,所以 Agent 只读第一页 20 条回答(2026-10-03 截取)。
完整程序
一个文件,标准库加 requests。第一次运行会抓数据并存成 items.json;把这个文件作为第四个参数传进去,就能让另一个模型给完全相同的内容打标签,我的对比就是这么做的。
export SANDBASE_API_KEY=... # 在 shell 里设置,别写进文件
python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6-luna
python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6.1-sol items.json
#!/usr/bin/env python3
"""Three-platform public opinion brief (Weibo, Zhihu, Bilibili) for one consumer/tech topic.
Usage: python3 opinion_brief.py <search keyword> <aliases, comma-separated> <model> [items.json]
Example: python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6-luna
Pass a saved items.json as the 4th argument to skip fetching and reuse the same items.
Writes items.json (with text, keep private), labels.csv (ids + labels only) and brief.md.
"""
import csv
import html
import json
import os
import re
import sys
import time
from collections import Counter
import requests
BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
PRICES = {"openai/gpt-6-luna": (0.1, 0.5), "openai/gpt-6.1-sol": (2.0, 10.0)} # USD per 1M tokens, 2026-10-03
STANCES = ["positive", "negative", "neutral", "off_topic"]
THEMES = ["chip_performance", "camera", "price_value", "design_build", "battery_charging",
"software_system", "availability_sales", "other"]
def call(model: str, params: dict, retries: int = 2) -> dict:
"""POST /v1/api/<model>. Prefer outputs[0].data, fall back to outputs[0], then output."""
for attempt in range(retries + 1):
try:
resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
body = resp.json()
except requests.RequestException as err:
if attempt == retries:
raise RuntimeError(f"{model}: {err}") from err
time.sleep(3 * (attempt + 1))
continue
if resp.status_code >= 500 and attempt < retries:
time.sleep(3 * (attempt + 1))
continue
if resp.status_code != 200 or body.get("status") != "completed":
raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
outputs = body.get("outputs") or []
if outputs and isinstance(outputs[0], dict):
return outputs[0].get("data", outputs[0]) or {}
if body.get("output") is not None:
return body["output"]
raise RuntimeError(f"{model}: completed but no payload")
raise RuntimeError(f"{model}: retries exhausted")
def matches(text: str, aliases: list[str]) -> bool:
t = (text or "").lower().replace(" ", "")
return any(a.lower().replace(" ", "") in t for a in aliases)
def clean(text: str, limit: int = 280) -> str:
text = re.sub(r"<[^>]+>", " ", html.unescape(text or ""))
return re.sub(r"\s+", " ", text).strip()[:limit]
def trending(aliases: list[str]) -> dict:
"""Where the topic sits on each platform's hot list right now (rank 1 = top)."""
wb = call("weibo/web-v2/hot-search", {}).get("realtime") or []
zh = (call("zhihu/web/hot-list", {"limit": "50"}).get("data") or [])
bl = ((call("bilibili/web/hot-search", {"limit": 30}).get("data") or {}).get("trending") or {}).get("list") or []
hits = {"weibo": [i + 1 for i, r in enumerate(wb) if matches(r.get("word"), aliases)],
"zhihu": [i + 1 for i, r in enumerate(zh) if matches((r.get("target") or {}).get("title"), aliases)],
"bilibili": [i + 1 for i, r in enumerate(bl) if matches(r.get("keyword"), aliases)]}
zhihu_qid = next((str(r["target"]["id"]) for r in zh
if matches((r.get("target") or {}).get("title"), aliases)), None)
return {"hot_list_ranks": hits, "zhihu_question_id": zhihu_qid}
def collect(keyword: str, aliases: list[str]) -> tuple[dict, list[dict]]:
trend, items = trending(aliases), []
for page in (1, 2):
data = call("weibo/web-v2/realtime-search", {"query": keyword, "page": page})
for r in (data.get("parsed_data") or {}).get("results") or []:
inter = r.get("interaction") or {}
items.append({"id": f"wb:{r.get('weibo_id')}", "platform": "weibo", "text": clean(r.get("content")),
"likes": int(inter.get("like_count") or 0)})
if trend["zhihu_question_id"]:
data = call("zhihu/web/question-answers", {"question_id": trend["zhihu_question_id"], "limit": 20})
for a in data.get("data") or []:
t = a.get("target") or {}
items.append({"id": f"zh:{t.get('id')}", "platform": "zhihu",
"text": clean(t.get("content") or t.get("excerpt")), "likes": int(t.get("voteup_count") or 0)})
data = call("bilibili/web/general-search", {"keyword": keyword, "order": "totalrank", "page": 1, "page_size": 20})
videos = [v for v in (data.get("data") or {}).get("result") or [] if matches(clean(v.get("title")), aliases)]
for v in sorted(videos, key=lambda v: int(v.get("play") or 0), reverse=True)[:2]:
for _ in range(3): # page 1 sometimes comes back short or empty; retry it
replies = (call("bilibili/web/video-comments", {"bv_id": v["bvid"], "pn": 1}).get("data") or {}).get("replies") or []
if len(replies) >= 10:
break
time.sleep(2)
for c in replies:
items.append({"id": f"bl:{c.get('rpid')}", "platform": "bilibili",
"text": clean((c.get("content") or {}).get("message")), "likes": int(c.get("like") or 0)})
items = [i for i in {i["id"]: i for i in items}.values() if i["text"]] # dedupe, drop empty
return trend, items
def chat(model: str, prompt: str, max_tokens: int) -> tuple[str, dict, float, str]:
t0 = time.time()
resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=300,
json={"model": model, "max_tokens": max_tokens,
"messages": [{"role": "user", "content": prompt}]})
resp.raise_for_status()
body = resp.json()
return body["choices"][0]["message"]["content"] or "", body.get("usage") or {}, time.time() - t0, body.get("id")
def classify(model: str, topic: str, items: list[dict], batch: int = 25) -> tuple[dict, list[dict]]:
labels, calls = {}, []
for start in range(0, len(items), batch):
chunk = [{"id": i["id"], "text": i["text"]} for i in items[start:start + batch]]
prompt = (f"Topic: {topic}. Label each post for public opinion monitoring.\n"
f"stance toward the topic product: one of {STANCES} (neutral = factual, question or mixed; "
f"off_topic = not about the product).\ntheme: the main aspect, one of {THEMES}.\n"
'Reply with JSON only: {"labels": [{"id": "...", "stance": "...", "theme": "..."}]}\n'
+ json.dumps(chunk, ensure_ascii=False))
text, usage, secs, rid = chat(model, prompt, 4000)
calls.append({"step": "classify", "id": rid, "usage": usage, "seconds": round(secs, 1)})
found = re.search(r"\{.*\}", text, re.S)
for lab in (json.loads(found.group(0)).get("labels") if found else []) or []:
if lab.get("stance") in STANCES and lab.get("theme") in THEMES:
labels[lab["id"]] = {"stance": lab["stance"], "theme": lab["theme"]}
return labels, calls
def aggregate(trend: dict, items: list[dict], labels: dict) -> dict:
stats = {"hot_list_ranks": trend["hot_list_ranks"], "platforms": {}}
for p in ("weibo", "zhihu", "bilibili"):
rows = [labels[i["id"]] | {"likes": i["likes"]} for i in items if i["platform"] == p and i["id"] in labels]
on = [r for r in rows if r["stance"] != "off_topic"]
st = Counter(r["stance"] for r in on)
stats["platforms"][p] = {
"items_labeled": len(rows), "on_topic": len(on),
"positive_share": round(st["positive"] / len(on), 2) if on else None,
"negative_share": round(st["negative"] / len(on), 2) if on else None,
"top_themes": Counter(r["theme"] for r in on).most_common(3),
"likes_on_negative_share": round(sum(r["likes"] for r in on if r["stance"] == "negative")
/ max(1, sum(r["likes"] for r in on)), 2)}
stats["unlabeled_items"] = len(items) - len(labels)
return stats
def unknown_numbers(text: str, stats: dict, names: str) -> list[str]:
"""Numbers in the brief not in the stats JSON or the topic names (percent forms of shares allowed)."""
known = set(re.findall(r"\d+", names))
for v in re.findall(r"\d+(?:\.\d+)?", json.dumps(stats)):
known |= {v, v.rstrip("0").rstrip(".") if "." in v else v}
if float(v) < 1:
known |= {str(round(float(v) * 100))}
found = [n.replace(",", "") for n in re.findall(r"\d[\d,]*(?:\.\d+)?", text)]
return sorted({n for n in found if n not in known})
def cost(model: str, calls: list[dict]) -> float:
pin, pout = PRICES[model]
return round(sum(c["usage"].get("prompt_tokens", 0) * pin + c["usage"].get("completion_tokens", 0) * pout
for c in calls) / 1e6, 5)
if __name__ == "__main__":
keyword, aliases, model = sys.argv[1], sys.argv[2].split(","), sys.argv[3]
if len(sys.argv) > 4:
saved = json.load(open(sys.argv[4], encoding="utf-8"))
trend, items = saved["trend"], saved["items"]
else:
trend, items = collect(keyword, aliases)
json.dump({"trend": trend, "items": items}, open("items.json", "w", encoding="utf-8"), ensure_ascii=False)
labels, calls = classify(model, keyword, items)
stats = aggregate(trend, items, labels)
text, usage, secs, rid = chat(model, (
f"Write a 120-180 word public opinion brief in English about '{keyword}' for a product team: where it trends, "
"overall stance per platform, main themes, one risk to watch. Shares are fractions of on-topic items. "
"Use only numbers in this JSON; do not compute new ones.\n" + json.dumps(stats)), 1500)
calls.append({"step": "summary", "id": rid, "usage": usage, "seconds": round(secs, 1)})
flags = unknown_numbers(text, stats, keyword + " " + " ".join(aliases))
with open("labels.csv", "w", newline="", encoding="utf-8") as f:
w = csv.writer(f)
w.writerow(["id", "platform", "likes", "stance", "theme"])
for i in items:
lab = labels.get(i["id"], {})
w.writerow([i["id"], i["platform"], i["likes"], lab.get("stance", ""), lab.get("theme", "")])
with open("brief.md", "w", encoding="utf-8") as f:
f.write(f"# Opinion brief: {keyword} ({model})\n\n```json\n{json.dumps(stats, indent=2)}\n```\n\n{text}\n\n"
+ (f"Check these numbers: {', '.join(flags)}\n" if flags else "Numbers check: OK\n"))
run = {"model": model, "items": len(items), "llm_calls": calls, "llm_seconds": round(sum(c["seconds"] for c in calls), 1),
"est_cost_usd": cost(model, calls), "unverified_numbers": flags}
json.dump(run, open("run.json", "w"), indent=1)
print(json.dumps(stats, indent=2))
print(text)
print(json.dumps({k: v for k, v in run.items() if k != "llm_calls"}))
改代码前先了解这几处设计:
- 每批 25 条。 78 条拆成 4 次分类调用,单次 prompt 最多约 2,500 token。返回的 JSON 用正则取出,不在允许范围内的标签直接丢弃,计入
unlabeled_items。我所有运行里这个数都是 0。 - 热榜位置是列表位置。 微博热搜列表里可能夹着没有排名的推广条目,所以位置不一定等于官方排名。测试期间微博热搜一直没命中别名,简报里这一项是空列表。
max_tokens是 Chat Completions 文档里的字段。 Luna 单次分类调用最多花了 1,386 个推理 token,所以分类预算给到 4,000。est_cost_usd按标价算。 用量乘以每百万输入/输出 token 的价格:Luna 0.10/0.50 美元,Sol 2/10 美元。实际扣费可能不同,见下文成本一节。
想自己跑,先看 bilibili/web/video-comments 接口文档和 Chat Completions 指南,然后获取 SandBase API Key。
怎么测便宜模型
选 Luna 的依据来自我们的大模型分档实测:GPT-6 Luna 在短工具任务上 51/51 全对,单任务成本约为 GPT-6.1 Sol 的 1/20。但给玩梗评论打标签是另一回事,得单独测。
- 先用 Luna 跑一遍程序,抓到 61 条(微博 18、知乎 20、B站 23),把这份
items.json固定下来。 - 取微博前 16 条、知乎前 17 条、B站前 17 条,共 50 条,在看任何模型输出之前,我自己标好立场和主题。标注人只有我一个,所以这是参照,不是标准答案。其中 11 条我标成了「有歧义」(褒贬参半、反讽、圈内梗)。
- 两个模型在这份固定数据上各跑 3 次,用下面的脚本打分。
- 之后每个模型各做一次完整的端到端运行。两次都重新抓数据,拿到的是同样的 78 条、文本完全一致,所以打标签的提示词也一致。开头引用的两份简报就来自这里。
我的 50 条标注:正面 24、中性 12、负面 7、无关 7。主题第一是芯片性能(17),其次 other(12)、相机(7)、价格与性价比(6)。打分脚本算准确率和 Cohen’s kappa,后者会扣掉碰巧一致的部分:
#!/usr/bin/env python3
"""Compare a run's labels.csv with hand labels: accuracy and Cohen's kappa for stance.
Usage: python3 score_labels.py hand_labels.json labels.csv [labels.csv ...]
hand_labels.json maps item id -> {"stance": ..., "theme": ...}.
"""
import csv
import json
import sys
from collections import Counter
def kappa(pairs: list[tuple[str, str]]) -> float:
n = len(pairs)
observed = sum(a == b for a, b in pairs) / n
ca, cb = Counter(a for a, _ in pairs), Counter(b for _, b in pairs)
expected = sum(ca[k] * cb[k] for k in ca) / (n * n)
return round((observed - expected) / (1 - expected), 3) if expected < 1 else 1.0
gold = json.load(open(sys.argv[1], encoding="utf-8"))
for path in sys.argv[2:]:
pred = {r["id"]: r for r in csv.DictReader(open(path, encoding="utf-8"))}
ids = [i for i in gold if i in pred]
stance = [(gold[i]["stance"], pred[i]["stance"] or "missing") for i in ids]
theme = [(gold[i]["theme"], pred[i]["theme"] or "missing") for i in ids]
print(json.dumps({"file": path, "scored": len(ids),
"stance_acc": round(sum(a == b for a, b in stance) / len(ids), 3),
"stance_kappa": kappa(stance),
"theme_acc": round(sum(a == b for a, b in theme) / len(ids), 3),
"both_acc": round(sum(gold[i]["stance"] == pred[i]["stance"] and
gold[i]["theme"] == pred[i]["theme"] for i in ids) / len(ids), 3)}))
结果:GPT-6 Luna 对 GPT-6.1 Sol
同一份 61 条数据,每个模型 3 次,在我标注的 50 条上打分:
| 指标(3 次运行的范围) | GPT-6 Luna | GPT-6.1 Sol |
|---|---|---|
| 立场准确率 | 0.70 到 0.80 | 0.78 到 0.84 |
| 立场 kappa | 0.556 到 0.704 | 0.670 到 0.755 |
| 主题准确率 | 0.80 到 0.82 | 0.86 到 0.90 |
| 立场和主题都对 | 0.62 到 0.74 | 0.70 到 0.78 |
| 立场准确率,39 条无歧义 | 0.769 到 0.821 | 0.846 到 0.897 |
| 立场准确率,11 条有歧义 | 0.455 到 0.818 | 0.545 到 0.636 |
| 运行之间立场一致率,全部 61 条 | 0.852 到 0.902 | 0.934 到 0.951 |
| 每份简报大模型耗时(秒) | 45.6 到 49.8 | 39.9 到 51.5 |
| 每份简报标价成本(美元) | 0.00257 到 0.00273 | 0.02962 到 0.03268 |
除了 11 条有歧义的内容,Sol 每一行都更好;在有歧义的那几条上,Luna 三次从 0.455 跳到 0.818,Sol 稳定在 0.545 到 0.636。整体差距不大:50 条里平均多对两三条。更大的差别在稳定性。Luna 三次运行之间,61 条里有 6 到 9 条立场互相对不上;Sol 只有 3 到 4 条。Luna 有一次立场准确率掉到 0.70,多出来的错大部分在有歧义的那几条上。
两个模型最常见的错一样。每次都有 4 条我标为无关的内容被标成中性:只有一个表情的回复、讨论话题热度本身的评论、和产品不相干的玩笑。这更像分类体系的问题,可以给「无关」写更明确的定义。
产品团队真正看的是汇总数字。在这 50 条上,按我的标注,B站正面占比 0.42、负面 0.25。Sol 三次给出正面 0.38 到 0.43、负面 0.21 到 0.25;Luna 给出正面 0.33 到 0.45、负面 0.08 到 0.27。Sol 也有自己的偏差:三次运行都没找出一条负面的知乎回答,而我在 15 条话题内回答里数到 2 条。
Luna 栽在哪
端到端运行把 B站的问题摆得很清楚。两个模型在 B站有 14 条分歧,其中 4 条 Luna 判负面、Sol 判正面,另有 4 条 Luna 判无关、Sol 判正面。好几条是用竞品芯片口吻写的「凡尔赛」式自夸,B站用户这么写其实是在夸新芯片,Luna 按字面意思读了。也有一条把这颗芯片贬为老款骁龙翻版的评论,Luna 判负面、Sol 判正面,这条我站 Luna。两个模型对反讽都不可靠,只是 Luna 漏得更多;在反讽成风的平台上,这足以改变头条数字。

截图:GPT-6 Luna 模型页(已滚动到价格一栏)标明每百万输入 token 0.10 美元、每百万输出 token 0.50 美元,即成本一列的标价(2026-10-03 截取)。
一份简报多少钱
2026-10-03,7 个数据端点通过 GET /v1/models/<model> 返回的都是 base_price: "0",B站评论的模型页显示 Free,顶部挂着「API Free Week」横幅。这是促销期间某一天的标价,不是承诺。大模型这边,model_card.price_formula 写的是:prompt 少于 272K token 时,Luna 每百万输入/输出 0.10/0.50 美元,Sol 2/10 美元。

截图:bilibili/web/video-comments 模型页显示基础价格 Free、同步执行、2 个输入字段(bv_id 和 pn),页面顶部是 API Free Week 横幅(2026-10-03 截取)。
每份简报的实测成本,按用量和标价算的 est_cost_usd,对比 GET /v1/tasks/<id>/cost 返回的实际扣费:
| 运行 | Luna 标价 | Luna 实扣 | Sol 标价 | Sol 实扣 |
|---|---|---|---|---|
| 端到端,78 条,全新 prompt | $0.00331 | $0.003457 | $0.03749 | $0.040513 |
| 固定 61 条上的重复运行(各一次) | $0.00257 | $0.00214 | $0.02962 | $0.020614 |
差额在成本记录里看得到。全新 prompt 带有 cache_creation_tokens,实扣比标价高 4% 到 8%;完全相同的 prompt 重跑时带有 cached_tokens,实扣比标价低 17%(Luna)到 30%(Sol)。代码没算缓存,est_cost_usd 只是估算。能对上账的只有这四次运行:同样的 id,运行后头几分钟能查到,之后再查就返回「task not found」,所以其他运行只有标价估算。
GET /v1/models/<id> 和 GET /v1/tasks/<id>/cost 用的是同一把 Bearer Key。没有 Key 的话,公开的模型页上也能看到同样的标价。
Luna 单次输出比 Sol 长,因为推理 token 花得多(端到端运行的三个满批次里,Luna 每次 1,125 到 1,897 个 completion token,Sol 是 618 到 717)。所以一份 Luna 简报大约是 Sol 的 1/11 到 1/12,而不是价目表上的 1/20。
怎么选模型
| 场景 | 选 |
|---|---|
| 每天看微博帖子和知乎回答 | GPT-6 Luna |
| B站评论,或任何玩梗多的评论区 | GPT-6.1 Sol,或 Luna 跑两遍、分歧条目交给 Sol |
| 头条数字要报给管理层 | GPT-6.1 Sol,再人工看一遍负面 |
| 每天几千条 | Luna,每个新话题先人工抽查 50 条 |
省钱又求稳的做法:Luna 跑两遍,只把两次不一致的条目交给 Sol。按我的数据,61 条里只有 6 到 9 条。
局限
测试范围:一个话题、一天,打分对比用 61 条,端到端用 78 条,每个模型 3 次,标注人一个。没测其他话题、中文简报,也没评估简报质量,只做了数字核对(最终运行全部通过)。热榜和搜索结果每分钟都在变,B站热搜位置一小时内从 21 位挪到 25 位。简报描述的是那一刻端点返回的内容,不代表全网舆论。
常见问题
GPT-6 Luna 做舆情监测够用吗?
在这次测试里,微博帖子和知乎回答够用:立场和主题比 Sol 低几个点,成本约 1/11 到 1/12。B站评论上它不够稳定,反讽漏判多到会改变头条占比。
为什么不让大模型直接算占比?
数数和算平均是模型容易出错的地方。所有占比和计数都由 Python 算,最终运行里简报的数字核对一次都没报警。
为什么知乎回答和 B站评论只读一页?
我实测知乎下一页返回为空,B站第 2 页有时什么都不回。所以每个来源只读一页,要更多数据先验证翻页。
抓到的评论能直接发吗?
不能照搬原文,那是个人写的。只发汇总和泛化转述,items.json 留在内部。
怎么查一次运行花了多少钱?
拿到 chat completion 返回的 id,带上 Key 调 GET /v1/tasks/<id>/cost。最好运行完就查:我实测同样的 id,几分钟内能查到成本记录,大约半小时后再查就返回「task not found」,查到就存下来。