Blog/教程/

舆情监测 Agent:微博、知乎、B站,便宜模型够不够用

用 Python 写一个舆情监测 Agent:微博、知乎、B站三平台抓同一话题,大模型打立场和主题标签,代码做统计。再用 50 条人工标注对比 GPT-6 Luna 和 GPT-6.1 Sol 的准确率、成本与耗时。

舆情监测 Agent 教程封面,覆盖微博、知乎和 B站,使用低价大模型

同一批 78 条华为 Mate 90 相关的帖子、回答和评论,我分别交给两个模型打标签。GPT-6.1 Sol 说 B站在话题内的评论里 45% 是正面的;GPT-6 Luna 说只有 17%,还把 B站写成了「主要风险区」。内容、打标签的提示词、代码完全一样,两份简报对 B站的判断却差得很远。微博加知乎的 38 条,两个模型有 33 条立场一致;B站的 40 条,只有 26 条一致。

「便宜模型够不够用」,答案就藏在这个差距里。本文搭一个舆情监测 Agent(微博、知乎、B站):针对一个消费科技话题,先看它在三个平台热榜上的位置,再拉公开的微博帖子、知乎回答和 B站评论,用大模型给每条内容打立场和主题,在 Python 里汇总,最后生成一页简报。再用我手工标注的 50 条对比 GPT-6 Luna 和 GPT-6.1 Sol。适合想每天看新品口碑、又不想每条标签都付旗舰价的产品、公关和研究团队。

先说结论

  • 一把 SandBase API Key 覆盖 7 个公开数据端点(三个热榜、微博实时搜索、知乎问题回答、B站搜索和评论)以及两个大模型。2026-10-03 这些数据端点在目录里标为 Free,当时页面挂着「API Free Week」横幅。
  • 50 条人工标注上,GPT-6 Luna 三次运行的立场准确率是 70% 到 80%(Cohen’s kappa 0.56 到 0.70);GPT-6.1 Sol 是 78% 到 84%(kappa 0.67 到 0.76)。主题准确率分别是 80% 到 82% 和 86% 到 90%。
  • 一份简报 Luna 约 0.003 美元,Sol 约 0.03 到 0.04 美元,约 1/11 到 1/12。耗时差不多,每份简报大模型时间约 40 到 56 秒。
  • 微博和知乎用 Luna 足够。B站评论里反讽和玩梗多,Luna 把夸奖读成批评的次数多到足以改变结论。要么跑两遍,要么 B站交给 Sol。

这个 Agent 做什么

话题选的是华为 Mate 90 系列,测试前几天刚发布。2026-10-03 当天,知乎上一个关于该机芯片解析视频的问题排在热榜第 12 到 13 位,B站热搜里对应关键词排在第 21 到 25 位。选它是因为它属于消费电子,多个平台都有热度,且不涉及时政或敏感议题。代码里搜索词用 华为Mate90,热榜标题匹配用别名 Mate 90 和 Mate90。

步骤端点代码保留什么
看热度weibo/web-v2/hot-search、zhihu/web/hot-list、bilibili/web/hot-search标题命中别名的条目位置;知乎问题 id
微博帖子weibo/web-v2/realtime-search2 页,帖子 id、正文、点赞数
知乎回答zhihu/web/question-answers热榜问题下 20 条回答的 id、正文、赞同数
B站视频bilibili/web/general-search播放量最高的 2 个相关视频
B站评论bilibili/web/video-comments每个视频第 1 页评论的 id、正文、点赞数
打标签、写简报openai/gpt-6-luna 或 openai/gpt-6.1-sol,走 /v1/chat/completions每条的立场和主题;120 到 180 词简报

立场四选一:positive、negative、neutral(陈述事实、提问或褒贬参半)、off_topic。主题八选一:chip_performance、camera、price_value、design_build、battery_charging、software_system、availability_sales、other。所有数字都由 Python 算:话题内条数、正负面占比、前三主题、负面内容拿到的点赞占比。大模型只根据这份 JSON 写简报,代码再检查简报里有没有 JSON 之外的数字。

它和我们之前的微博加抖音舆情监测 Agent 不同:那篇按关键词长期追踪,这篇是单话题三平台简报,外加模型选型实测。想在单个平台挖深,可以看知乎问答挖掘和 B站视频评论分析。

数据边界

只读公开数据,需要 SandBase API Key。SandBase 不是微博、知乎或 B站的官方合作方。不碰私有账号、私信、账号后台数据,也不做任何账号操作。帖子和评论是个人写的,所以本文只放汇总结果和泛化转述,不贴原文、不出现用户名。程序生成的 labels.csv 只有 id、点赞数和标签;items.json 里有模型看过的原文,请只留在内部。

需求用什么
看某个产品或发布会的公开讨论SandBase 公开数据端点(本文)
自己微博账号的帖子和粉丝数据微博开放平台
自己 B站频道的数据B站开放平台

实测记录:2026-10-03(UTC)

每个请求都是 POST https://api.sandbase.ai/v1/api/<vendor>/<path>,带 Authorization: Bearer $SANDBASE_API_KEY,body 只放该端点自己的字段。参考文档把数据放在 outputs[0].data,但有些端点被观察到放在别处,所以读取函数优先取 outputs[0].data,取不到再退到 outputs[0],最后才看顶层 output;状态不是 completed 就直接报错。当天我记录的每一次调用,这 7 个端点都返回在 outputs[0].data。下面提到的字段名都来自这些调用,属于实测观察,不是文档保证,所以代码一律用 .get() 读。

curl -s https://api.sandbase.ai/v1/api/bilibili/web/hot-search \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"limit": 30}'

运行 a8eec5b3-ae9f-46a7-8500-94c6e8415aaa(12:44 UTC)返回 {"id", "status": "completed", "model", "outputs"},outputs[0].data 下面的列表在 data.trending.list。命中的那一条,完整如下:

{"goto": "", "heat_score": 438461, "icon": "", "keyword": "华为Mate90系列韬定律解析", "show_name": "华为Mate90系列韬定律解析", "uri": ""}

zhihu/web/hot-list 传 {"limit": "50"}(运行 eb994e55-fe36-42a4-959e-d80e80796f4e),条目在 data 下,每条有个 target。命中问题的 target 一共 12 个键,这里只展示 4 个(author、excerpt 等其余键省略):

{"id": 2089412797440455404, "answer_count": 178, "follower_count": 447, "created": 1790934830}

Agent 把这个 id 传给 question-answers:{"question_id": "2089412797440455404", "limit": 20}。运行 30eb337a-c6e2-41e6-92a8-5de83df30245 返回 20 条,每条 target 里有 content、voteup_count、comment_count,另有一个 paging 对象,is_end: false,带一个 next 链接。我从链接里解析出 cursor、offset、session_id 再请求一次(运行 9ba6fdcf-3b01-4188-9379-22c3caa50589),状态 completed,但 0 条回答。所以程序只读第 1 页。

weibo/web-v2/realtime-search 传 {"query": "华为Mate90", "page": 1},结果在 data.parsed_data.results,我的调用里每页 9 到 10 条,每条有 weibo_id、content,以及包含 comment_count、like_count、repost_count 的 interaction。抓到的微博全都是 0 赞(发出才几分钟),所以「负面点赞占比」这一项实际上只反映知乎和 B站。第 2、3 页能拿到新的帖子 id。

最不稳的是 bilibili/web/video-comments,参数 {"bv_id": "<视频 id>", "pn": 1}。对播放量最高的那个视频,有一次调用(运行 93cbc90a-332c-4b97-8206-3b3d4889e926)只回了 3 条评论,data.page 却写着:

{"acount": 26460, "count": 26460, "num": 1, "size": 20}

之后再请求 pn: 1 回了 20 条,pn: 2 回 0 条且 page 各字段全是 0,pn: 3 又回 20 条。所以程序对第 1 页最多重试三次,直到拿到至少 10 条,不再往后翻。

SandBase weibo/web-v2/realtime-search 端点参考页,显示 POST 路由、query 和 page 字段以及 outputs[0].data 响应信封

截图:微博实时搜索参考页写明了 POST 路由、必填的 query 和可选的 page,响应信封里 outputs[0].data 的示例是空的,所以上文的帖子字段都来自实测调用(2026-10-03 截取)。

SandBase zhihu/web/question-answers 端点参考页,列出 cursor、limit、offset、order、question_id 和 session_id

截图:知乎问题回答参考页列出了用于翻页的 cursor、offset 和 session_id;我实测第二页返回为空,所以 Agent 只读第一页 20 条回答(2026-10-03 截取)。

完整程序

一个文件,标准库加 requests。第一次运行会抓数据并存成 items.json;把这个文件作为第四个参数传进去,就能让另一个模型给完全相同的内容打标签,我的对比就是这么做的。

export SANDBASE_API_KEY=...   # 在 shell 里设置,别写进文件
python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6-luna
python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6.1-sol items.json
#!/usr/bin/env python3
"""Three-platform public opinion brief (Weibo, Zhihu, Bilibili) for one consumer/tech topic.

Usage: python3 opinion_brief.py <search keyword> <aliases, comma-separated> <model> [items.json]
Example: python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6-luna
Pass a saved items.json as the 4th argument to skip fetching and reuse the same items.
Writes items.json (with text, keep private), labels.csv (ids + labels only) and brief.md.
"""
import csv
import html
import json
import os
import re
import sys
import time
from collections import Counter

import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
PRICES = {"openai/gpt-6-luna": (0.1, 0.5), "openai/gpt-6.1-sol": (2.0, 10.0)}  # USD per 1M tokens, 2026-10-03
STANCES = ["positive", "negative", "neutral", "off_topic"]
THEMES = ["chip_performance", "camera", "price_value", "design_build", "battery_charging",
          "software_system", "availability_sales", "other"]


def call(model: str, params: dict, retries: int = 2) -> dict:
    """POST /v1/api/<model>. Prefer outputs[0].data, fall back to outputs[0], then output."""
    for attempt in range(retries + 1):
        try:
            resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
            body = resp.json()
        except requests.RequestException as err:
            if attempt == retries:
                raise RuntimeError(f"{model}: {err}") from err
            time.sleep(3 * (attempt + 1))
            continue
        if resp.status_code >= 500 and attempt < retries:
            time.sleep(3 * (attempt + 1))
            continue
        if resp.status_code != 200 or body.get("status") != "completed":
            raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
        outputs = body.get("outputs") or []
        if outputs and isinstance(outputs[0], dict):
            return outputs[0].get("data", outputs[0]) or {}
        if body.get("output") is not None:
            return body["output"]
        raise RuntimeError(f"{model}: completed but no payload")
    raise RuntimeError(f"{model}: retries exhausted")


def matches(text: str, aliases: list[str]) -> bool:
    t = (text or "").lower().replace(" ", "")
    return any(a.lower().replace(" ", "") in t for a in aliases)


def clean(text: str, limit: int = 280) -> str:
    text = re.sub(r"<[^>]+>", " ", html.unescape(text or ""))
    return re.sub(r"\s+", " ", text).strip()[:limit]


def trending(aliases: list[str]) -> dict:
    """Where the topic sits on each platform's hot list right now (rank 1 = top)."""
    wb = call("weibo/web-v2/hot-search", {}).get("realtime") or []
    zh = (call("zhihu/web/hot-list", {"limit": "50"}).get("data") or [])
    bl = ((call("bilibili/web/hot-search", {"limit": 30}).get("data") or {}).get("trending") or {}).get("list") or []
    hits = {"weibo": [i + 1 for i, r in enumerate(wb) if matches(r.get("word"), aliases)],
            "zhihu": [i + 1 for i, r in enumerate(zh) if matches((r.get("target") or {}).get("title"), aliases)],
            "bilibili": [i + 1 for i, r in enumerate(bl) if matches(r.get("keyword"), aliases)]}
    zhihu_qid = next((str(r["target"]["id"]) for r in zh
                      if matches((r.get("target") or {}).get("title"), aliases)), None)
    return {"hot_list_ranks": hits, "zhihu_question_id": zhihu_qid}


def collect(keyword: str, aliases: list[str]) -> tuple[dict, list[dict]]:
    trend, items = trending(aliases), []
    for page in (1, 2):
        data = call("weibo/web-v2/realtime-search", {"query": keyword, "page": page})
        for r in (data.get("parsed_data") or {}).get("results") or []:
            inter = r.get("interaction") or {}
            items.append({"id": f"wb:{r.get('weibo_id')}", "platform": "weibo", "text": clean(r.get("content")),
                          "likes": int(inter.get("like_count") or 0)})
    if trend["zhihu_question_id"]:
        data = call("zhihu/web/question-answers", {"question_id": trend["zhihu_question_id"], "limit": 20})
        for a in data.get("data") or []:
            t = a.get("target") or {}
            items.append({"id": f"zh:{t.get('id')}", "platform": "zhihu",
                          "text": clean(t.get("content") or t.get("excerpt")), "likes": int(t.get("voteup_count") or 0)})
    data = call("bilibili/web/general-search", {"keyword": keyword, "order": "totalrank", "page": 1, "page_size": 20})
    videos = [v for v in (data.get("data") or {}).get("result") or [] if matches(clean(v.get("title")), aliases)]
    for v in sorted(videos, key=lambda v: int(v.get("play") or 0), reverse=True)[:2]:
        for _ in range(3):  # page 1 sometimes comes back short or empty; retry it
            replies = (call("bilibili/web/video-comments", {"bv_id": v["bvid"], "pn": 1}).get("data") or {}).get("replies") or []
            if len(replies) >= 10:
                break
            time.sleep(2)
        for c in replies:
            items.append({"id": f"bl:{c.get('rpid')}", "platform": "bilibili",
                          "text": clean((c.get("content") or {}).get("message")), "likes": int(c.get("like") or 0)})
    items = [i for i in {i["id"]: i for i in items}.values() if i["text"]]  # dedupe, drop empty
    return trend, items


def chat(model: str, prompt: str, max_tokens: int) -> tuple[str, dict, float, str]:
    t0 = time.time()
    resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=300,
                         json={"model": model, "max_tokens": max_tokens,
                               "messages": [{"role": "user", "content": prompt}]})
    resp.raise_for_status()
    body = resp.json()
    return body["choices"][0]["message"]["content"] or "", body.get("usage") or {}, time.time() - t0, body.get("id")


def classify(model: str, topic: str, items: list[dict], batch: int = 25) -> tuple[dict, list[dict]]:
    labels, calls = {}, []
    for start in range(0, len(items), batch):
        chunk = [{"id": i["id"], "text": i["text"]} for i in items[start:start + batch]]
        prompt = (f"Topic: {topic}. Label each post for public opinion monitoring.\n"
                  f"stance toward the topic product: one of {STANCES} (neutral = factual, question or mixed; "
                  f"off_topic = not about the product).\ntheme: the main aspect, one of {THEMES}.\n"
                  'Reply with JSON only: {"labels": [{"id": "...", "stance": "...", "theme": "..."}]}\n'
                  + json.dumps(chunk, ensure_ascii=False))
        text, usage, secs, rid = chat(model, prompt, 4000)
        calls.append({"step": "classify", "id": rid, "usage": usage, "seconds": round(secs, 1)})
        found = re.search(r"\{.*\}", text, re.S)
        for lab in (json.loads(found.group(0)).get("labels") if found else []) or []:
            if lab.get("stance") in STANCES and lab.get("theme") in THEMES:
                labels[lab["id"]] = {"stance": lab["stance"], "theme": lab["theme"]}
    return labels, calls


def aggregate(trend: dict, items: list[dict], labels: dict) -> dict:
    stats = {"hot_list_ranks": trend["hot_list_ranks"], "platforms": {}}
    for p in ("weibo", "zhihu", "bilibili"):
        rows = [labels[i["id"]] | {"likes": i["likes"]} for i in items if i["platform"] == p and i["id"] in labels]
        on = [r for r in rows if r["stance"] != "off_topic"]
        st = Counter(r["stance"] for r in on)
        stats["platforms"][p] = {
            "items_labeled": len(rows), "on_topic": len(on),
            "positive_share": round(st["positive"] / len(on), 2) if on else None,
            "negative_share": round(st["negative"] / len(on), 2) if on else None,
            "top_themes": Counter(r["theme"] for r in on).most_common(3),
            "likes_on_negative_share": round(sum(r["likes"] for r in on if r["stance"] == "negative")
                                             / max(1, sum(r["likes"] for r in on)), 2)}
    stats["unlabeled_items"] = len(items) - len(labels)
    return stats


def unknown_numbers(text: str, stats: dict, names: str) -> list[str]:
    """Numbers in the brief not in the stats JSON or the topic names (percent forms of shares allowed)."""
    known = set(re.findall(r"\d+", names))
    for v in re.findall(r"\d+(?:\.\d+)?", json.dumps(stats)):
        known |= {v, v.rstrip("0").rstrip(".") if "." in v else v}
        if float(v) < 1:
            known |= {str(round(float(v) * 100))}
    found = [n.replace(",", "") for n in re.findall(r"\d[\d,]*(?:\.\d+)?", text)]
    return sorted({n for n in found if n not in known})


def cost(model: str, calls: list[dict]) -> float:
    pin, pout = PRICES[model]
    return round(sum(c["usage"].get("prompt_tokens", 0) * pin + c["usage"].get("completion_tokens", 0) * pout
                     for c in calls) / 1e6, 5)


if __name__ == "__main__":
    keyword, aliases, model = sys.argv[1], sys.argv[2].split(","), sys.argv[3]
    if len(sys.argv) > 4:
        saved = json.load(open(sys.argv[4], encoding="utf-8"))
        trend, items = saved["trend"], saved["items"]
    else:
        trend, items = collect(keyword, aliases)
        json.dump({"trend": trend, "items": items}, open("items.json", "w", encoding="utf-8"), ensure_ascii=False)
    labels, calls = classify(model, keyword, items)
    stats = aggregate(trend, items, labels)
    text, usage, secs, rid = chat(model, (
        f"Write a 120-180 word public opinion brief in English about '{keyword}' for a product team: where it trends, "
        "overall stance per platform, main themes, one risk to watch. Shares are fractions of on-topic items. "
        "Use only numbers in this JSON; do not compute new ones.\n" + json.dumps(stats)), 1500)
    calls.append({"step": "summary", "id": rid, "usage": usage, "seconds": round(secs, 1)})
    flags = unknown_numbers(text, stats, keyword + " " + " ".join(aliases))
    with open("labels.csv", "w", newline="", encoding="utf-8") as f:
        w = csv.writer(f)
        w.writerow(["id", "platform", "likes", "stance", "theme"])
        for i in items:
            lab = labels.get(i["id"], {})
            w.writerow([i["id"], i["platform"], i["likes"], lab.get("stance", ""), lab.get("theme", "")])
    with open("brief.md", "w", encoding="utf-8") as f:
        f.write(f"# Opinion brief: {keyword} ({model})\n\n```json\n{json.dumps(stats, indent=2)}\n```\n\n{text}\n\n"
                + (f"Check these numbers: {', '.join(flags)}\n" if flags else "Numbers check: OK\n"))
    run = {"model": model, "items": len(items), "llm_calls": calls, "llm_seconds": round(sum(c["seconds"] for c in calls), 1),
           "est_cost_usd": cost(model, calls), "unverified_numbers": flags}
    json.dump(run, open("run.json", "w"), indent=1)
    print(json.dumps(stats, indent=2))
    print(text)
    print(json.dumps({k: v for k, v in run.items() if k != "llm_calls"}))

改代码前先了解这几处设计:

  • 每批 25 条。 78 条拆成 4 次分类调用,单次 prompt 最多约 2,500 token。返回的 JSON 用正则取出,不在允许范围内的标签直接丢弃,计入 unlabeled_items。我所有运行里这个数都是 0。
  • 热榜位置是列表位置。 微博热搜列表里可能夹着没有排名的推广条目,所以位置不一定等于官方排名。测试期间微博热搜一直没命中别名,简报里这一项是空列表。
  • max_tokens 是 Chat Completions 文档里的字段。 Luna 单次分类调用最多花了 1,386 个推理 token,所以分类预算给到 4,000。
  • est_cost_usd 按标价算。 用量乘以每百万输入/输出 token 的价格:Luna 0.10/0.50 美元,Sol 2/10 美元。实际扣费可能不同,见下文成本一节。

想自己跑,先看 bilibili/web/video-comments 接口文档和 Chat Completions 指南,然后获取 SandBase API Key。

怎么测便宜模型

选 Luna 的依据来自我们的大模型分档实测:GPT-6 Luna 在短工具任务上 51/51 全对,单任务成本约为 GPT-6.1 Sol 的 1/20。但给玩梗评论打标签是另一回事,得单独测。

  1. 先用 Luna 跑一遍程序,抓到 61 条(微博 18、知乎 20、B站 23),把这份 items.json 固定下来。
  2. 取微博前 16 条、知乎前 17 条、B站前 17 条,共 50 条,在看任何模型输出之前,我自己标好立场和主题。标注人只有我一个,所以这是参照,不是标准答案。其中 11 条我标成了「有歧义」(褒贬参半、反讽、圈内梗)。
  3. 两个模型在这份固定数据上各跑 3 次,用下面的脚本打分。
  4. 之后每个模型各做一次完整的端到端运行。两次都重新抓数据,拿到的是同样的 78 条、文本完全一致,所以打标签的提示词也一致。开头引用的两份简报就来自这里。

我的 50 条标注:正面 24、中性 12、负面 7、无关 7。主题第一是芯片性能(17),其次 other(12)、相机(7)、价格与性价比(6)。打分脚本算准确率和 Cohen’s kappa,后者会扣掉碰巧一致的部分:

#!/usr/bin/env python3
"""Compare a run's labels.csv with hand labels: accuracy and Cohen's kappa for stance.

Usage: python3 score_labels.py hand_labels.json labels.csv [labels.csv ...]
hand_labels.json maps item id -> {"stance": ..., "theme": ...}.
"""
import csv
import json
import sys
from collections import Counter


def kappa(pairs: list[tuple[str, str]]) -> float:
    n = len(pairs)
    observed = sum(a == b for a, b in pairs) / n
    ca, cb = Counter(a for a, _ in pairs), Counter(b for _, b in pairs)
    expected = sum(ca[k] * cb[k] for k in ca) / (n * n)
    return round((observed - expected) / (1 - expected), 3) if expected < 1 else 1.0


gold = json.load(open(sys.argv[1], encoding="utf-8"))
for path in sys.argv[2:]:
    pred = {r["id"]: r for r in csv.DictReader(open(path, encoding="utf-8"))}
    ids = [i for i in gold if i in pred]
    stance = [(gold[i]["stance"], pred[i]["stance"] or "missing") for i in ids]
    theme = [(gold[i]["theme"], pred[i]["theme"] or "missing") for i in ids]
    print(json.dumps({"file": path, "scored": len(ids),
                      "stance_acc": round(sum(a == b for a, b in stance) / len(ids), 3),
                      "stance_kappa": kappa(stance),
                      "theme_acc": round(sum(a == b for a, b in theme) / len(ids), 3),
                      "both_acc": round(sum(gold[i]["stance"] == pred[i]["stance"] and
                                            gold[i]["theme"] == pred[i]["theme"] for i in ids) / len(ids), 3)}))

结果:GPT-6 Luna 对 GPT-6.1 Sol

同一份 61 条数据,每个模型 3 次,在我标注的 50 条上打分:

指标(3 次运行的范围)GPT-6 LunaGPT-6.1 Sol
立场准确率0.70 到 0.800.78 到 0.84
立场 kappa0.556 到 0.7040.670 到 0.755
主题准确率0.80 到 0.820.86 到 0.90
立场和主题都对0.62 到 0.740.70 到 0.78
立场准确率,39 条无歧义0.769 到 0.8210.846 到 0.897
立场准确率,11 条有歧义0.455 到 0.8180.545 到 0.636
运行之间立场一致率,全部 61 条0.852 到 0.9020.934 到 0.951
每份简报大模型耗时(秒)45.6 到 49.839.9 到 51.5
每份简报标价成本(美元)0.00257 到 0.002730.02962 到 0.03268

除了 11 条有歧义的内容,Sol 每一行都更好;在有歧义的那几条上,Luna 三次从 0.455 跳到 0.818,Sol 稳定在 0.545 到 0.636。整体差距不大:50 条里平均多对两三条。更大的差别在稳定性。Luna 三次运行之间,61 条里有 6 到 9 条立场互相对不上;Sol 只有 3 到 4 条。Luna 有一次立场准确率掉到 0.70,多出来的错大部分在有歧义的那几条上。

两个模型最常见的错一样。每次都有 4 条我标为无关的内容被标成中性:只有一个表情的回复、讨论话题热度本身的评论、和产品不相干的玩笑。这更像分类体系的问题,可以给「无关」写更明确的定义。

产品团队真正看的是汇总数字。在这 50 条上,按我的标注,B站正面占比 0.42、负面 0.25。Sol 三次给出正面 0.38 到 0.43、负面 0.21 到 0.25;Luna 给出正面 0.33 到 0.45、负面 0.08 到 0.27。Sol 也有自己的偏差:三次运行都没找出一条负面的知乎回答,而我在 15 条话题内回答里数到 2 条。

Luna 栽在哪

端到端运行把 B站的问题摆得很清楚。两个模型在 B站有 14 条分歧,其中 4 条 Luna 判负面、Sol 判正面,另有 4 条 Luna 判无关、Sol 判正面。好几条是用竞品芯片口吻写的「凡尔赛」式自夸,B站用户这么写其实是在夸新芯片,Luna 按字面意思读了。也有一条把这颗芯片贬为老款骁龙翻版的评论,Luna 判负面、Sol 判正面,这条我站 Luna。两个模型对反讽都不可靠,只是 Luna 漏得更多;在反讽成风的平台上,这足以改变头条数字。

SandBase openai/gpt-6-luna 模型页,显示每百万 token 输入 0.10 美元、输出 0.50 美元,最大输出 128K

截图:GPT-6 Luna 模型页(已滚动到价格一栏)标明每百万输入 token 0.10 美元、每百万输出 token 0.50 美元,即成本一列的标价(2026-10-03 截取)。

一份简报多少钱

2026-10-03,7 个数据端点通过 GET /v1/models/<model> 返回的都是 base_price: "0",B站评论的模型页显示 Free,顶部挂着「API Free Week」横幅。这是促销期间某一天的标价,不是承诺。大模型这边,model_card.price_formula 写的是:prompt 少于 272K token 时,Luna 每百万输入/输出 0.10/0.50 美元,Sol 2/10 美元。

SandBase bilibili/web/video-comments 模型页,显示 Free 基础价格、同步执行、2 个输入字段,顶部有 API Free Week 横幅

截图:bilibili/web/video-comments 模型页显示基础价格 Free、同步执行、2 个输入字段(bv_id 和 pn),页面顶部是 API Free Week 横幅(2026-10-03 截取)。

每份简报的实测成本,按用量和标价算的 est_cost_usd,对比 GET /v1/tasks/<id>/cost 返回的实际扣费:

运行Luna 标价Luna 实扣Sol 标价Sol 实扣
端到端,78 条,全新 prompt$0.00331$0.003457$0.03749$0.040513
固定 61 条上的重复运行(各一次)$0.00257$0.00214$0.02962$0.020614

差额在成本记录里看得到。全新 prompt 带有 cache_creation_tokens,实扣比标价高 4% 到 8%;完全相同的 prompt 重跑时带有 cached_tokens,实扣比标价低 17%(Luna)到 30%(Sol)。代码没算缓存,est_cost_usd 只是估算。能对上账的只有这四次运行:同样的 id,运行后头几分钟能查到,之后再查就返回「task not found」,所以其他运行只有标价估算。

GET /v1/models/<id> 和 GET /v1/tasks/<id>/cost 用的是同一把 Bearer Key。没有 Key 的话,公开的模型页上也能看到同样的标价。

Luna 单次输出比 Sol 长,因为推理 token 花得多(端到端运行的三个满批次里,Luna 每次 1,125 到 1,897 个 completion token,Sol 是 618 到 717)。所以一份 Luna 简报大约是 Sol 的 1/11 到 1/12,而不是价目表上的 1/20。

怎么选模型

场景选
每天看微博帖子和知乎回答GPT-6 Luna
B站评论,或任何玩梗多的评论区GPT-6.1 Sol,或 Luna 跑两遍、分歧条目交给 Sol
头条数字要报给管理层GPT-6.1 Sol,再人工看一遍负面
每天几千条Luna,每个新话题先人工抽查 50 条

省钱又求稳的做法:Luna 跑两遍,只把两次不一致的条目交给 Sol。按我的数据,61 条里只有 6 到 9 条。

局限

测试范围:一个话题、一天,打分对比用 61 条,端到端用 78 条,每个模型 3 次,标注人一个。没测其他话题、中文简报,也没评估简报质量,只做了数字核对(最终运行全部通过)。热榜和搜索结果每分钟都在变,B站热搜位置一小时内从 21 位挪到 25 位。简报描述的是那一刻端点返回的内容,不代表全网舆论。

常见问题

GPT-6 Luna 做舆情监测够用吗?

在这次测试里,微博帖子和知乎回答够用:立场和主题比 Sol 低几个点,成本约 1/11 到 1/12。B站评论上它不够稳定,反讽漏判多到会改变头条占比。

为什么不让大模型直接算占比?

数数和算平均是模型容易出错的地方。所有占比和计数都由 Python 算,最终运行里简报的数字核对一次都没报警。

为什么知乎回答和 B站评论只读一页?

我实测知乎下一页返回为空,B站第 2 页有时什么都不回。所以每个来源只读一页,要更多数据先验证翻页。

抓到的评论能直接发吗?

不能照搬原文,那是个人写的。只发汇总和泛化转述,items.json 留在内部。

怎么查一次运行花了多少钱?

拿到 chat completion 返回的 id,带上 Key 调 GET /v1/tasks/<id>/cost。最好运行完就查:我实测同样的 id,几分钟内能查到成本记录,大约半小时后再查就返回「task not found」,查到就存下来。