Blog/Tutorials/

Public Opinion Agent: Weibo, Zhihu, Bilibili on GPT-6 Luna

Build a public opinion monitoring agent for Weibo, Zhihu and Bilibili, then test GPT-6 Luna vs GPT-6.1 Sol on 50 hand-labeled posts for accuracy and cost.

Cover for a public opinion agent tutorial covering Weibo, Zhihu and Bilibili with a cheap LLM

I ran the same 78 posts and comments about the Huawei Mate 90 launch through two models. GPT-6.1 Sol said 45% of on-topic Bilibili comments were positive. GPT-6 Luna said 17% and called Bilibili “the main area of concern”. Same items, same labeling prompt, same code, two very different readings of Bilibili. On Weibo and Zhihu the two models agreed on 33 of 38 items. On Bilibili they agreed on 26 of 40.

That gap is the real answer to “is the cheap model good enough”. This tutorial builds a public opinion agent for one consumer tech topic across Weibo, Zhihu and Bilibili, which is what people usually want when they search for a public opinion monitoring agent Weibo Zhihu Bilibili setup. It finds where the topic trends, pulls public posts, answers and comments, labels each item’s stance and theme with an LLM, aggregates in Python and writes a one-page brief. Then it measures GPT-6 Luna against GPT-6.1 Sol on 50 items I labeled by hand. It’s for product, PR and research teams who want a daily read on a launch without paying flagship prices for every label.

Key takeaway

  • One SandBase API key covers seven public data endpoints (hot lists, Weibo search, Zhihu answers, Bilibili search and comments) and both LLMs. The data endpoints were listed as Free on 2026-10-03, during an “API Free Week” banner.
  • On 50 hand-labeled items, GPT-6 Luna matched my stance label 70 to 80% of the time over 3 runs (Cohen’s kappa 0.56 to 0.70). GPT-6.1 Sol matched 78 to 84% (kappa 0.67 to 0.76). Theme accuracy was 80 to 82% vs 86 to 90%.
  • A brief cost about $0.003 with Luna and about $0.03 to $0.04 with Sol, roughly 1/11 to 1/12. Latency was the same: about 40 to 56 seconds of LLM time per brief.
  • Luna is good enough for Weibo and Zhihu. On Bilibili’s ironic, meme-heavy comments it flipped praise to criticism often enough to change the headline. Run it twice, or send Bilibili to Sol.

What the agent does

The topic is the Huawei Mate 90 series, which launched days before my test. On 2026-10-03 a Zhihu question about a chip analysis video for the phone sat at position 12 or 13 of the Zhihu hot list, and the video’s keyword sat at position 21 to 25 on Bilibili’s trending searches. I picked it because it’s a consumer product, it was trending on more than one platform, and it isn’t a political or sensitive topic. The search keyword in the code is 华为Mate90 (“Huawei Mate90”), with aliases Mate 90 and Mate90 for matching hot-list titles.

StepEndpointWhat the code keeps
Where it trendsweibo/web-v2/hot-search, zhihu/web/hot-list, bilibili/web/hot-searchposition of any entry whose title matches an alias; the Zhihu question id
Weibo postsweibo/web-v2/realtime-search2 pages, post id, text, like count
Zhihu answerszhihu/web/question-answers20 answers to the trending question, id, text, upvotes
Bilibili videosbilibili/web/general-searchthe 2 most played matching videos
Bilibili commentsbilibili/web/video-commentspage 1 of comments per video, id, text, likes
Labels and briefopenai/gpt-6-luna or openai/gpt-6.1-sol via /v1/chat/completionsstance and theme per item; a 120 to 180 word brief

Stance is one of positive, negative, neutral (factual, a question, or mixed) and off_topic. Theme is one of chip_performance, camera, price_value, design_build, battery_charging, software_system, availability_sales and other. Python counts everything: on-topic items, positive and negative share, top three themes, and the share of likes that went to negative items. The LLM writes the brief from that JSON and a check flags any number in its text that isn’t in the stats.

This is distinct from our Weibo and Douyin social listening agent, which tracks keywords over time. This one is a single-topic brief across three platforms with a measured model choice. For deeper single-platform work, see Zhihu Q&A mining and Bilibili video comment analysis.

Data boundary

This reads public, read-only data. It needs a SandBase API key. SandBase isn’t an official partner of Weibo, Zhihu or Bilibili. There are no private accounts, DMs, owner analytics or account actions. Posts and comments are written by individuals, so the article shows only aggregates and generic paraphrases, never their text or usernames. The program’s labels.csv keeps ids, likes and labels only. Its items.json holds the text the model saw, so keep it private.

NeedUse
A read on public discussion of a product or launchSandBase public-data endpoints (this tutorial)
Your own Weibo account’s posts and fansWeibo open platform
Your own Bilibili channel’s dataBilibili open platform

Tested on 2026-10-03 (UTC)

Every request is POST https://api.sandbase.ai/v1/api/<vendor>/<path> with Authorization: Bearer $SANDBASE_API_KEY and a JSON body of that endpoint’s fields. The reference documents the payload at outputs[0].data, but some endpoints have been seen returning it elsewhere, so the reader prefers outputs[0].data, falls back to outputs[0] and then a top-level output, and raises if the status isn’t completed. All seven endpoints here returned outputs[0].data on every call I logged that day. The field names below come from those calls. They’re observed-only, not documented guarantees, which is why the code reads them with .get().

curl -s https://api.sandbase.ai/v1/api/bilibili/web/hot-search \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"limit": 30}'

Run a8eec5b3-ae9f-46a7-8500-94c6e8415aaa (12:44 UTC) returned {"id", "status": "completed", "model", "outputs"}. Under outputs[0].data, the list sat at data.trending.list. The matching entry, complete:

{"goto": "", "heat_score": 438461, "icon": "", "keyword": "华为Mate90系列韬定律解析", "show_name": "华为Mate90系列韬定律解析", "uri": ""}

The keyword reads “Huawei Mate90 series Tao’s Law chip analysis”. zhihu/web/hot-list with {"limit": "50"} (run eb994e55-fe36-42a4-959e-d80e80796f4e) put items under data. Each had a target; for the matching question I show 4 of its 12 keys (the others, including author and excerpt, are omitted):

{"id": 2089412797440455404, "answer_count": 178, "follower_count": 447, "created": 1790934830}

The agent sends that id to question-answers with {"question_id": "2089412797440455404", "limit": 20}. Run 30eb337a-c6e2-41e6-92a8-5de83df30245 returned 20 items with a target holding content, voteup_count and comment_count, plus a paging object with is_end: false and a next URL. I parsed cursor, offset and session_id from that URL and sent them back (run 9ba6fdcf-3b01-4188-9379-22c3caa50589). It came back completed with 0 answers, so the program stays on page 1.

weibo/web-v2/realtime-search with {"query": "华为Mate90", "page": 1} returned data.parsed_data.results, 9 or 10 posts per page in my calls, each with weibo_id, content and an interaction object of comment_count, like_count and repost_count. Every Weibo post in my fetches showed 0 likes (they were minutes old), so the likes-weighted share effectively reflects Zhihu and Bilibili. Pages 2 and 3 returned new post ids.

bilibili/web/video-comments with {"bv_id": "<video id>", "pn": 1} was the flaky one. On the most played video, one call (run 93cbc90a-332c-4b97-8206-3b3d4889e926) returned only 3 replies, though its data.page object said:

{"acount": 26460, "count": 26460, "num": 1, "size": 20}

A later pn: 1 call returned 20 replies, pn: 2 returned 0 with every page field at 0, and pn: 3 returned 20. So the program retries page 1 until it has at least 10 replies, up to three times, and doesn’t page further.

SandBase endpoint reference for weibo/web-v2/realtime-search showing the POST route with query and page fields and an outputs[0].data response envelope

Caption: The Weibo realtime search reference documents the POST route, a required query and an optional page, and an envelope whose outputs[0].data example is empty, so the post fields above come from live calls (captured 2026-10-03).

SandBase endpoint reference for zhihu/web/question-answers listing cursor, limit, offset, order, question_id and session_id

Caption: The Zhihu question-answers reference lists cursor, offset and session_id for paging; in my test the second page came back empty, which is why the agent reads one page of 20 answers (captured 2026-10-03).

The complete program

One file, standard library plus requests. The first run fetches and saves items.json. Pass that file as a fourth argument to label the exact same items with another model, which is how I compared the two.

export SANDBASE_API_KEY=...   # set it in your shell, never in the file
python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6-luna
python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6.1-sol items.json
#!/usr/bin/env python3
"""Three-platform public opinion brief (Weibo, Zhihu, Bilibili) for one consumer/tech topic.

Usage: python3 opinion_brief.py <search keyword> <aliases, comma-separated> <model> [items.json]
Example: python3 opinion_brief.py "华为Mate90" "Mate 90,Mate90" openai/gpt-6-luna
Pass a saved items.json as the 4th argument to skip fetching and reuse the same items.
Writes items.json (with text, keep private), labels.csv (ids + labels only) and brief.md.
"""
import csv
import html
import json
import os
import re
import sys
import time
from collections import Counter

import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
PRICES = {"openai/gpt-6-luna": (0.1, 0.5), "openai/gpt-6.1-sol": (2.0, 10.0)}  # USD per 1M tokens, 2026-10-03
STANCES = ["positive", "negative", "neutral", "off_topic"]
THEMES = ["chip_performance", "camera", "price_value", "design_build", "battery_charging",
          "software_system", "availability_sales", "other"]


def call(model: str, params: dict, retries: int = 2) -> dict:
    """POST /v1/api/<model>. Prefer outputs[0].data, fall back to outputs[0], then output."""
    for attempt in range(retries + 1):
        try:
            resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
            body = resp.json()
        except requests.RequestException as err:
            if attempt == retries:
                raise RuntimeError(f"{model}: {err}") from err
            time.sleep(3 * (attempt + 1))
            continue
        if resp.status_code >= 500 and attempt < retries:
            time.sleep(3 * (attempt + 1))
            continue
        if resp.status_code != 200 or body.get("status") != "completed":
            raise RuntimeError(f"{model}: HTTP {resp.status_code} {body.get('status')} {body.get('error')}")
        outputs = body.get("outputs") or []
        if outputs and isinstance(outputs[0], dict):
            return outputs[0].get("data", outputs[0]) or {}
        if body.get("output") is not None:
            return body["output"]
        raise RuntimeError(f"{model}: completed but no payload")
    raise RuntimeError(f"{model}: retries exhausted")


def matches(text: str, aliases: list[str]) -> bool:
    t = (text or "").lower().replace(" ", "")
    return any(a.lower().replace(" ", "") in t for a in aliases)


def clean(text: str, limit: int = 280) -> str:
    text = re.sub(r"<[^>]+>", " ", html.unescape(text or ""))
    return re.sub(r"\s+", " ", text).strip()[:limit]


def trending(aliases: list[str]) -> dict:
    """Where the topic sits on each platform's hot list right now (rank 1 = top)."""
    wb = call("weibo/web-v2/hot-search", {}).get("realtime") or []
    zh = (call("zhihu/web/hot-list", {"limit": "50"}).get("data") or [])
    bl = ((call("bilibili/web/hot-search", {"limit": 30}).get("data") or {}).get("trending") or {}).get("list") or []
    hits = {"weibo": [i + 1 for i, r in enumerate(wb) if matches(r.get("word"), aliases)],
            "zhihu": [i + 1 for i, r in enumerate(zh) if matches((r.get("target") or {}).get("title"), aliases)],
            "bilibili": [i + 1 for i, r in enumerate(bl) if matches(r.get("keyword"), aliases)]}
    zhihu_qid = next((str(r["target"]["id"]) for r in zh
                      if matches((r.get("target") or {}).get("title"), aliases)), None)
    return {"hot_list_ranks": hits, "zhihu_question_id": zhihu_qid}


def collect(keyword: str, aliases: list[str]) -> tuple[dict, list[dict]]:
    trend, items = trending(aliases), []
    for page in (1, 2):
        data = call("weibo/web-v2/realtime-search", {"query": keyword, "page": page})
        for r in (data.get("parsed_data") or {}).get("results") or []:
            inter = r.get("interaction") or {}
            items.append({"id": f"wb:{r.get('weibo_id')}", "platform": "weibo", "text": clean(r.get("content")),
                          "likes": int(inter.get("like_count") or 0)})
    if trend["zhihu_question_id"]:
        data = call("zhihu/web/question-answers", {"question_id": trend["zhihu_question_id"], "limit": 20})
        for a in data.get("data") or []:
            t = a.get("target") or {}
            items.append({"id": f"zh:{t.get('id')}", "platform": "zhihu",
                          "text": clean(t.get("content") or t.get("excerpt")), "likes": int(t.get("voteup_count") or 0)})
    data = call("bilibili/web/general-search", {"keyword": keyword, "order": "totalrank", "page": 1, "page_size": 20})
    videos = [v for v in (data.get("data") or {}).get("result") or [] if matches(clean(v.get("title")), aliases)]
    for v in sorted(videos, key=lambda v: int(v.get("play") or 0), reverse=True)[:2]:
        for _ in range(3):  # page 1 sometimes comes back short or empty; retry it
            replies = (call("bilibili/web/video-comments", {"bv_id": v["bvid"], "pn": 1}).get("data") or {}).get("replies") or []
            if len(replies) >= 10:
                break
            time.sleep(2)
        for c in replies:
            items.append({"id": f"bl:{c.get('rpid')}", "platform": "bilibili",
                          "text": clean((c.get("content") or {}).get("message")), "likes": int(c.get("like") or 0)})
    items = [i for i in {i["id"]: i for i in items}.values() if i["text"]]  # dedupe, drop empty
    return trend, items


def chat(model: str, prompt: str, max_tokens: int) -> tuple[str, dict, float, str]:
    t0 = time.time()
    resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=300,
                         json={"model": model, "max_tokens": max_tokens,
                               "messages": [{"role": "user", "content": prompt}]})
    resp.raise_for_status()
    body = resp.json()
    return body["choices"][0]["message"]["content"] or "", body.get("usage") or {}, time.time() - t0, body.get("id")


def classify(model: str, topic: str, items: list[dict], batch: int = 25) -> tuple[dict, list[dict]]:
    labels, calls = {}, []
    for start in range(0, len(items), batch):
        chunk = [{"id": i["id"], "text": i["text"]} for i in items[start:start + batch]]
        prompt = (f"Topic: {topic}. Label each post for public opinion monitoring.\n"
                  f"stance toward the topic product: one of {STANCES} (neutral = factual, question or mixed; "
                  f"off_topic = not about the product).\ntheme: the main aspect, one of {THEMES}.\n"
                  'Reply with JSON only: {"labels": [{"id": "...", "stance": "...", "theme": "..."}]}\n'
                  + json.dumps(chunk, ensure_ascii=False))
        text, usage, secs, rid = chat(model, prompt, 4000)
        calls.append({"step": "classify", "id": rid, "usage": usage, "seconds": round(secs, 1)})
        found = re.search(r"\{.*\}", text, re.S)
        for lab in (json.loads(found.group(0)).get("labels") if found else []) or []:
            if lab.get("stance") in STANCES and lab.get("theme") in THEMES:
                labels[lab["id"]] = {"stance": lab["stance"], "theme": lab["theme"]}
    return labels, calls


def aggregate(trend: dict, items: list[dict], labels: dict) -> dict:
    stats = {"hot_list_ranks": trend["hot_list_ranks"], "platforms": {}}
    for p in ("weibo", "zhihu", "bilibili"):
        rows = [labels[i["id"]] | {"likes": i["likes"]} for i in items if i["platform"] == p and i["id"] in labels]
        on = [r for r in rows if r["stance"] != "off_topic"]
        st = Counter(r["stance"] for r in on)
        stats["platforms"][p] = {
            "items_labeled": len(rows), "on_topic": len(on),
            "positive_share": round(st["positive"] / len(on), 2) if on else None,
            "negative_share": round(st["negative"] / len(on), 2) if on else None,
            "top_themes": Counter(r["theme"] for r in on).most_common(3),
            "likes_on_negative_share": round(sum(r["likes"] for r in on if r["stance"] == "negative")
                                             / max(1, sum(r["likes"] for r in on)), 2)}
    stats["unlabeled_items"] = len(items) - len(labels)
    return stats


def unknown_numbers(text: str, stats: dict, names: str) -> list[str]:
    """Numbers in the brief not in the stats JSON or the topic names (percent forms of shares allowed)."""
    known = set(re.findall(r"\d+", names))
    for v in re.findall(r"\d+(?:\.\d+)?", json.dumps(stats)):
        known |= {v, v.rstrip("0").rstrip(".") if "." in v else v}
        if float(v) < 1:
            known |= {str(round(float(v) * 100))}
    found = [n.replace(",", "") for n in re.findall(r"\d[\d,]*(?:\.\d+)?", text)]
    return sorted({n for n in found if n not in known})


def cost(model: str, calls: list[dict]) -> float:
    pin, pout = PRICES[model]
    return round(sum(c["usage"].get("prompt_tokens", 0) * pin + c["usage"].get("completion_tokens", 0) * pout
                     for c in calls) / 1e6, 5)


if __name__ == "__main__":
    keyword, aliases, model = sys.argv[1], sys.argv[2].split(","), sys.argv[3]
    if len(sys.argv) > 4:
        saved = json.load(open(sys.argv[4], encoding="utf-8"))
        trend, items = saved["trend"], saved["items"]
    else:
        trend, items = collect(keyword, aliases)
        json.dump({"trend": trend, "items": items}, open("items.json", "w", encoding="utf-8"), ensure_ascii=False)
    labels, calls = classify(model, keyword, items)
    stats = aggregate(trend, items, labels)
    text, usage, secs, rid = chat(model, (
        f"Write a 120-180 word public opinion brief in English about '{keyword}' for a product team: where it trends, "
        "overall stance per platform, main themes, one risk to watch. Shares are fractions of on-topic items. "
        "Use only numbers in this JSON; do not compute new ones.\n" + json.dumps(stats)), 1500)
    calls.append({"step": "summary", "id": rid, "usage": usage, "seconds": round(secs, 1)})
    flags = unknown_numbers(text, stats, keyword + " " + " ".join(aliases))
    with open("labels.csv", "w", newline="", encoding="utf-8") as f:
        w = csv.writer(f)
        w.writerow(["id", "platform", "likes", "stance", "theme"])
        for i in items:
            lab = labels.get(i["id"], {})
            w.writerow([i["id"], i["platform"], i["likes"], lab.get("stance", ""), lab.get("theme", "")])
    with open("brief.md", "w", encoding="utf-8") as f:
        f.write(f"# Opinion brief: {keyword} ({model})\n\n```json\n{json.dumps(stats, indent=2)}\n```\n\n{text}\n\n"
                + (f"Check these numbers: {', '.join(flags)}\n" if flags else "Numbers check: OK\n"))
    run = {"model": model, "items": len(items), "llm_calls": calls, "llm_seconds": round(sum(c["seconds"] for c in calls), 1),
           "est_cost_usd": cost(model, calls), "unverified_numbers": flags}
    json.dump(run, open("run.json", "w"), indent=1)
    print(json.dumps(stats, indent=2))
    print(text)
    print(json.dumps({k: v for k, v in run.items() if k != "llm_calls"}))

A few choices worth knowing before you change it:

  • Batches of 25. 78 items became four classify calls of up to about 2,500 prompt tokens each. The JSON reply is parsed with a regex, and any label outside the allowed sets is dropped and counted in unlabeled_items. That stayed 0 in all my runs.
  • Hot-list positions are list positions. Weibo’s list can include an unranked promoted entry, so position isn’t always the official rank. Weibo’s hot list never matched the alias during my runs, so the brief shows an empty list there.
  • max_tokens is the documented Chat Completions field. Luna spent up to 1,386 reasoning tokens on a single classify call in my runs, so the classify budget is 4,000.
  • est_cost_usd is list price. It multiplies usage by $0.10/$0.50 (Luna) or $2/$10 (Sol) per million input/output tokens. The billed amount can differ; see the cost section.

To try it, open the bilibili/web/video-comments API reference and the Chat Completions guide, and get a SandBase API key.

How I tested the cheap model

The model choice comes from our cheap LLM tiers benchmark, where GPT-6 Luna scored 51/51 on short tool tasks at about 1/20 of GPT-6.1 Sol’s cost per task. Classifying jokey comments is a different job, so I measured it here.

  1. I ran the program once with Luna to fetch 61 items (18 Weibo posts, 20 Zhihu answers, 23 Bilibili comments) and froze that items.json.
  2. I took the first 16 Weibo, 17 Zhihu and 17 Bilibili items, 50 in all, and labeled stance and theme myself before opening any model output. I’m one annotator, so these labels are a reference, not ground truth. I marked 11 of the 50 as ambiguous (mixed reviews, sarcasm, in-jokes).
  3. I ran each model 3 times on the frozen file and scored each run with the script below.
  4. Later I ran the program end to end once per model. Each run fetched fresh, and both fetches returned the same 78 items with identical text, so the labeling prompts were identical. Those runs gave the briefs quoted at the top.

My 50 labels: 24 positive, 12 neutral, 7 negative, 7 off-topic. Chip performance was the top theme (17), then other (12), camera (7) and price/value (6). The scorer computes accuracy and Cohen’s kappa, which discounts agreement you’d get by chance:

#!/usr/bin/env python3
"""Compare a run's labels.csv with hand labels: accuracy and Cohen's kappa for stance.

Usage: python3 score_labels.py hand_labels.json labels.csv [labels.csv ...]
hand_labels.json maps item id -> {"stance": ..., "theme": ...}.
"""
import csv
import json
import sys
from collections import Counter


def kappa(pairs: list[tuple[str, str]]) -> float:
    n = len(pairs)
    observed = sum(a == b for a, b in pairs) / n
    ca, cb = Counter(a for a, _ in pairs), Counter(b for _, b in pairs)
    expected = sum(ca[k] * cb[k] for k in ca) / (n * n)
    return round((observed - expected) / (1 - expected), 3) if expected < 1 else 1.0


gold = json.load(open(sys.argv[1], encoding="utf-8"))
for path in sys.argv[2:]:
    pred = {r["id"]: r for r in csv.DictReader(open(path, encoding="utf-8"))}
    ids = [i for i in gold if i in pred]
    stance = [(gold[i]["stance"], pred[i]["stance"] or "missing") for i in ids]
    theme = [(gold[i]["theme"], pred[i]["theme"] or "missing") for i in ids]
    print(json.dumps({"file": path, "scored": len(ids),
                      "stance_acc": round(sum(a == b for a, b in stance) / len(ids), 3),
                      "stance_kappa": kappa(stance),
                      "theme_acc": round(sum(a == b for a, b in theme) / len(ids), 3),
                      "both_acc": round(sum(gold[i]["stance"] == pred[i]["stance"] and
                                            gold[i]["theme"] == pred[i]["theme"] for i in ids) / len(ids), 3)}))

Results: GPT-6 Luna vs GPT-6.1 Sol

Three runs per model on the same 61 items, scored on the 50 I labeled:

Metric (range over 3 runs)GPT-6 LunaGPT-6.1 Sol
Stance accuracy0.70 to 0.800.78 to 0.84
Stance kappa0.556 to 0.7040.670 to 0.755
Theme accuracy0.80 to 0.820.86 to 0.90
Stance and theme both right0.62 to 0.740.70 to 0.78
Stance accuracy, 39 clear items0.769 to 0.8210.846 to 0.897
Stance accuracy, 11 ambiguous items0.455 to 0.8180.545 to 0.636
Run-to-run stance agreement, all 61 items0.852 to 0.9020.934 to 0.951
LLM time per brief (s)45.6 to 49.839.9 to 51.5
Cost per brief at list price (USD)0.00257 to 0.002730.02962 to 0.03268

Sol is better on every row except the 11 ambiguous items, where Luna’s runs swung from 0.455 to 0.818 and Sol stayed at 0.545 to 0.636. Overall the gap is small: on average two or three more correct stances out of 50. The bigger difference is stability. Luna’s three runs disagreed with each other on 6 to 9 of 61 stances; Sol’s on 3 or 4. One Luna run dropped to 0.70 stance accuracy, and most of its extra errors were on the ambiguous items.

Both models made the same most common mistake. In every run, 4 items I’d marked off-topic got labeled neutral: a sticker-only reply, remarks about how much attention the topic was getting, and unrelated jokes. That’s arguably a taxonomy problem, not a model problem. If off-topic matters to you, give it a sharper definition or examples.

What a product team reads is the aggregate. On the 50 labeled items, my labels gave Bilibili 0.42 positive and 0.25 negative share. Sol’s three runs gave 0.38 to 0.43 positive and 0.21 to 0.25 negative. Luna’s gave 0.33 to 0.45 positive and 0.08 to 0.27 negative. Sol has its own bias: it found no negative Zhihu answers in any run, where I’d counted 2 of 15 on-topic ones.

Where Luna slipped

The end-to-end run made the Bilibili problem concrete. Of 14 Bilibili items where the models disagreed, 4 were negative under Luna and positive under Sol, and another 4 were off-topic under Luna and positive under Sol. Several were mock boasts written in the voice of a rival chip, which Bilibili users post as praise for the new one. Luna read them literally. One dismissive comment comparing the chip to an older Snapdragon part was negative under Luna and positive under Sol, and there I agree with Luna. Neither model is reliable on irony. Luna just misses it more often, and on a platform where a lot of the comments are irony, that moves the headline number.

SandBase model page for openai/gpt-6-luna showing $0.10 input price and $0.50 output price per million tokens and 128K max output

Caption: The GPT-6 Luna model page, scrolled to its price row, lists $0.10 per million input tokens and $0.50 per million output tokens, the list prices behind the cost column (captured 2026-10-03).

What a brief costs

The seven data endpoints returned base_price: "0" from GET /v1/models/<model> on 2026-10-03, and the Bilibili comments model page showed Free under an “API Free Week” banner. That’s a dated listing during a promotion, not a promise. The LLM side, from model_card.price_formula: Luna $0.10/$0.50 and Sol $2/$10 per million input/output tokens below 272K prompt tokens.

SandBase model page for bilibili/web/video-comments showing a Free base price, sync execution and 2 input fields under an API Free Week banner

Caption: The bilibili/web/video-comments model page shows a Free base price, sync execution and 2 input fields (bv_id and pn), with the API Free Week banner across the top (captured 2026-10-03).

Measured per brief, list price from usage (est_cost_usd) vs what GET /v1/tasks/<id>/cost reported:

RunLuna list priceLuna billedSol list priceSol billed
End to end, 78 items, fresh prompts$0.00331$0.003457$0.03749$0.040513
Repeat on the frozen 61 items (one run each)$0.00257$0.00214$0.02962$0.020614

The cost records explain the gap. Fresh prompts showed cache_creation_tokens and billed 4 to 8% above list. Repeats of identical prompts showed cached_tokens and billed 17% (Luna) to 30% (Sol) below list. My code ignores caching, so treat est_cost_usd as a ballpark. I could reconcile billing for those four runs only. Lookups for the same ids worked in the first minutes after the runs and returned “task not found” later, so the other runs are list-price estimates.

GET /v1/models/<id> and GET /v1/tasks/<id>/cost need the same Bearer key as the calls themselves. Without a key, the public model pages show the same list prices.

Luna’s per-call output was larger than Sol’s, because it spent more reasoning tokens (on the three full batches of the end-to-end runs, completion tokens ran 1,125 to 1,897 per call for Luna vs 618 to 717 for Sol). So a Luna brief cost about 1/11 to 1/12 of a Sol brief, not the 1/20 the price list suggests.

Which model to use

SituationPick
Daily brief on Weibo posts and Zhihu answersGPT-6 Luna
Bilibili comments, or any meme-heavy comment sectionGPT-6.1 Sol, or Luna run twice with disagreements sent to Sol
One headline number goes to leadershipGPT-6.1 Sol, plus a manual look at negatives
Thousands of items a dayLuna, with a 50-item hand check per new topic

A cheap way to get most of Sol’s stability: run Luna twice and send only the items where the two runs disagree to Sol. In my runs that would have been 6 to 9 items out of 61.

Limits

Scope: one topic, one day, 61 items for the scored comparison and 78 for the end-to-end runs, 3 runs per model, one annotator. I didn’t test other topics, other languages for the brief, or the brief’s quality beyond the numbers check, which passed in every final run. Hot lists and search results change by the minute; the Bilibili trending position moved from 21 to 25 over an hour. The brief describes what the endpoints returned at that time, not public opinion at large.

FAQ

Is GPT-6 Luna good enough for public opinion monitoring?

For Weibo posts and Zhihu answers in this test, yes: within a few points of Sol on stance and theme at about 1/11 to 1/12 of the cost. For Bilibili comments it was less stable and missed irony often enough to change the headline share.

Why not let the LLM compute the shares?

Counting is where models slip, as our earlier benchmarks showed. Python computes every share and count, and the brief’s numbers check flagged nothing in the final runs.

Why only one page of Zhihu answers and Bilibili comments?

In my calls the Zhihu next-page request came back empty, and Bilibili’s page 2 sometimes returned nothing. One page per source kept runs predictable. More data means more calls to verify first.

Can I publish the comments the agent collected?

Not verbatim. They’re written by individuals. Publish aggregates and generic paraphrases, as this article does, and keep items.json private.

How do I check what a run cost?

Read the chat completion’s id and call GET /v1/tasks/<id>/cost with your key, soon after the run. In my test the same ids returned the cost record within minutes but “task not found” about half an hour later, so store the result when you get it.