Blog/Developer Tools/

Bilibili Video Comment Analysis API Tutorial | SandBase

Analyze a Bilibili video's audience: read stats from a BV id, page hot comments, sample replies, and bin danmaku by minute. Four SandBase endpoints, one API key.

Dark cinematic render of a glass video screen streaming bullet-comment light dashes into branching comment bubbles that flow into an agent core

A tech channel posts a two-minute recap of a phone launch, and by the next morning it has 2,199 comments and 1,404 danmaku, the bullet comments that scroll across the player. The comments argue about price. The danmaku react to specific seconds of the video. If you want to know how an audience took a launch, a trailer, or a competitor’s ad, Bilibili gives you two signals, and you need both as data. This Bilibili Video Comment Analysis API tutorial starts from a BV id, reads the video’s stats, pages through hot comments, samples the busiest reply threads, and bins danmaku by minute, using four SandBase endpoints. It builds on the Bilibili public data API hub; read that for the full endpoint map.

All of this is public, read-only data. You need no Bilibili login, no cookie, and no SDK. You do authenticate with a SandBase API key. The four endpoints are currently listed as Free in the SandBase catalog.

The endpoint API reference is the source of truth for parameters and the response envelope, and the envelope is all it guarantees. The payload field names below come from calls I ran (tested on 2026-10-01, UTC). Treat them as illustrative and observed-only, not documented guarantees, and check them against a live response before you build on them.

Key takeaway

  • One BV id is enough. bilibili/web/one-video returns the title, channel, a stat block (views, likes, coins, favorites, replies, danmaku), the aid, and the cid that danmaku needs.
  • bilibili/app/video-comments pages hot comments 20 at a time. Send the returned cursor.next as the next request’s integer next_offset.
  • bilibili/web/comment-reply reads a thread from bv_id plus the comment’s rpid, paged with pn.
  • bilibili/web/video-danmaku takes the cid and returns raw XML as a string, not JSON. On a larger video it returned a sample, not every danmaku.

Why comments and danmaku, and why not the official route

Bilibili’s open platform is built for creators and partners who manage their own content through an authorized app. For reading reaction on someone else’s public video, most developers end up with scrapers, cookies, and request signing that changes without notice. That maintenance is most of the work for a one-off launch review.

The SandBase route is narrower. It reads public video data, comments, replies, and danmaku through one key, and it does not post, like, or touch anything account-bound. The creator research tutorial covers trends, search, and creator profiles. This one covers what viewers say about a single video.

Your needUse
Public comments, replies, and danmaku for research or reaction summariesSandBase Bilibili public-data API
Manage, reply to, or moderate content on your own accountBilibili’s official creator and open-platform tools
Private, members-only, or account-bound dataNeither public workflow

The workflow at a glance

  1. Get the BV id from the URL (bilibili.com/video/BV...).
  2. Read the video with bilibili/web/one-video (bv_id) for title, channel, stat, aid, and cid.
  3. Page hot comments with bilibili/app/video-comments (bv_id, mode, then next_offset).
  4. Expand the busiest threads with bilibili/web/comment-reply (bv_id, rpid, pn).
  5. Read danmaku with bilibili/web/video-danmaku (cid) and bin it by playback minute.
  6. Summarize comments, replies, and the danmaku timeline with a model.

SandBase Bilibili API catalog page listing Bilibili endpoints with Free status The Bilibili catalog on SandBase: 38 endpoints shown as GET /apis/v1/bilibili/..., with the selected “Get video playurl” endpoint marked Available, Free. This tutorial calls the Model API POST /v1/api/bilibili/... routes instead.

A note on surfaces before the code. The catalog page shows each operation as GET /apis/v1/bilibili/<path>. This tutorial uses the Model API surface from the endpoint reference: POST /v1/api/bilibili/<path> with a JSON body that holds only that endpoint’s parameters. One generated cURL example in the reference (for app/video-comments) puts a model field in the body. My calls worked without it, because the path already names the endpoint. Keep the POST method and the /v1/api/ prefix, and don’t mix them with the catalog GET paths.

Step 0: one helper for every call

import os
import re
import time
import xml.etree.ElementTree as ET
from collections import Counter

import requests

API = "https://api.sandbase.ai/v1/api"
HEADERS = {
    "Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}",
    "Content-Type": "application/json",
}

def call(path: str, payload: dict, retries: int = 2):
    for attempt in range(retries + 1):
        try:
            resp = requests.post(f"{API}/{path}", headers=HEADERS, json=payload, timeout=90)
        except requests.exceptions.ConnectionError:
            if attempt < retries:
                time.sleep(2 * (attempt + 1))  # dropped TLS connection, try again
                continue
            raise
        if resp.status_code >= 500 and attempt < retries:
            time.sleep(2 * (attempt + 1))
            continue
        resp.raise_for_status()
        body = resp.json()
        if body.get("status") != "completed":
            raise RuntimeError(body.get("error", {}).get("message", f"{path} did not complete"))
        # Reference documents outputs[0].data; some calls return a top-level `output`.
        output = body.get("output")
        if output is None and body.get("outputs"):
            output = body["outputs"][0].get("data")
        print(f"  {path}: run {body.get('id')}")
        return output
    return None

def bvid_from_url(url: str) -> str:
    m = re.search(r"(BV[0-9A-Za-z]{10})", url)
    if not m:
        raise ValueError(f"no BV id in {url}")
    return m.group(1)

def clean_text(message: str) -> str:
    # Replies start with "回复 @someone :" — drop the handle, keep what was said.
    return re.sub(r"^回复\s*@[^::]+[::]\s*", "", message or "").strip()

The reference documents completed responses as outputs[0].data, and every Bilibili call in this test returned that shape. I’ve seen a top-level output on other SandBase platform endpoints, so the helper reads both and prefers output when present. Two details are different from most platforms:

  • The helper returns whatever data holds, not always a dict. For the JSON endpoints it’s Bilibili’s own wrapper, {code, data, message, ttl}, so the useful fields sit one level deeper under data. For danmaku it’s a plain XML string.
  • The retry catches connection errors, not only 5xx. My first one-video call for the test video failed with a TLS EOF before any HTTP status came back, and the same request succeeded on the next try.

clean_text exists because reply text embeds the handle of the person being answered. You want the opinion, not the name.

Step 1: read the video

def read_video(bv_id: str) -> dict:
    v = (call("bilibili/web/one-video", {"bv_id": bv_id}) or {}).get("data") or {}
    stat = v.get("stat") or {}
    return {
        "bv_id": bv_id,
        "aid": v.get("aid"),
        "cid": v.get("cid"),  # first part; multi-part videos list more under `pages`
        "title": v.get("title"),
        "channel": (v.get("owner") or {}).get("name"),
        "published": v.get("pubdate"),  # unix seconds
        "duration": v.get("duration"),
        "views": stat.get("view"),
        "likes": stat.get("like"),
        "coins": stat.get("coin"),
        "favorites": stat.get("favorite"),
        "shares": stat.get("share"),
        "comments": stat.get("reply"),
        "danmaku": stat.get("danmaku"),
    }

My test video was a public recap from the tech media channel Kejimeixue, a two-minute launch-event recap of the Xiaomi 18 Pro / Pro Max (BV19uhb6mEWd, published 2026-09-23 UTC). Run 11f78f98-533c-463e-8b1f-c8de5c165752 returned this, trimmed from outputs[0].data:

{
  "code": 0,
  "data": {
    "aid": 117321125400871,
    "bvid": "BV19uhb6mEWd",
    "cid": 42143190519,
    "title": "两分钟发布会 | 小米 18 Pro / Pro Max亮相 小米平板9 小米手环11 小米手表S5 还有咖啡机??",
    "owner": { "name": "科技美学" },
    "pubdate": 1790178608,
    "duration": 710,
    "pages": [{ "cid": 42143190519, "page": 1, "duration": 710 }],
    "stat": {
      "view": 247824, "like": 6310, "coin": 445, "favorite": 741,
      "share": 266, "reply": 2199, "danmaku": 1404
    }
  },
  "message": "OK"
}

Unlike YouTube, the counts are plain integers, so no parsing is needed. stat.reply sizes the comment job and stat.danmaku tells you what to expect from step 4. The cid is the one id you can’t get from the URL; one-video saves you a call to bilibili/web/video-parts, which returned the same cid for this single-part video (run f3862a4d-034b-451e-b924-c452af230b24). For multi-part videos, loop over pages and read danmaku per part.

I didn’t need bilibili/web/bv-to-aid either. The comment endpoints I used accept bv_id directly, and one-video returns the aid anyway. One small note: tname (the category) came back empty on this video.

SandBase API reference for the Bilibili web one-video endpoint The bilibili/web/one-video reference (“Get single video data”): POST /v1/api/bilibili/web/one-video with one required string, bv_id. The documented response example leaves outputs[0].data empty.

Step 2: page through hot comments

def top_comments(bv_id: str, max_pages: int = 3, mode: int = 3) -> list[dict]:
    payload = {"bv_id": bv_id, "mode": mode}  # 3 = hot, 2 = newest
    out, seen = [], set()
    for _ in range(max_pages):
        data = (call("bilibili/app/video-comments", payload) or {}).get("data") or {}
        for r in data.get("replies") or []:
            if r.get("rpid") in seen:
                continue
            seen.add(r.get("rpid"))
            out.append({
                "rpid": str(r.get("rpid")),
                "text": clean_text((r.get("content") or {}).get("message", "")),
                "likes": r.get("like", 0),
                "replies": r.get("rcount", 0),
                "ctime": r.get("ctime"),
            })
        cursor = data.get("cursor") or {}
        if cursor.get("is_end") or not cursor.get("next"):
            break
        payload = {"bv_id": bv_id, "mode": mode, "next_offset": cursor["next"]}
    return out

I started with bilibili/web/video-comments, which takes bv_id and a page number pn, and switched after it misbehaved. Page 1 (run b1b5caeb-43f0-43de-b4e9-44f20245f03c) returned 20 comments and page.count: 2199. Page 2 (run fa847910-aa82-4166-b3b4-17cf5357f45d) returned zero comments and page.count: 0. A retry of page 2 (6f7e6258-f781-4ffb-b00b-a6b1bc168090) worked, and page 3 (139f2dca-660c-4fa9-9c25-d5caef347687) was empty again. An empty page that looks exactly like “the end” is hard to page against.

The app endpoint was steadier. Its first page (run 98654147-fe5a-4a47-8a2a-b23282e23575) returned 20 comments and a cursor object:

{
  "all_count": 2199,
  "is_begin": true,
  "is_end": false,
  "mode": 3,
  "name": "热门评论",
  "next": 2,
  "pagination_reply": { "next_offset": "CAEiAggC" },
  "support_mode": [2, 3]
}

There are two candidate cursors here, and only one fits the schema. The reference types next_offset as an integer, and sending the string "CAEiAggC" failed with HTTP 400, expected integer, but got string. Sending cursor.next worked: next_offset: 2 (run 9706405a-a879-4769-a3b5-918945e400b0) returned 20 new comments with no rpid overlap and next: 3, and next_offset: 3 (run af057287-9ad9-44b6-b434-4bb7c0ae93b3) returned another 20 with next: 4. The helper stops on is_end and still dedupes, because nothing documents that pages never overlap.

Each comment carried rpid, like, rcount, ctime, and content.message, plus a member object with the commenter’s name, avatar, and mid. The helper drops all identity fields. One comment, as kept:

{ "rpid": "314848758273", "text": "哈哈,我花这钱来买安卓?哈哈哈哈哈", "likes": 92, "replies": 21 }

Use rcount for reply counts, not count. On that comment rcount was 21 and count was 47. The replies endpoint in step 3 reported 21, matching rcount. I can’t tell from the API what the extra 26 in count represent.

SandBase API reference for the Bilibili app video-comments endpoint The bilibili/app/video-comments reference: optional av_id or bv_id (choose one), mode (3 = hot, 2 = time, default 3), and an integer next_offset described as the pagination cursor, default 1.

Step 3: expand the busiest reply threads

def thread_replies(bv_id: str, rpid: str, max_pages: int = 2) -> list[dict]:
    out = []
    for pn in range(1, max_pages + 1):
        data = (call("bilibili/web/comment-reply", {"bv_id": bv_id, "rpid": rpid, "pn": pn}) or {}).get("data") or {}
        if not (data.get("page") or {}).get("count"):
            # Seen once: a page came back with count 0 and no replies, then filled on retry.
            data = (call("bilibili/web/comment-reply", {"bv_id": bv_id, "rpid": rpid, "pn": pn}) or {}).get("data") or {}
        batch = data.get("replies") or []
        out += [{"text": clean_text((r.get("content") or {}).get("message", "")), "likes": r.get("like", 0)}
                for r in batch]
        page = data.get("page") or {}
        if not batch or pn * (page.get("size") or 20) >= (page.get("count") or 0):
            break
    return out

comment-reply needs the video’s bv_id and the parent comment’s rpid; pn is optional. For the 21-reply thread above, page 1 (run b3d5734d-0e61-4890-9cab-ab4e37b4800b) returned 20 replies with page: {count: 21, num: 1, size: 20} and a root object holding the parent comment. Page 2 (run 54035a3e-044e-4bfa-8247-48a10125db05) came back empty with count: 0, the same failure as the web comments endpoint. The retry (1d3c6dc2-a2ba-4fd3-b883-bbd68c9b0c72) returned page.num: 2 and count: 21. Hence the single retry when count is zero. The loop stops when pn * size covers count, so it doesn’t request a page that shouldn’t exist.

Replies are where clean_text matters. On that page, 15 of 20 began with 回复 @<name> : (“Reply @:”), which puts another user’s handle into your data. Stripping that prefix keeps the reply readable (in one case, a short note on the typical price of a mainland-China iPhone Air 12GB+256GB) without the handle.

SandBase API reference for the Bilibili web comment-reply endpoint The bilibili/web/comment-reply reference (“Get reply to the specified comment”): required bv_id and rpid strings, plus an optional integer pn for the page number.

Step 4: read danmaku and bin it by minute

def danmaku(cid) -> list[dict]:
    xml_text = call("bilibili/web/video-danmaku", {"cid": str(cid)})
    if not isinstance(xml_text, str) or not xml_text.strip():
        return []
    root = ET.fromstring(xml_text.encode("utf-8"))
    rows = []
    for d in root.findall("d"):
        attrs = (d.get("p") or "").split(",")
        # Keep only the playback offset (first field) and the text; drop the rest.
        rows.append({"t": float(attrs[0]) if attrs and attrs[0] else 0.0, "text": d.text or ""})
    return rows

This is the step that surprised me most. The reference types outputs[0].data as “object | array”. For danmaku (run 3a0efcf1-8b87-43bd-a60f-b3393d4f8ce2) it was a 123,221-character XML string. Trimmed:

<i>
  <chatid>42143190519</chatid>
  <maxlimit>1500</maxlimit>
  <d p="99.22400,1,25,16777215,…">懂了:pro性能不足 max烫手</d>
  <d p="390.45500,5,25,16765698,…">我为什么要在手机背屏上弹吉他?</d>
</i>

Each <d> element is one danmaku. Its p attribute is a comma-separated list. On this video the first field always fell between 0.7 and 707.9, inside the 710-second duration, so I read it as the playback offset in seconds. Later fields include what looks like a per-sender hash. The helper keeps only the offset and the text, and drops the rest.

Count what you get. Here the XML held 1,404 <d> elements, exactly stat.danmaku. On a second, older video (BV1M1421t7hT, stat.danmaku 3,281), the same endpoint returned 1,200 elements with maxlimit 1,000 (run 196f4179-b89e-457b-ae43-88828b7bef66). So for busy videos, treat danmaku as a sample, and don’t report its totals as the video’s totals.

SandBase API reference for the Bilibili web video-danmaku endpoint The bilibili/web/video-danmaku reference (“Get Video Danmaku”): one required string, cid. The example response shows outputs[0].data as an empty object, while my live call returned an XML string.

Putting it together: one reaction record per video

def analyze(bv_id: str, comment_pages: int = 3, threads: int = 3) -> dict:
    video = read_video(bv_id)
    comments = top_comments(bv_id, comment_pages)
    busiest = sorted(comments, key=lambda c: c["replies"], reverse=True)[:threads]
    for c in busiest:
        c["reply_sample"] = thread_replies(bv_id, c["rpid"])
    dm = danmaku(video["cid"]) if video.get("cid") else []
    per_minute = Counter(int(d["t"] // 60) for d in dm)
    repeated = Counter(d["text"].strip() for d in dm if d["text"].strip()).most_common(10)
    return {
        "video": video,
        "comments": comments,
        "danmaku_count": len(dm),
        "danmaku_per_minute": dict(sorted(per_minute.items())),
        "danmaku_repeated": repeated,
        "danmaku_sample": [d["text"] for d in dm[:200]],
    }

report = analyze(bvid_from_url("https://www.bilibili.com/video/BV19uhb6mEWd/"))

I ran this end to end. It made 11 calls: one-video (a8abc9c6-d802-410a-8e47-01192919cc77), three comment pages (ca89f260-162e-429b-a9c4-2bb79ea01195, 7bcadf99-7938-47d5-80fd-2a7505e0b9ca, 647f6ca4-0535-46f2-a59b-ae913ef4373e), six reply pages across three threads (first d8fe09de-1123-4096-a067-33bab9194d8f), and one danmaku call (e3c4c8af-d0b7-44ba-84b4-1c630fc8fc72). It collected 60 hot comments out of 2,199, sampled 40, 27, and 30 replies from threads of 54, 27, and 30, found 9 comments with a question mark, and read all 1,404 danmaku.

The danmaku timeline was front-loaded: 210 in the first minute, 184 in the second, a low of 55 in the tenth minute, and a bump to 163 in the eighth. The most repeated danmaku were digits (“0” 99 times, “2” 33 times, “1” 30 times). The API can’t tell me what viewers were answering, but the “0”s clustered around 40 and 460 seconds, so that’s where I’d look in the video.

My reading of the 60 comments, not a measured distribution: the price dominated. The most-liked comment (368 likes) argued that brands now build the national subsidy into launch pricing, others compared the price with iPhones, and a few questioned the rear screen and privacy display. That kind of reading is what a model does well at scale. Pass the report dict with a prompt like “group these comments into themes, mark each positive, negative, or mixed, list distinct questions, and note which playback minutes drew the most danmaku”. Keep like counts in so the model can weight a 368-like comment above a 0-like one.

Cost scales with pages: one call for the video, one per 20 hot comments, one per 20 replies, and one per cid for danmaku.

Documented vs. observed

ItemStatus
POST /v1/api/bilibili/web/one-video with bv_idDocumented in the reference
POST /v1/api/bilibili/app/video-comments with bv_id/av_id, mode, integer next_offsetDocumented in the reference
POST /v1/api/bilibili/web/comment-reply with bv_id, rpid, pnDocumented in the reference
POST /v1/api/bilibili/web/video-danmaku with cidDocumented in the reference
Envelope id / status / model / outputs[0].dataDocumented in the reference
Upstream wrapper {code, data, message} with stat, aid, cid, owner, pagesObserved only
Comment rpid, like, rcount, content.message, and cursor.next / is_endObserved only
Reply page.count / num / size and rootObserved only
Danmaku as an XML string, first p field as playback secondsObserved only
Danmaku capped below stat.danmaku on a busier videoObserved only
Empty pages with count: 0 on web comment and reply endpointsObserved only

Common use cases

Launch reaction reviews

Run the pipeline on a brand’s own launch video and on the media recaps that follow. Compare comment themes, and line up the danmaku timeline against the video’s segments. Input: a handful of BV ids. Output: one reaction record each. Endpoints: all four.

Trailer and ad moment analysis

Danmaku is timestamped, comments aren’t. Bin danmaku by 10 or 30 seconds to see which moments of a trailer drew the most reaction, then read the danmaku text in the busiest bins. Input: BV id. Output: a per-second-bin histogram with sample text. Endpoints: one-video, video-danmaku.

FAQ mining

Collect comments with question marks across a product’s tutorial or review videos, then cluster them. Input: a list of BV ids. Output: deduplicated questions. Endpoint: app/video-comments.

Sponsorship checks

Before a sponsorship, sample hot comments on a creator’s recent videos and see whether the audience engages with the content or mostly jokes. Input: recent BV ids. Output: comment samples with like counts. Endpoints: one-video, app/video-comments.

Practical notes

  • Unwrap twice. Read outputs[0].data (or output), then Bilibili’s own data.
  • Page hot comments with the integer cursor.next. The string pagination_reply.next_offset is rejected by the schema.
  • Retry empty pages once. On the web comment and reply endpoints, an empty page with count: 0 was sometimes transient.
  • Use rcount, not count, for reply totals.
  • Parse danmaku as XML, and compare its length with stat.danmaku before you report totals.
  • Strip identity. Drop member, mid, and danmaku p fields beyond the offset, and remove 回复 @name : prefixes.
  • Retry connection errors, not only 5xx. One call failed at the TLS layer and succeeded on retry.
  • Public, read-only data only. No posting, liking, or moderation.

FAQ

Do I need a Bilibili account or cookie? No. You authenticate to SandBase with SANDBASE_API_KEY. These read endpoints need no Bilibili login on your side.

Is it free? The four endpoints are currently listed as Free in the SandBase catalog. Check the catalog for the current status.

Do I need to convert BV to aid first? Not for this chain. one-video returns the aid and cid, and the comment endpoints accept bv_id. bilibili/web/bv-to-aid exists if another endpoint needs an aid.

Does the danmaku endpoint return every danmaku? Not always. It matched stat.danmaku on a video with 1,404, and returned 1,200 of 3,281 on a busier one.

Can I get the newest comments instead of hot ones? The reference lists mode: 2 for time order. I used the default hot order (mode: 3) for this tutorial.

Wrap up

A BV id gets you both of Bilibili’s reaction signals. one-video sizes the job and hands you the cid, app/video-comments pages the argument with an integer cursor, comment-reply opens the threads where people push back, and video-danmaku shows which seconds made people type. Unwrap twice, retry empty pages, and treat danmaku as a sample on busy videos, and you have a record a model can turn into themes, questions, and a moment-by-moment timeline. For the rest of the Bilibili endpoints, see the Bilibili public data API hub. When you’re ready: