Blog/Developer Tools/

Xiaohongshu Note Research API Tutorial | SandBase

See which Xiaohongshu notes perform in a category: search notes, page results, open top notes, and sample comments with SandBase. No RED login; one API key.

Dark cinematic render of a search lens scanning a grid of note cards, the brightest flowing into a note detail card and a comment stream feeding an agent core

Before a brand briefs creators on Xiaohongshu (also called RED), someone usually spends an afternoon scrolling search results for the category. Which notes got saved, not just liked? Are the winners image posts or videos? Which hashtags keep showing up, and what are commenters actually asking? This Xiaohongshu note research API tutorial turns that afternoon into a script: search notes by keyword, page through results, open the top notes for full text, tags, and engagement, then sample their comments. It builds on the Xiaohongshu public data API hub, which has the full endpoint map.

This is the note angle. If you want shoppable products, SKUs, and product reviews instead, the Xiaohongshu product research workflow covers that side.

Everything here is public, read-only data. You need no Xiaohongshu login and no SDK, just a SandBase API key. The note endpoints are currently listed as Free in the SandBase catalog.

The endpoint API reference is the source of truth for parameters and the response envelope, and that envelope is all it guarantees. The payload field names below come from calls I ran (tested on 2026-10-01, UTC). Treat them as illustrative and observed-only, not documented guarantees, and confirm them against a live response before you depend on them.

Key takeaway

  • xiaohongshu/app-v2/search-notes takes a keyword and returns note cards with id, type, title, and like/collect/comment counts. Page it with page plus the search_id and search_session_id it returns.
  • Open image notes with app-v2/image-note-detail and video notes with app-v2/video-note-detail, both by note_id. The video route returns a list, so match on the id.
  • app-v2/note-comments samples comments by note_id; sort by like_count and page by sending back the returned cursor string.
  • Compare collects to likes. In my runs that ratio separated save-worthy guides from entertainment and giveaway posts better than likes alone.

SandBase or the official route

Your needUse
Public note, engagement, and comment data for category researchSandBase Xiaohongshu public-data API
Publishing, managing a brand account, ads, or licensed dataXiaohongshu’s official channels
Private, follower-only, or account-gated contentNeither public workflow

The workflow at a glance

  1. Expand the seed keyword (optional) with xiaohongshu/web-v3/search-suggest.
  2. Search notes with xiaohongshu/app-v2/search-notes (keyword, page) and page with the returned ids.
  3. Rank the cards by likes plus collects and pick the top few.
  4. Open each top note with image-note-detail or video-note-detail, depending on the card’s type.
  5. Sample comments with xiaohongshu/app-v2/note-comments.
  6. Aggregate: tag frequency, video share, median likes, collect-to-like ratio.

SandBase Xiaohongshu API catalog page listing the Xiaohongshu endpoints, with Search notes selected and shown as Available, Free The Xiaohongshu catalog on SandBase. It shows 34 endpoints as GET /apis/v1/xiaohongshu/..., with Search notes selected as Available, Free. This tutorial calls the Model API POST /v1/api/xiaohongshu/... routes instead.

A note on surfaces before the code. The catalog lists each operation as GET /apis/v1/xiaohongshu/<path>. This tutorial uses the Model API surface from the endpoint reference: POST /v1/api/xiaohongshu/<path> with a JSON body that holds only the endpoint’s parameters. Some generated reference examples include a model field in the body; the path already identifies the operation, so I left it out. Keep the POST method and the /v1/api/ prefix when you copy the examples.

Step 0: one helper for every call

import os
import time
import requests

API = "https://api.sandbase.ai/v1/api"
HEADERS = {
    "Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}",
    "Content-Type": "application/json",
}

def call(path: str, payload: dict, retries: int = 2) -> dict:
    for attempt in range(retries + 1):
        try:
            resp = requests.post(f"{API}/{path}", headers=HEADERS, json=payload, timeout=90)
            resp.raise_for_status()
            break
        except requests.RequestException:
            if attempt == retries:
                raise
            time.sleep(2 * (attempt + 1))
    body = resp.json()
    if body.get("status") != "completed":
        raise RuntimeError(body.get("error", {}).get("message", f"{path} did not complete"))
    # Reference documents outputs[0].data; also accept a top-level `output`.
    output = body.get("output")
    if output is None and body.get("outputs"):
        output = body["outputs"][0].get("data", {})
    output = output or {}
    # app-v2 routes used code 0, web-v3 used code 1000; both carried success: true.
    if output.get("success") is False:
        raise RuntimeError(f"{path}: upstream code {output.get('code')} {output.get('msg')}")
    return output

The reference documents completed responses as outputs[0].data, and every Xiaohongshu call in this test returned that shape. The helper also reads a top-level output, because I have seen that shape on other SandBase platform endpoints. Inside the payload sits an upstream wrapper: code, msg, success, and the real content under data.

My first version treated any non-zero code as a failure. That broke on the first call. web-v3/search-suggest answered with code: 1000 and msg: "成功" (success), while the app-v2 routes used code: 0. Both carried success: true, so the helper checks that instead. It also retries transient network errors, since one of my polling calls in this session died with a connection reset.

Step 1 (optional): expand the seed keyword

def suggest_keywords(seed: str) -> list[str]:
    data = call("xiaohongshu/web-v3/search-suggest", {"keyword": seed}).get("data", {})
    return [s.get("text") for s in data.get("sug_items", []) if s.get("text")]

For the seed 防晒霜 (sunscreen), the run I logged (74a3e360-568a-4688-b9c4-f347d630a1a0) suggested five longer phrases: sunscreen recommendations, light non-greasy sunscreen recommendations, how to apply sunscreen correctly, sunscreen recommendations for men, and sunscreen recommendations for military training. That’s already useful research: “recommend”, “non-greasy”, “how to apply”, “for men”, and “for military training” are the angles people search. Run the next step for each phrase if you want a wider sample.

Step 2: search notes and page through results

def search_notes(keyword: str, pages: int = 2) -> list[dict]:
    notes, params = [], {"keyword": keyword, "page": 1}
    for _ in range(pages):
        out = call("xiaohongshu/app-v2/search-notes", params)
        for item in out.get("data", {}).get("items", []):
            if item.get("model_type") != "note":
                continue  # skip ads and other cards
            n = item.get("note", {})
            notes.append({
                "note_id": n.get("id"),
                "type": n.get("type"),  # "normal" (image) or "video"
                "title": n.get("title"),
                "likes": n.get("liked_count", 0),
                "collects": n.get("collected_count", 0),
                "comments": n.get("comments_count", 0),
            })
        if not out.get("next_page"):
            break
        params = {
            "keyword": keyword,
            "page": out["next_page"],
            "search_id": out.get("search_id", ""),
            "search_session_id": out.get("search_session_id", ""),
        }
    return notes

The reference says search_id and search_session_id should be passed from the first response for pagination. The responses confirmed it. Page 1 (run 829c7829-b1d4-4588-abcb-bdabd4f2a16a) returned 20 items plus page: 1, next_page: 2, search_id, and search_session_id beside data. Sending those back with page: 2 (run c578ddab-1bfe-42f9-ac33-509e5bb39206) returned 20 more items, the same two ids, and next_page: 3.

Two things showed up on page 2. One item had model_type: "ads" instead of "note", which is why the loop filters on it. And results drift between runs: two pages gave me 40 unique notes in one run and 30 in the next. Dedupe by note_id, and don’t treat a single run as a stable ranking.

Each note card carried id, type, title, desc, liked_count, collected_count, comments_count, shared_count, a user object, and an xsec_token, all observed-only. The user object includes creator nicknames and ids. For category research you rarely need them, so the function leaves them out.

SandBase API reference for the Xiaohongshu app-v2 search-notes endpoint The xiaohongshu/app-v2/search-notes reference: POST /v1/api/xiaohongshu/app-v2/search-notes with a required keyword, page starting at 1, and optional search_id and search_session_id for pagination. The response example leaves outputs[0].data empty.

Step 3: open the top notes

def fetch_note(note_id: str, note_type: str) -> dict | None:
    if note_type == "video":
        items = call("xiaohongshu/app-v2/video-note-detail", {"note_id": note_id}).get("data", [])
        note = next((x for x in items if x.get("id") == note_id), None)
    else:
        items = call("xiaohongshu/app-v2/image-note-detail", {"note_id": note_id}).get("data", [])
        notes = items[0].get("note_list", []) if items else []
        note = next((x for x in notes if x.get("id") == note_id), None)
    if not note:
        return None  # deleted, private, or the id did not come back
    return {
        "note_id": note_id,
        "type": note.get("type"),
        "title": note.get("title"),
        "text": (note.get("desc") or "")[:1500],
        "tags": [t.get("name") for t in note.get("hash_tag", []) if t.get("name")],
        "likes": note.get("liked_count"),
        "collects": note.get("collected_count"),
        "comments": note.get("comments_count"),
        "shares": note.get("shared_count"),
        "posted_at": note.get("time"),
        "video_seconds": (note.get("video_info_v2") or {}).get("capa", {}).get("duration"),
    }

The two detail routes return different shapes, and this is the part that took the most poking. image-note-detail returned data as a list with one wrapper object, and the note sat inside its note_list. video-note-detail returned data as a flat list of video notes. The requested note came first, followed by other videos.

That second shape has a trap. When I sent an image note’s id to video-note-detail, it still answered code: 0, but with two unrelated videos and not the note I asked for (run fd0b3c04-618a-4a9b-8098-153badd96660). The next(...) match on id turns that into a clean None instead of silently returning the wrong note. Route by the card’s type, and always match the id.

Here is a trimmed excerpt from a public note by La Roche-Posay’s brand account about a spokesperson merchandise giveaway (run f90a6662-0eb4-49b5-a76b-88ffe488d085, read again with identical counts in 640fbaba-3390-4935-a269-08ba2162937c):

{
  "code": 0,
  "success": true,
  "data": [
    {
      "model_type": "note",
      "note_list": [
        {
          "id": "6a749cd70000000005021fd8",
          "type": "normal",
          "title": "莎莎周边已就位!晒单抽亲签💙",
          "desc": "夏日养肤,谁还没用这套!\n维稳抗应激、深层修护、清洁净肤……",
          "liked_count": 4297,
          "collected_count": 136,
          "comments_count": 193,
          "shared_count": 86,
          "time": 1786071657,
          "ip_location": "Shanghai",
          "view_count": 0,
          "hash_tag": [
            {"name": "理肤泉B5面膜PRO", "type": "topic"},
            {"name": "理肤泉超级B5精华", "type": "topic"}
          ],
          "user": {"nickname": "理肤泉larocheposay", "red_official_verified": false}
        }
      ]
    }
  ]
}

A few observations from this note and the others I opened. Hashtags appear twice: inline in desc as #name[话题]# and as structured hash_tag entries. Use the structured list. view_count came back 0 here, so don’t read it as a view metric. time is a Unix timestamp in seconds. And red_official_verified was false for this brand account in both search and detail, so it can’t tell you which notes come from brands. If you need that split, keep your own list of brand account ids.

image-note-detail also accepted a video note’s id. It returned the right note with type: "video" and its counts, just without the video_info_v2 block. That makes it a reasonable single fallback when you only need text and engagement.

SandBase API reference for the Xiaohongshu app-v2 image-note-detail endpoint The xiaohongshu/app-v2/image-note-detail reference: optional note_id and share_text (a share link). Its generated cURL example sends only a model field, so add a note_id yourself.

Step 4: sample comments

def sample_comments(note_id: str, max_pages: int = 2, sort: str = "like_count") -> list[dict]:
    params, out = {"note_id": note_id, "sort_strategy": sort}, []
    for _ in range(max_pages):
        data = call("xiaohongshu/app-v2/note-comments", params).get("data", {})
        for c in data.get("comments", []):
            if c.get("content"):
                out.append({"text": c["content"], "likes": c.get("like_count"),
                            "replies": c.get("sub_comment_count")})
        if not data.get("has_more") or not data.get("cursor"):
            break
        params = {"note_id": note_id, "sort_strategy": sort, "cursor": data["cursor"]}
    return out

The schema default for sort_strategy is latest_v2. The response listed three options in all_sort_strategies: default, latest_v2, and like_count. For research, like_count surfaces the comments other readers agreed with.

Pagination took one experiment. The reference lists cursor, index, and pageArea as request fields to copy from the previous response. In practice, the response’s cursor was itself a JSON string holding cursor, index, and pageArea. I tried both ways on the brand note. Sending the whole string back as cursor (run dbfada10-3360-4a16-a0fd-e0cd1b2db866) and sending the three parsed fields separately (run a01020cc-24ea-4d94-aea9-64bfa31b921e) returned the same next 10 comments. The function uses the simpler form.

Each comment carried content, like_count, sub_comment_count, time, and ip_location, plus commenter identity under user, all observed-only. The function keeps only text, likes, and reply counts, which is enough for theme analysis and keeps personal data out of your store. With like_count sorting on the brand note, the first comment had 0 likes but 18 replies, and the next ones had 97, 41, and 30 likes. So the order isn’t strictly by likes. Sort again on your side if it matters.

SandBase API reference for the Xiaohongshu app-v2 note-comments endpoint The xiaohongshu/app-v2/note-comments reference: optional cursor, index, note_id, pageArea, share_text, and sort_strategy (default latest_v2).

Putting it together: one category report

from collections import Counter
from statistics import median

def research_topic(keyword: str, top_n: int = 5) -> dict:
    notes = search_notes(keyword, pages=2)
    unique = list({n["note_id"]: n for n in notes if n["note_id"]}.values())
    ranked = sorted(unique, key=lambda n: n["likes"] + n["collects"], reverse=True)
    top = []
    for n in ranked[:top_n]:
        detail = fetch_note(n["note_id"], n["type"])
        if detail is None:
            continue
        detail["comment_sample"] = sample_comments(n["note_id"], max_pages=1)
        top.append(detail)
    return {
        "keyword": keyword,
        "searched": len(ranked),
        "video_share": round(sum(n["type"] == "video" for n in ranked) / max(len(ranked), 1), 2),
        "median_likes": median([n["likes"] for n in ranked]) if ranked else 0,
        "top_notes": top,
        "top_tags": Counter(t for d in top for t in d["tags"]).most_common(10),
    }

I ran research_topic("防晒霜") on 2026-10-01 UTC. Two search pages (be1bcdab-01f0-41c8-ba6c-7d64f74ab394, 1ad8a233-75c2-4054-afa0-14e8396eb072) gave 30 unique notes. Of those, 23% were videos, and the median note had 84 likes. Three of the top five were videos.

The engagement mix was the interesting part. The top note, an untitled video tagged comedy, had about 156,000 likes but only about 6,700 collects. The second, a sunscreen round-up video, had about 22,000 likes and 25,500 collects. The brand giveaway note above sat at 4,297 likes and 136 collects. Likes reward entertainment and giveaways; collects look closer to “I’ll come back to this before I buy”. If your goal is a content brief for a product category, rank by collects or by the collect-to-like ratio, not by likes. That’s one keyword on one day, so treat it as a hypothesis to check on your own categories, not a rule.

The tag counter is the other quick signal. In this run the tags for sunscreen, sunscreen for sensitive skin, and commuter sunscreen led. Hand the whole report to a model and ask it to describe the formats, hooks, and objections in the top notes and their comments.

Cost-wise, each run is two search calls plus two calls per top note. At five top notes that’s 12 calls, or 13 with the suggest step.

Documented vs. observed

ItemStatus
POST /v1/api/xiaohongshu/app-v2/search-notes with keyword, page, search_id, search_session_idDocumented in the reference
POST /v1/api/xiaohongshu/app-v2/image-note-detail and video-note-detail with note_id or share_textDocumented in the reference
POST /v1/api/xiaohongshu/app-v2/note-comments with note_id, cursor, sort_strategyDocumented in the reference
Envelope id / status / model / outputs[0].dataDocumented in the reference
Upstream wrapper code / msg / success / dataObserved only
Search items[].model_type, note.liked_count, collected_count, next_pageObserved only
Image detail data[0].note_list[]; video detail flat data[] listObserved only
hash_tag[].name, desc, time, shared_count, video_info_v2.capa.durationObserved only
Comments comments[], has_more, JSON-string cursor, all_sort_strategiesObserved only

Common use cases

Content briefs for a product category

Search the category and its suggested variants, then rank by collects. Hand the top notes’ titles, tags, and text to a model and ask for the patterns. Input: a seed keyword. Output: a brief with formats, hooks, and tags. Endpoints: search-suggest, search-notes, image-note-detail.

Brand vs. category benchmarking

Compare your brand’s notes against the category median for likes, collects, and comments. Keep your own list of brand account ids, since the verified flag didn’t help. Input: note ids. Output: a benchmark table. Endpoints: image-note-detail, video-note-detail.

Objection and question mining

Sample the most-liked comments on top notes and cluster them by theme: texture, whitening cast, price, where to buy. Input: top note ids. Output: a ranked list of buyer questions. Endpoint: note-comments.

Format tracking over time

Run the same keyword weekly and store video share, median likes, and top tags, so you can see when a format starts winning. Input: keywords. Output: a time series. Endpoint: search-notes.

Practical notes

  • Route detail calls by type. normal goes to image-note-detail and video to video-note-detail. Match on id either way.
  • Filter on model_type. Search pages can include ads cards.
  • Send back what the response gives you. next_page, search_id, and search_session_id for search, and the cursor string for comments.
  • Check success, not code. The success code differed between app-v2 and web-v3 in my runs.
  • Don’t trust view_count or red_official_verified for this research. One read 0 and the other was false for a brand account.
  • Expect drift. Search results changed between runs, so dedupe and store snapshots.
  • One route I couldn’t use. xiaohongshu/web-v3/note-detail needs a note_id plus an xsec_token. With the token from the app-v2 search result, three attempts all came back HTTP 503 wrapping an upstream 400. Stick to the app-v2 detail routes.
  • Public, read-only data only. No posting, and no private or follower-only content.

FAQ

Do I need a Xiaohongshu account or login? No. You authenticate to SandBase with SANDBASE_API_KEY. These read endpoints need no Xiaohongshu account or OAuth on your side.

Is it free? The note endpoints are currently listed as Free in the SandBase catalog. Check the catalog for the current status.

How do I get more search results? Send page from the previous response’s next_page, along with the search_id and search_session_id it returned.

Can I open a note from a share link? The detail and comment references also accept share_text, a Xiaohongshu share link. I only tested with note_id in this tutorial.

How is this different from the product research workflow? That one works on products: search-products, product-detail, and product reviews keyed by sku_id. This one works on notes, the posts people write about products.

Wrap up

A keyword is enough to see what performs in a Xiaohongshu category. Search notes and page with the returned ids. Open the top notes through the detail route that matches their type, and sample their most-liked comments. Then compare collects to likes before you decide what “performs” means. For the other Xiaohongshu endpoints, see the Xiaohongshu public data API hub. When you’re ready: