Blog/Developer Tools/

Instagram Hashtag Research API Tutorial | SandBase

Research Instagram hashtags: find related tags, page top posts, and rank them by engagement with three SandBase endpoints. No Instagram login; one SandBase API key.

Dark cinematic render of a glass hash symbol emitting photo tiles that sort into a rising stack and feed an agent core

Hashtag research on Instagram usually starts with a seed word and a vague question. Which tags around “vegan recipes” actually carry volume? What kind of post performs under the big tag right now? Which smaller tags keep showing up next to the winners? This Instagram Hashtag Research API tutorial answers those three questions with three SandBase endpoints and about 150 lines of Python. It builds on the Instagram public data API hub; read that first for the full endpoint map.

If you want to research a specific account rather than a tag, the Instagram profile research tutorial covers that path. This one stays on hashtags.

Everything here is public, read-only data. You need no Instagram login and no SDK, but you authenticate with a SandBase API key. The endpoints used here are currently listed as Free in the SandBase catalog.

The endpoint API reference is the source of truth for parameters and the response envelope, and it guarantees only that envelope. The payload field names below come from calls I ran (tested on 2026-10-01, UTC). Treat them as illustrative and observed-only, not documented guarantees, and confirm them against a live response.

Key takeaway

  • instagram/v2/search-hashtags (keyword) returns related tags with a media_count, which is enough to pick a shortlist.
  • instagram/v2/hashtag-posts (keyword, feed_type) returns posts with like, comment, and play counts. Page it by sending back the returned pagination_token.
  • Rank by likes plus comments, and keep posts with hidden likes flagged instead of treating them as zero.
  • instagram/v2/post-info (code_or_url) re-reads one shortlisted post. Its counts sit under metrics, not at the top level.

Why the v2 hashtag routes, not v3

The hub points to instagram/v3/hashtag-posts for tag feeds, so I tried that first. An earlier test in this series had returned nothing from it. This time it worked: run e0a2628b-a7b9-4dc7-b678-a6a83989af9b returned 27 posts, more_available: true, and a next_max_id cursor. But every post in that page was from the last few hours, and most had zero to eight likes. That’s a recency feed. Fine for monitoring, useless for “what performs under this tag”.

instagram/v1/hashtag-posts had the same problem plus a GraphQL-style shape (data.hashtag.edge_hashtag_to_media.edges[].node) and only an owner id per post, no username.

instagram/v2/hashtag-posts takes a feed_type of top, recent, or reels, with top as the default. The top feed (run ef34d367-313d-4470-a4d9-505e238103cc) returned 24 posts spanning roughly three months, with like counts into the thousands and play counts on Reels. Each item also carries the poster’s username and verification flag and the caption’s hashtags, already parsed. That’s the shape research needs, so this tutorial uses the v2 family throughout.

Your needUse
Public hashtag search, tag feeds, and post metrics for researchSandBase Instagram public-data API
Publish, manage your own account, or read your own insightsInstagram’s official Graph API with your business account
Private accounts or data behind a loginNeither public workflow

The workflow at a glance

  1. Expand the seed with instagram/v2/search-hashtags (keyword) and keep tags above a volume floor.
  2. Read the top posts under a tag with instagram/v2/hashtag-posts (keyword, feed_type: "top").
  3. Page by sending back pagination_token until you have enough posts or the token disappears.
  4. Rank by engagement and count which other hashtags co-occur on those posts.
  5. Re-check a shortlisted post with instagram/v2/post-info (code_or_url).

SandBase Instagram API catalog page listing Instagram endpoints with Available and Free status The Instagram catalog on SandBase. It shows 81 endpoints as GET /apis/v1/instagram/... paths, with the selected endpoint marked Available, Free. This tutorial calls the Model API POST /v1/api/instagram/... routes instead.

A note on surfaces before the code. The catalog page lists each operation as GET /apis/v1/instagram/<path>. This tutorial uses the Model API surface from the endpoint reference: POST /v1/api/instagram/<path> with a JSON body that holds only the endpoint’s own parameters. Keep the POST method and the /v1/api/ prefix when you copy the examples.

Step 0: one helper for every call

import os
import time
from datetime import datetime, timezone

import requests

API = "https://api.sandbase.ai/v1/api"
HEADERS = {
    "Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}",
    "Content-Type": "application/json",
}


def call(path: str, payload: dict, retries: int = 2) -> dict:
    for attempt in range(retries + 1):
        try:
            resp = requests.post(f"{API}/{path}", headers=HEADERS, json=payload, timeout=90)
            resp.raise_for_status()
            break
        except requests.RequestException:
            if attempt == retries:
                raise
            time.sleep(2 * (attempt + 1))
    body = resp.json()
    if body.get("status") != "completed":
        raise RuntimeError(body.get("error", {}).get("message", f"{path} did not complete"))
    # Documented shape is outputs[0].data; also accept a top-level `output`.
    output = body.get("output")
    if output is None and body.get("outputs"):
        output = body["outputs"][0].get("data", {})
    print(f"  {path} run id: {body.get('id')}")
    return output or {}

The reference documents completed responses as outputs[0].data, and every Instagram call in this test came back in that shape. The helper still accepts a top-level output, because I have seen that shape on other SandBase platform endpoints and the two can vary between calls. Inside outputs[0].data, the v2 routes wrapped their payload in another data key. That’s why the steps below read .get("data") again. The helper prints each run id, which is handy when you need to compare a strange result against the raw response later.

Step 1: expand the seed into candidate hashtags

def find_hashtags(seed: str, min_posts: int = 1000) -> list[dict]:
    out = call("instagram/v2/search-hashtags", {"keyword": seed})
    items = (out.get("data") or {}).get("items", []) or []
    tags = [
        {"name": t.get("name"), "media_count": t.get("media_count") or 0}
        for t in items
        if (t.get("media_count") or 0) >= min_posts
    ]
    return sorted(tags, key=lambda t: t["media_count"], reverse=True)

With the seed veganrecipes, the search (run 0ae668b6-3580-498d-ae55-55e34a7fb84c) returned 55 tags. Here’s a trimmed excerpt of outputs[0].data:

{
  "data": {
    "count": 55,
    "items": [
      {"id": "17843646523025713", "name": "veganrecipeshare", "media_count": 332329, "allow_following": false},
      {"id": "17842274242068426", "name": "veganrecipes", "media_count": 9249683, "allow_following": false},
      {"id": "17842769464147759", "name": "veganrecipesforhealth", "media_count": 7781, "allow_following": false}
    ]
  }
}

The list is not sorted by volume, so sort it yourself. A floor of 1,000 posts drops the long tail of near-empty spelling variants. In the full run below, the top of the list was veganrecipes (about 9.2 million posts), then veganrecipeshare, easyveganrecipes, alkalineveganrecipes, and rawveganrecipes, each above 100,000.

I also tried instagram/v3/search-hashtags, which takes query instead of keyword. It returned 20 tags with the same ids and counts, and a rank_token field for paging. In my run that rank_token was null, so there was no second page to fetch. The v2 route gave me more tags in one call, so I kept it.

SandBase API reference for the Instagram v2 search-hashtags endpoint The instagram/v2/search-hashtags reference: POST /v1/api/instagram/v2/search-hashtags with one required keyword field. The documented response example leaves outputs[0].data empty.

Step 2: read and page the top posts under a tag

def to_utc(value) -> str | None:
    # The top feed returned ISO strings; the recent feed returned Unix seconds.
    if isinstance(value, (int, float)):
        return datetime.fromtimestamp(value, tz=timezone.utc).isoformat()
    return value


def hashtag_posts(tag: str, feed_type: str = "top", max_pages: int = 2) -> list[dict]:
    posts, token = [], None
    for _ in range(max_pages):
        payload = {"keyword": tag, "feed_type": feed_type}
        if token:
            payload["pagination_token"] = token
        out = call("instagram/v2/hashtag-posts", payload)
        for item in (out.get("data") or {}).get("items", []) or []:
            user = item.get("user") or {}
            posts.append({
                "code": item.get("code"),
                "url": f"https://www.instagram.com/p/{item.get('code')}/",
                "account": user.get("username"),
                "verified": user.get("is_verified"),
                "media_type": item.get("media_type"),
                "likes": item.get("like_count"),  # None when likes are hidden
                "comments": item.get("comment_count") or 0,
                "plays": item.get("play_count") or 0,
                "taken_at": to_utc(item.get("taken_at")),
                "hashtags": item.get("caption_hashtags") or [],
                "paid_partnership": item.get("is_paid_partnership"),
            })
        token = out.get("pagination_token")
        if not token:
            break
    return posts

The reference documents keyword, feed_type, and pagination_token. The live responses agreed on where the token lives: it came back at the top of outputs[0].data, next to the inner data object, not inside it. Sending it back as pagination_token worked. Page one of the top feed returned 24 posts; page two (run 4c96135e-d0c9-4642-8979-bd47599fe0be) returned 30 more with no overlap and a fresh token. I stopped at two pages. I didn’t test how deep the top feed goes.

Here’s one post item from the top feed, trimmed. It’s a Reel from a verified recipe account; I’ve left the account name out and kept the counts:

{
  "code": "DcZlsXtzS-X",
  "media_type": 2,
  "product_type": "clips",
  "like_count": 5719,
  "comment_count": 1763,
  "play_count": 661027,
  "taken_at": "2026-08-23T22:44:44Z",
  "taken_at_ts": 1787525084,
  "is_paid_partnership": false,
  "like_and_view_counts_disabled": false,
  "caption_hashtags": ["#pineapplechutney", "#spicychutney", "#easyrecipes", "#veganrecipes"]
}

Three things tripped me up here. First, like_count was null on two of the 24 posts. Both had like_and_view_counts_disabled: true, which means the creator hid their likes. Don’t read that as zero. Second, taken_at was an ISO string in the top feed but a Unix timestamp in the recent feed (run a249ddfd-f55a-4f94-809f-f4cf42482341), which is why to_utc handles both. Third, media_type came back as 1 (photo), 2 (video or Reel), or 8 (carousel). Carousels showed play_count: 0, so plays only make sense for comparing videos with other videos.

SandBase API reference for the Instagram v2 hashtag-posts endpoint The instagram/v2/hashtag-posts reference: a required keyword (without #), an optional feed_type that defaults to top, and an optional pagination_token from the previous response.

Step 3: rank by engagement and find co-occurring tags

def rank_by_engagement(posts: list[dict]) -> list[dict]:
    seen, ranked = set(), []
    for p in posts:
        if p["code"] in seen:
            continue
        seen.add(p["code"])
        p["likes_hidden"] = p["likes"] is None
        # hidden likes: no score, sorted after posts with full counts
        p["engagement"] = None if p["likes_hidden"] else p["likes"] + p["comments"]
        ranked.append(p)
    return sorted(ranked, key=lambda p: (p["engagement"] is not None, p["engagement"] or 0), reverse=True)


def co_hashtags(posts: list[dict], tag: str, top: int = 10) -> list[tuple[str, int]]:
    counts: dict[str, int] = {}
    for p in posts:
        for h in {h.lower().lstrip("#") for h in p["hashtags"]}:
            if h != tag.lower():
                counts[h] = counts.get(h, 0) + 1
    return sorted(counts.items(), key=lambda kv: kv[1], reverse=True)[:top]

“Engagement” here means likes plus comments, as absolute numbers. The hashtag feed has no follower counts for the posters, so you can’t compute an engagement rate from this endpoint alone. If you need a rate, look up each shortlisted account with the profile workflow from the other tutorial, and accept the extra calls.

Posts with hidden likes get no engagement score. They sort after every post with full counts, and the likes_hidden flag tells you why, so a partial number never competes with a complete one. The tag counts come from caption_hashtags, which arrived with the # prefix and mixed case (#DairyFree and #dairyFree on the same post), so the helper lowercases and de-duplicates per post before counting.

Step 4: re-check one shortlisted post

def refresh_post(code: str) -> dict:
    data = call("instagram/v2/post-info", {"code_or_url": code}).get("data", {}) or {}
    metrics = data.get("metrics") or {}
    caption = data.get("caption") or {}
    return {
        "code": data.get("code"),
        "account": (data.get("user") or {}).get("username"),
        "product_type": data.get("product_type"),
        "taken_at": data.get("taken_at_date"),
        "likes": metrics.get("like_count"),
        "comments": metrics.get("comment_count"),
        "hashtags": caption.get("hashtags") or [],
    }

This is the step where my first draft read the wrong field. In the hashtag feed, like_count sits at the top of each item. In post-info, the item has no top-level like_count at all. The counts live in a metrics object. I tested it on a carousel from the Instant Pot brand account that showed up under #veganrecipes in the recent feed (run f714c418-01ef-4074-9cf1-94f6241169d8). Trimmed:

{
  "data": {
    "code": "Dd7t9arlwZC",
    "product_type": "carousel_container",
    "carousel_media_count": 3,
    "taken_at_date": "2026-10-01T01:22:00+00:00",
    "metrics": {"like_count": 4, "comment_count": 0, "play_count": null, "share_count": null},
    "caption": {
      "text": "Yes, ice cream can be vegan. We tested it - and it works perfectly. ...",
      "hashtags": ["#instantpot", "#instantchill", "#veganicecream", "#vegan", "#veganrecipes"]
    },
    "user": {"username": "instantpotofficial", "full_name": "Instant Pot®", "is_verified": true}
  }
}

The post was a few hours old, which explains the small numbers. Re-reading a post later is the point of this step: run it a day or a week after the first pass to see whether a shortlisted post kept growing. I also tried instagram/v3/post-info-by-code on the same shortcode (run bc654757-6209-44de-9ff6-de74b144316c). It returned an items list in a different shape, so I stayed with v2 for consistency.

SandBase API reference for the Instagram v2 post-info endpoint The instagram/v2/post-info reference: one required code_or_url field that takes a post shortcode or a full post URL.

Putting it together

if __name__ == "__main__":
    seed = "veganrecipes"
    print("1) related hashtags")
    tags = find_hashtags(seed)
    for t in tags[:8]:
        print(f"   #{t['name']:<28} {t['media_count']:>10,}")

    print("2) top posts under the seed tag, two pages")
    posts = hashtag_posts(seed, "top", max_pages=2)
    print(f"   collected {len(posts)} posts")

    print("3) ranked by likes + comments")
    ranked = rank_by_engagement(posts)
    for p in ranked[:5]:
        flag = " (likes hidden)" if p["likes_hidden"] else ""
        print(f"   {p['code']}  eng={str(p['engagement']):>6}  plays={p['plays']:>7}  type={p['media_type']}{flag}")
    print("   co-occurring tags:", co_hashtags(posts, seed, 8))

    print("4) refresh one post")
    print("  ", refresh_post("Dd7t9arlwZC"))

I ran the whole file as written. The search call was run 4323d2d9-d773-447b-adaa-297cc5982320; the two hashtag pages were 45830a7e-d588-4674-b629-5e2a5b5d9309 and 27d391bb-1aeb-4af4-a9fa-a88a9bcc1bcd; the refresh was f714c418-01ef-4074-9cf1-94f6241169d8. Two pages gave 54 posts. Four of the top five by engagement were Reels (media_type 2), and the top post had about 20,000 likes plus comments on roughly 500,000 plays. The most frequent co-occurring tags were vegan (11 of 54 posts), plantbased (10), plantbasedrecipes (6), and easyrecipes (6).

That’s one tag, one afternoon, two pages. It tells you what the top feed looked like for that tag at that moment, not what Instagram ranks in general. Repeat the run for each shortlisted tag from step 1 and you get a small comparison table an agent can summarize: which tags have Reel-heavy top feeds, which co-tags recur, and which posts are worth a closer look.

Each tag costs one call per page, plus one search call per seed and one call per refreshed post. Cache posts by code so a re-run doesn’t duplicate them.

Documented vs. observed

ItemStatus
POST /v1/api/instagram/v2/search-hashtags with keywordDocumented in the reference
POST /v1/api/instagram/v2/hashtag-posts with keyword, feed_type, pagination_tokenDocumented in the reference
POST /v1/api/instagram/v2/post-info with code_or_urlDocumented in the reference
Envelope id / status / model / outputs[0].dataDocumented in the reference
Search data.items[] with name, id, media_countObserved only
pagination_token at the top of outputs[0].dataObserved only
Post code, like_count, comment_count, play_count, media_type, caption_hashtags, userObserved only
taken_at as ISO in top, Unix seconds in recentObserved only
post-info counts under metrics, caption.hashtags, taken_at_dateObserved only

Common use cases

Niche tag shortlisting

Start from a seed word, keep the tags above a volume floor, and compare their top feeds. Input: a seed keyword. Output: a ranked list of candidate tags with post volume. Endpoint: v2/search-hashtags.

Content format research

Count media_type in the top feed of each tag to see whether Reels, carousels, or photos dominate. Input: tag names. Output: a format mix per tag. Endpoint: v2/hashtag-posts.

Co-tag discovery for campaigns

Tally the hashtags that co-occur on top posts to find adjacent tags worth testing. Input: a tag. Output: a frequency list. Endpoint: v2/hashtag-posts.

Creator and brand discovery

List the accounts behind the highest-engagement posts, then hand them to a profile lookup for follower counts. Input: ranked posts. Output: a shortlist of accounts to review. Endpoints: v2/hashtag-posts, then the profile workflow.

Practical notes

  • Param names differ by version. v2 search takes keyword, v3 search takes query; v2 tag posts take keyword, v3 takes tag, v1 takes hashtag.
  • Pick the feed on purpose. top for “what performs”, recent for monitoring. The v3 tag route behaved like a recency feed in my test.
  • Hidden likes are null, not zero. Check like_and_view_counts_disabled and flag those posts.
  • Normalize timestamps and hashtags. ISO vs. Unix by feed type; #-prefixed, mixed-case tags.
  • Counts move in post-info. Read metrics.like_count, not a top-level field.
  • Keep personal data out of your store. Posts under a tag include private individuals’ public accounts. Store post codes and counts; keep usernames only for business or brand accounts you plan to contact.
  • Business fields are observed-only. Confirm them against a live response before you depend on them.

FAQ

Do I need an Instagram account or login? No. You authenticate to SandBase with SANDBASE_API_KEY. These read endpoints need no Instagram account or OAuth on your side.

Is it free? The endpoints used here are currently listed as Free in the SandBase catalog. Check the catalog for the current status.

Can I get an engagement rate? Not from the hashtag feed alone. It returns counts per post but no follower counts. Look up shortlisted accounts separately and divide.

How do I get more posts under a tag? Send the pagination_token from the previous response. Stop when it’s missing.

Should I include the # in the keyword? No. The reference asks for the tag without #, and that’s how I called it.

Wrap up

A seed word is enough to research a hashtag niche on Instagram. One call expands it into tags with volume, a few paged calls show what performs under a tag, and a final call re-checks the posts you care about. Rank by likes plus comments, flag hidden likes, and count co-tags, and an agent has something concrete to summarize. For the rest of the Instagram endpoints, see the Instagram public data API hub. When you’re ready: