Blog/教程/

DataForSEO 关键词调研 Agent:用 SandBase 实测搭建

用 Python 搭一个 DataForSEO 关键词调研 Agent:拿关键词搜索量和难度、实时 SERP、竞品已排名的词,数字在代码里算,最后由 LLM 写计划。全部通过 SandBase 一个 Key 调用。

DataForSEO 关键词调研 Agent 封面;示意插图,不是测试结果

种子词填 “ai agent”,DataForSEO 一口气返回了 7 个月搜索量都是 49,500 的词:“ai agent”、“agent ai”、“agent in ai”、“ai intelligent agent”,还有另外三个。再一看,它们过去 12 个月的逐月数据一模一样。要是直接把这几行加起来,你会得到 346,500 次根本不存在的搜索。

这是做 DataForSEO 关键词调研 Agent 时碰到的第一个坑,也正好说明一件事:数字得在代码里算,别交给模型。本文用大约 170 行 Python 搭一个这样的 Agent。输入一个种子主题和一个竞品域名,它会输出一份排好优先级的关键词表,包含搜索量、难度、前三个关键词当前的自然搜索结果,以及竞品已经在排名的词。最后由 openai/gpt-6.1-sol 根据算好的数字写一份简短计划。所有调用都走 SandBase,一个 API Key 搞定。

先说结论

  • 三个 DataForSEO 端点就够用:keyword_suggestions(关键词、搜索量、难度、意图)、ranked_keywords(竞品排了哪些词)、serp/google/organic/live/advanced(实时自然结果)。截至 2026-10-03,三个在 SandBase 上都是启用状态。
  • 完整跑一次是 5 次数据调用加 1 次 LLM 调用。2026-10-03 数据调用的计费是 $0.00(模型页标 Free,站点正在搞 “API Free Week”),LLM 那次花了 $0.0104。
  • 近义变体共用同一个搜索量。按 12 个月搜索量序列分组后,50 条建议被合并成更少、更真实的行,光 “ai agent” 一组就吸收了 7 个变体。
  • 实时 SERP 变得很快:相隔约 6 分钟的两次查询,“ai agent” 的前五个自然结果域名一个都不重合。一次抓取只能当样本看。
  • 这些是 DataForSEO 提供的第三方 SEO 估算数据,不是 Google 自己的数据。

Agent 产出什么

产出来源SandBase 价格(2026-10-03)
含种子词的关键词,带 search_volume、keyword_difficulty、main_intentdataforseo_labs/google/keyword_suggestions/liveFree(模型页)
竞品已排名的词,以及排名位置和 URLdataforseo_labs/google/ranked_keywords/liveFree(模型页)
前 3 个关键词的自然结果域名serp/google/organic/live/advancedFree(模型页)
变体分组、机会分、合并、导出 CSVPython—
不超过 8 条的计划openai/gpt-6.1-sol输入 $2/M、输出 $10/M token

最终输出是 keyword_plan.csv,外加终端里打印的表格和计划。

SandBase 上的 DataForSEO 目录页截至 2026-10-03 列出 74 个端点。选中的卡片(一个 OnPage 端点,本文没用到)显示 Available 和 Free:

SandBase DataForSEO API 目录页,显示 74 个端点,其中包括 SERP Google Organic Live,选中卡片标为 Available 和 Free 截图:SandBase 的 DataForSEO 目录页列出 74 个端点,包括 SERP Google Organic Live,页面顶部挂着 “API Free Week” 横幅(2026-10-03 截取)。

目录页展示的是 GET /apis/v1/... 这一套接口。本文用的是端点参考页里的 Model API 路由:POST https://api.sandbase.ai/v1/api/dataforseo/<path>。

数据边界,以及直连 DataForSEO 和走 SandBase 的区别

这里用的是公开、只读的 SEO 数据。关键词搜索量、难度分和 SERP 快照,都是上游 DataForSEO 给出的第三方估算,不是 Google 自己的数据,SandBase 和 Google 也没有任何关联。Agent 需要一个 SandBase API Key。它不碰 Google Search Console、Google Ads 账户、站长后台数据,也不会对任何网站或账户做改动。

你也可以直接调 DataForSEO,区别主要在接入方式:

直连 DataForSEO通过 SandBase 调 DataForSEO
鉴权用 DataForSEO 的 API login 和 password 做 HTTP Basic(鉴权文档)Authorization: Bearer $SANDBASE_API_KEY
请求体任务对象组成的 JSON 数组单个任务对象,字段相同
返回DataForSEO 的 tasks[].result[]同样的 DataForSEO 返回体,放在 SandBase 的 outputs[0] 里
LLM 那一步另一家供应商、另一把 Key、另一张账单同一把 Key,POST /v1/chat/completions
适合只用 DataForSEO、量大、想自己签合同想把 SEO 数据和模型放在一把 Key、一份合同后面

两边的价格我没有做对比,DataForSEO 自己公开了定价。

第 1 步:拿关键词、搜索量和难度

SandBase 的 DataForSEO 参考页没有列请求字段,模型页上写的是 “0 input fields”,请求体会原样转给 DataForSEO。所以参数以 DataForSEO 的 keyword_suggestions 文档为准:keyword(必填)、location_code、language_code、limit、filters、order_by、include_seed_keyword。

SandBase 上 DataForSEO Labs Google Keyword Suggestions 的模型页,显示基础价格 Free、同步执行、api 类型、0 个输入字段 截图:keyword_suggestions/live 模型页显示基础价格 Free、同步执行、0 个输入字段,所以请求字段要看 DataForSEO 自己的文档(2026-10-03 截取)。

实测:2026-10-03(UTC)

公开输入:通用种子词 “ai agent”,Google 美国(location_code 2840),英文,竞品域名用 Microsoft Learn(learn.microsoft.com)。

curl -s https://api.sandbase.ai/v1/api/dataforseo/v3/dataforseo_labs/google/keyword_suggestions/live \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"keyword": "ai agent", "location_code": 2840, "language_code": "en",
       "include_seed_keyword": true, "limit": 50,
       "filters": [["keyword_info.search_volume", ">=", 100]],
       "order_by": ["keyword_info.search_volume,desc"]}'

截短后的返回(run id 6ed0cb41-ad1f-4bf3-a10e-bb1c046e2223,50 条中的 1 条):

{
  "id": "6ed0cb41-ad1f-4bf3-a10e-bb1c046e2223",
  "status": "completed",
  "model": "dataforseo/v3/dataforseo_labs/google/keyword_suggestions/live",
  "outputs": [{
    "status_code": 20000,
    "tasks": [{"status_code": 20000, "result": [{
      "total_count": 615, "items_count": 50,
      "items": [{
        "keyword": "ai agent",
        "keyword_info": {"search_volume": 49500, "competition_level": "MEDIUM", "cpc": 20.9,
          "monthly_searches": [{"year": 2026, "month": 8, "search_volume": 49500}, "…"]},
        "keyword_properties": {"core_keyword": "agentic ai", "keyword_difficulty": 70},
        "search_intent_info": {"main_intent": "commercial"}
      }]
    }]}]
  }]
}

抄解析代码之前,先说个坑。端点参考页写的完成态返回是 outputs[0].data,代码也优先按这个约定读。但我 2026-10-03 调的 22 次 DataForSEO,全部是把 DataForSEO 自己的返回体直接放在 outputs[0](version、status_code、tasks),没有 data 这一层。所以函数先看 outputs[0].data 里有没有 DataForSEO 返回体,没有再退回读 outputs[0],两处都找不到 tasks 就直接报错。这个不一致我已经反馈给 SandBase 团队。

上面这些业务字段(search_volume、keyword_difficulty、monthly_searches、core_keyword、main_intent)来自这几次实测和 DataForSEO 文档,属于观察到的字段,不是 SandBase 的保证。代码里一律用 .get() 读。返回里还有一个 DataForSEO 自己的 cost 字段,你的 SandBase 账户实际扣多少,以 GET /v1/tasks/<id>/cost 为准。

第 2 步:竞品已经排了哪些词

ranked_keywords 接收一个 target 域名。它的数据是 DataForSEO 的快照,官方文档说每周更新,所以这里的排名不等于当天的 Google 实时排名(后面的 SERP 那一步才是实时查的)。我只筛含种子词的关键词,按搜索量排序:

{"target": "learn.microsoft.com", "location_code": 2840, "language_code": "en", "limit": 50,
 "filters": [["keyword_data.keyword", "like", "%ai agent%"]],
 "order_by": ["keyword_data.keyword_info.search_volume,desc"]}

run id d38178c1-e21d-4610-9763-02c222363b14 报告 total_count 为 254,返回了 50 条。截一条:

{"keyword_data": {"keyword": "ai agent course",
   "keyword_info": {"search_volume": 1000},
   "keyword_properties": {"keyword_difficulty": 12}},
 "ranked_serp_element": {"serp_item": {"type": "organic", "rank_group": 2, "rank_absolute": 3,
   "domain": "learn.microsoft.com",
   "url": "https://learn.microsoft.com/en-us/shows/ai-agents-for-beginners/"}}}

keyword_data 的字段名和第 1 步的关键词条目一致,所以一个 row_from() 就能把两边都展平。rank_group 是在自然结果里的名次,rank_absolute 则把页面上所有元素都算进去。

SandBase 上 dataforseo ranked_keywords live 的 API 参考页,显示 POST 路由和带 outputs data 对象的完成态示例 截图:ranked_keywords/live 参考页显示 POST /v1/api/dataforseo/v3/dataforseo_labs/google/ranked_keywords/live,完成态示例是 outputs[0].data,和实测返回的结构不一样(2026-10-03 截取)。

参考页自动生成的 curl 示例在请求体里放了 model 字段。路径里已经写明了模型,所以代码只发 DataForSEO 的字段。

定下这一步之前,我先试了个看起来最省事的办法:把之前一次测试拿到的 30 条建议词用 in 过滤丢给 ranked_keywords(run id 75cad2ec-fd7d-4eda-ace5-8274c295eff0)。结果 Microsoft Learn 只命中 1 个(“no code ai agent builder”,第 58 名)。它真正有排名的 “build ai agents”、“ai agent course” 这类词,压根不在建议列表里。做竞品差距分析,得把两张表合并,而不是取交集。

第 3 步:前几个关键词的实时 SERP

SERP 端点(DataForSEO 文档)接收 keyword、location_code、language_code 和 depth。depth: 10 时,三次调用各返回 12 到 13 个元素,其中 organic 只有 6 到 8 个,第一个元素每次都是 ai_overview。所以代码按 type == "organic" 过滤。“ai voice agent” 这次(run id a6618702-02a4-47bd-8f85-e848f2309252):

[{"type": "ai_overview", "rank_group": 1, "rank_absolute": 1},
 {"type": "organic", "rank_group": 1, "rank_absolute": 2, "domain": "www.retellai.com"},
 {"type": "organic", "rank_group": 2, "rank_absolute": 3, "domain": "elevenlabs.io"},
 {"type": "organic", "rank_group": 3, "rank_absolute": 4, "domain": "www.reddit.com"}]

6 分钟后我把整个程序又跑了一遍。“ai agent” 在 04:10 UTC 那次的前五个自然结果是 IBM、aiagent.app、AWS、Wikipedia、Meta;04:16 那次变成了 Reddit、LinkedIn、retresco.de、airia.com、airbyte.com,一个都不重合。“what is an ai agent” 也完全不重合,“ai voice agent” 只剩 Reddit 是共同的。两次样本分不出这是 Google 本身的波动、AI Overview 版式,还是 DataForSEO 的采集方式。所以计划里 SERP 域名只用来判断页面类型,不当作排名目标清单。

第 4 步:在代码里打分

规则都在 build_table() 里,一共三条:

  1. 按关键词合并两张表,把竞品的 rank_group 和 URL 写到对应行上。
  2. 去掉搜索量低于 100、没有难度值、或意图是 navigational 的行(比如 “n8n ai agent” 这种品牌词)。
  3. 按完全相同的 12 个月搜索量序列,把近义变体归为一组。代表词优先选含种子词的,其次选最短的。variants 记录合并了几行。

然后 opportunity = volume × (100 − difficulty) / 100。说白了就是用 DataForSEO 的 0–100 难度给搜索量打个折,是个启发式分数,不是流量预测。表里保留机会分最高的 12 行,再补最多 5 行竞品有排名的词,免得竞品那部分被挤到表外。

为什么按搜索量序列分组,而不用 DataForSEO 的 core_keyword?这次数据里,core_keyword 把 “ai agent” 归到了 “agentic ai”,却把 “ai intelligent agent” 分到另一组,可这两个词的 49,500 和逐月数字完全一样。按序列分组更粗糙,但它针对的正是真问题:搜索量被重复计算。两个不相干的词也可能碰巧序列相同,不过在搜索量 ≥ 100、12 个月的条件下,我这次没遇到。

第 5 步:让 LLM 写计划

openai/gpt-6.1-sol 拿到的是 JSON 格式的表格和 SERP 域名,外加一条硬性要求:只引用给定数据里的数字,不要自己算新数字。选它,是因为我们的 Claude Sonnet 5.5 和 GPT-6.1 Sol Agent 对比测试里,它 30 个短工具任务全对,每个任务约 $0.0066。那次测试里出错的地方恰好都是对工具结果做算术,这也是本文把求和、分组、打分都留在 Python 里的原因。调用走 OpenAI 兼容的 Chat Completions 接口。

完整代码

#!/usr/bin/env python3
"""SEO keyword-research agent: DataForSEO data through SandBase, numbers in code, plan by an LLM.

Usage: python3 seo_research_agent.py "ai agent" learn.microsoft.com
Needs SANDBASE_API_KEY in the environment. Writes keyword_plan.csv and prints a plan.
"""
import csv
import json
import os
import sys

import requests

FIELDS = ["keyword", "variants", "volume", "difficulty", "intent", "opportunity",
          "competitor_rank", "competitor_url", "serp_top3"]
API = "https://api.sandbase.ai"
HEADERS = {
    "Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}",
    "Content-Type": "application/json",
}
LOCALE = {"location_code": 2840, "language_code": "en"}  # United States, English
TOP_N = 12            # best rows by opportunity
COMPETITOR_EXTRA = 5  # extra rows where the competitor ranks, even below the cut
SERP_CHECKS = 3       # live SERP lookups for the top keywords


def dataforseo(path: str, params: dict) -> dict:
    """POST /v1/api/dataforseo/<path> and return the first DataForSEO task result."""
    resp = requests.post(f"{API}/v1/api/dataforseo/{path}", headers=HEADERS, json=params, timeout=120)
    resp.raise_for_status()
    body = resp.json()
    outputs = body.get("outputs") or []
    if body.get("status") != "completed" or not outputs:
        raise RuntimeError(f"SandBase call did not complete: {body.get('status')} {body.get('error')}")
    # Documented contract: outputs[0].data. Observed on 2026-10-03: DataForSEO's body sat directly in outputs[0].
    data = outputs[0].get("data")
    payload = data if isinstance(data, dict) and "tasks" in data else outputs[0]
    tasks = payload.get("tasks") or []
    if not tasks or tasks[0].get("status_code") != 20000:
        task = tasks[0] if tasks else payload
        raise RuntimeError(f"DataForSEO task failed: {task.get('status_code')} {task.get('status_message')}")
    results = tasks[0].get("result") or []
    return results[0] if results else {}


def keyword_ideas(seed: str, limit: int = 50) -> list[dict]:
    """Keywords that contain the seed, with volume >= 100, biggest first."""
    result = dataforseo("v3/dataforseo_labs/google/keyword_suggestions/live", {
        "keyword": seed, **LOCALE, "include_seed_keyword": True, "limit": limit,
        "filters": [["keyword_info.search_volume", ">=", 100]],
        "order_by": ["keyword_info.search_volume,desc"],
    })
    return result.get("items") or []


def competitor_keywords(domain: str, seed: str, limit: int = 50) -> list[dict]:
    """Keywords containing the seed that the competitor domain ranks for."""
    result = dataforseo("v3/dataforseo_labs/google/ranked_keywords/live", {
        "target": domain, **LOCALE, "limit": limit,
        "filters": [["keyword_data.keyword", "like", f"%{seed}%"]],
        "order_by": ["keyword_data.keyword_info.search_volume,desc"],
    })
    return result.get("items") or []


def top_organic(keyword: str, n: int = 5) -> list[dict]:
    """Live Google organic results (first page) for one keyword."""
    result = dataforseo("v3/serp/google/organic/live/advanced", {"keyword": keyword, **LOCALE, "depth": 10})
    organic = [i for i in result.get("items") or [] if i.get("type") == "organic"]
    return [{"rank": i.get("rank_group"), "domain": i.get("domain"), "url": i.get("url")} for i in organic[:n]]


def row_from(kw: dict) -> dict:
    """Flatten the keyword fields both Labs endpoints share."""
    info = kw.get("keyword_info") or {}
    props = kw.get("keyword_properties") or {}
    # Close variants ("ai agent", "agent ai") came back with identical 12-month volume series.
    series = tuple(m.get("search_volume") for m in info.get("monthly_searches") or [])
    return {
        "keyword": kw.get("keyword"),
        "cluster": series or kw.get("keyword"),
        "volume": info.get("search_volume") or 0,
        "difficulty": props.get("keyword_difficulty"),
        "intent": (kw.get("search_intent_info") or {}).get("main_intent"),
        "competitor_rank": None,
        "competitor_url": "",
    }


def build_table(seed: str, ideas: list[dict], ranked: list[dict]) -> list[dict]:
    """Merge both sources, collapse close variants, score, and keep the top rows."""
    rows = {}
    for kw in ideas:
        rows[kw.get("keyword")] = row_from(kw)
    for item in ranked:
        kw = item.get("keyword_data") or {}
        serp_item = (item.get("ranked_serp_element") or {}).get("serp_item") or {}
        row = rows.setdefault(kw.get("keyword"), row_from(kw))
        row["competitor_rank"] = serp_item.get("rank_group")
        row["competitor_url"] = serp_item.get("url") or ""

    # One row per variant group. Representative: contains the seed, then the shortest keyword.
    keep = [r for r in rows.values()
            if r["volume"] >= 100 and r["difficulty"] is not None and r["intent"] != "navigational"]
    keep.sort(key=lambda r: (seed not in r["keyword"], len(r["keyword"])))
    best = {}
    for row in keep:
        winner = best.setdefault(row["cluster"], {**row, "variants": 0})
        winner["variants"] += 1
        if row["competitor_rank"] and not winner["competitor_rank"]:
            winner["competitor_rank"], winner["competitor_url"] = row["competitor_rank"], row["competitor_url"]

    # Opportunity = volume discounted by difficulty (0-100). A heuristic, not a traffic forecast.
    for row in best.values():
        row["opportunity"] = round(row["volume"] * (100 - row["difficulty"]) / 100)
    ranked_rows = sorted(best.values(), key=lambda r: r["opportunity"], reverse=True)
    # Top rows overall, plus the competitor's best keywords even if they fall below the cut.
    table = ranked_rows[:TOP_N]
    table += [r for r in ranked_rows if r["competitor_rank"] and r not in table][:COMPETITOR_EXTRA]
    return table


def write_plan(seed: str, domain: str, table: list[dict], serps: dict) -> tuple[str, dict]:
    """Ask openai/gpt-6.1-sol for a short plan. It gets computed numbers and must not invent new ones."""
    prompt = (
        f"Seed topic: {seed}. Competitor: {domain}. Market: Google US, English.\n"
        f"Keyword table (opportunity = volume x (100 - difficulty) / 100, computed in code):\n"
        f"{json.dumps([{k: r.get(k) for k in FIELDS} for r in table], ensure_ascii=False)}\n"
        f"Current top organic results for the top keywords:\n{json.dumps(serps, ensure_ascii=False)}\n\n"
        "Write a content plan in at most 8 bullets: which 3-5 keywords to target first and why, "
        "where the competitor already ranks, and what page type each target needs based on the SERP. "
        "Quote numbers only from the data above; do not compute or estimate new ones."
    )
    resp = requests.post(f"{API}/v1/chat/completions", headers=HEADERS, timeout=120, json={
        "model": "openai/gpt-6.1-sol",
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": 800,
    })
    resp.raise_for_status()
    body = resp.json()
    return body["choices"][0]["message"]["content"], body.get("usage") or {}


def main(seed: str, domain: str) -> None:
    ideas = keyword_ideas(seed)
    ranked = competitor_keywords(domain, seed)
    print(f"{len(ideas)} keyword suggestions, {len(ranked)} competitor keywords containing '{seed}'")
    table = build_table(seed, ideas, ranked)
    serps = {row["keyword"]: top_organic(row["keyword"]) for row in table[:SERP_CHECKS]}
    for row in table[:SERP_CHECKS]:
        row["serp_top3"] = " | ".join(r["domain"] for r in serps[row["keyword"]][:3])

    with open("keyword_plan.csv", "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=FIELDS, extrasaction="ignore")
        writer.writeheader()
        writer.writerows(table)

    print(f"\n{'keyword':28} {'var':>3} {'volume':>7} {'KD':>3} {'opp':>6} {'comp':>4}  intent")
    for row in table:
        rank = row["competitor_rank"] or "-"
        print(f"{row['keyword'][:28]:28} {row['variants']:>3} {row['volume']:>7} {row['difficulty']:>3} "
              f"{row['opportunity']:>6} {rank:>4}  {row['intent']}")

    plan, usage = write_plan(seed, domain, table, serps)
    print("\nPlan (openai/gpt-6.1-sol):\n" + plan)
    print(f"\nLLM usage: {usage.get('prompt_tokens')} input / {usage.get('completion_tokens')} output tokens")


if __name__ == "__main__":
    if len(sys.argv) != 3:
        sys.exit('usage: python3 seo_research_agent.py "<seed topic>" <competitor domain>')
    main(sys.argv[1], sys.argv[2])

需要 Python 3.9 以上和 requests。运行:python3 seo_research_agent.py "ai agent" learn.microsoft.com。

实际跑出来的结果

下面是上面这段程序在 2026-10-03 04:15 UTC 的输出(6 次调用全部 completed):

50 keyword suggestions, 50 competitor keywords containing 'ai agent'

keyword                      var  volume  KD    opp comp  intent
ai agent                       7   49500  70  14850    -  commercial
ai voice agent                 1   22200  38  13764    -  commercial
what is an ai agent            1   18100  49   9231    -  informational
open source ai agent           2    9900  20   7920    -  commercial
ai code agent                  3    8100  30   5670    -  commercial
ai powered coding agent        1    8100  32   5508    -  commercial
gemini spark ai agent          1    6600  18   5412    -  informational
autonomous ai agent            1    6600  22   5148    -  informational
ai agent moltbook              1    6600  43   3762    -  transactional
open-source ai coding agent    2    5400  31   3726    -  commercial
self-hosted ai agent           1    3600   0   3600    -  commercial
no-code ai agent builder       2    3600  10   3240   58  commercial
build ai agent                 4    2400  20   1920    4  commercial
how to build ai agents         1    1600  41    944    4  informational
ai agent course                2    1000  12    880    2  commercial
ai agent frameworks            1    1000  27    730    2  commercial
build an ai agent              2    1000  28    720    4  informational

LLM usage: 1957 input / 551 output tokens

计划把 “open source ai agent”、“ai voice agent”、“what is an ai agent”、“self-hosted ai agent”、“no-code ai agent builder” 列为首批目标。它指出 Microsoft Learn 的 AI Agents for Beginners 页面在 “build ai agent” 上已经排第 4、在 “ai agent course” 上排第 2,建议先别正面硬碰。有一条我挺认可:没做 SERP 检查的词,它把页面类型标成”待验证”,没有硬编一个。对于没匹配到排名的行,它说的是 Microsoft 的排名”未提供”,这个理解也是对的,因为 ranked_keywords 只返回 DataForSEO 收录到的数据。

表里有两处提醒你发布前一定要人工看一眼。“self-hosted ai agent” 的难度是 0;“ai agent moltbook” 和 “gemini spark ai agent” 看着像产品名,却被标成了信息型或交易型意图。代码抓不住这些,得靠人。

成本

调用每次运行的次数2026-10-03 计费
keyword_suggestions/live1$0.000000
ranked_keywords/live1$0.000000
serp/google/organic/live/advanced3每次 $0.000000
openai/gpt-6.1-sol1$0.010401

以上是上面那次运行通过 GET /v1/tasks/<id>/cost 查到的数字。2026-10-03 当天,三个 DataForSEO 模型页的基础价格都标 Free,站点还挂着 “API Free Week” 横幅,所以 Free 只能理解为当前的标价,不代表一直免费。打算每天定时跑之前,先去模型页确认一下。

LLM 这一项是费用记录里的实际扣费,不是我推算的。模型页标的价格是提示词 272K 以下输入 $2/M、输出 $10/M,按这个算 1,957 × $2/M + 551 × $10/M = $0.0094,实际扣费 $0.010401,多出约 $0.001。同一条费用记录的 usage 里有 cache_creation_tokens: 1954,但公开的计价公式在这个提示词长度下并不对缓存写入收费,所以光凭公开价格对不上这笔差额,我已经反馈给 SandBase 团队。做预算请以费用记录为准,别只套公式。我一共完整跑了三次,LLM 那一步每次扣费在 $0.0104 到 $0.0116 之间。

局限

  • 测试范围说在前面:一个种子词、一个市场(美国英文)、一个竞品、一天。搜索量和难度都是 DataForSEO 的估算,和 Google Keyword Planner、Search Console 的数字会有出入。
  • 两个 Labs 调用都设了 limit: 50。keyword_suggestions 报告共 615 条匹配,ranked_keywords 共 254 条,这里拿到的只是头部。DataForSEO 文档里有 offset 翻页参数,本文没有测。
  • like "%ai agent%" 是子串匹配,连 “openai agent” 也会匹配上,这次恰好被导航型意图过滤掉了。
  • 相隔 6 分钟的两次运行,SERP 域名几乎全换了。一次抓取不等于排名。
  • 机会分是我自己定的启发式公式。如果你更看重意图或 CPC,就改公式。

FAQ

没有 DataForSEO 账号,能用 DataForSEO API 做关键词调研吗?

走 SandBase 可以。请求字段和 DataForSEO 一样,鉴权换成 SandBase API Key,返回的就是 DataForSEO 自己的返回体,放在 outputs[0] 里(文档写在 outputs[0].data 下,2026-10-03 实测是直接在 outputs[0])。直连 DataForSEO 的话,需要它的 API login 和 password 做 Basic 鉴权。

DataForSEO 的关键词数据就是 Google 的数据吗?

不是。这是 DataForSEO 的第三方估算。搜索量和难度适合用来给自己的关键词列表排先后,别当成 Google 官方数字。

怎么避免搜索量被重复计算?

把近义变体归到一组。这次 “ai agent” 的 7 个变体都报 49,500,逐月序列也一模一样,按序列分组后合成一行,variants = 7。

为什么不让 LLM 直接打分?

因为模型最容易在算术上出错。在我们的 GPT-6.1 Sol 和 Claude Sonnet 5.5 对比测试里,三个错误答案全是对工具结果求平均或计数。Python 算得准,还不花钱;模型只负责写计划。

这类 Agent 还能搭配哪些搜索 API?

如果 Agent 需要通用网页搜索,可以看我们的网页搜索 API 盘点和 Exa、Tavily、Firecrawl、SerpAPI 对比。DataForSEO 在这里的优势是 SEO 指标:搜索量、难度,以及按域名查已排名关键词。

下一步

申请 SandBase API Key,打开 DataForSEO 目录页,或者看 ranked_keywords 参考页。把种子词和竞品域名换成你自己的,SERP 那一列至少跑两次再信。