DataForSEO 关键词调研 Agent:用 SandBase 实测搭建
用 Python 搭一个 DataForSEO 关键词调研 Agent:拿关键词搜索量和难度、实时 SERP、竞品已排名的词,数字在代码里算,最后由 LLM 写计划。全部通过 SandBase 一个 Key 调用。

种子词填 “ai agent”,DataForSEO 一口气返回了 7 个月搜索量都是 49,500 的词:“ai agent”、“agent ai”、“agent in ai”、“ai intelligent agent”,还有另外三个。再一看,它们过去 12 个月的逐月数据一模一样。要是直接把这几行加起来,你会得到 346,500 次根本不存在的搜索。
这是做 DataForSEO 关键词调研 Agent 时碰到的第一个坑,也正好说明一件事:数字得在代码里算,别交给模型。本文用大约 170 行 Python 搭一个这样的 Agent。输入一个种子主题和一个竞品域名,它会输出一份排好优先级的关键词表,包含搜索量、难度、前三个关键词当前的自然搜索结果,以及竞品已经在排名的词。最后由 openai/gpt-6.1-sol 根据算好的数字写一份简短计划。所有调用都走 SandBase,一个 API Key 搞定。
先说结论
- 三个 DataForSEO 端点就够用:
keyword_suggestions(关键词、搜索量、难度、意图)、ranked_keywords(竞品排了哪些词)、serp/google/organic/live/advanced(实时自然结果)。截至 2026-10-03,三个在 SandBase 上都是启用状态。- 完整跑一次是 5 次数据调用加 1 次 LLM 调用。2026-10-03 数据调用的计费是 $0.00(模型页标 Free,站点正在搞 “API Free Week”),LLM 那次花了 $0.0104。
- 近义变体共用同一个搜索量。按 12 个月搜索量序列分组后,50 条建议被合并成更少、更真实的行,光 “ai agent” 一组就吸收了 7 个变体。
- 实时 SERP 变得很快:相隔约 6 分钟的两次查询,“ai agent” 的前五个自然结果域名一个都不重合。一次抓取只能当样本看。
- 这些是 DataForSEO 提供的第三方 SEO 估算数据,不是 Google 自己的数据。
Agent 产出什么
| 产出 | 来源 | SandBase 价格(2026-10-03) |
|---|---|---|
含种子词的关键词,带 search_volume、keyword_difficulty、main_intent | dataforseo_labs/google/keyword_suggestions/live | Free(模型页) |
| 竞品已排名的词,以及排名位置和 URL | dataforseo_labs/google/ranked_keywords/live | Free(模型页) |
| 前 3 个关键词的自然结果域名 | serp/google/organic/live/advanced | Free(模型页) |
| 变体分组、机会分、合并、导出 CSV | Python | — |
| 不超过 8 条的计划 | openai/gpt-6.1-sol | 输入 $2/M、输出 $10/M token |
最终输出是 keyword_plan.csv,外加终端里打印的表格和计划。
SandBase 上的 DataForSEO 目录页截至 2026-10-03 列出 74 个端点。选中的卡片(一个 OnPage 端点,本文没用到)显示 Available 和 Free:
截图:SandBase 的 DataForSEO 目录页列出 74 个端点,包括 SERP Google Organic Live,页面顶部挂着 “API Free Week” 横幅(2026-10-03 截取)。
目录页展示的是 GET /apis/v1/... 这一套接口。本文用的是端点参考页里的 Model API 路由:POST https://api.sandbase.ai/v1/api/dataforseo/<path>。
数据边界,以及直连 DataForSEO 和走 SandBase 的区别
这里用的是公开、只读的 SEO 数据。关键词搜索量、难度分和 SERP 快照,都是上游 DataForSEO 给出的第三方估算,不是 Google 自己的数据,SandBase 和 Google 也没有任何关联。Agent 需要一个 SandBase API Key。它不碰 Google Search Console、Google Ads 账户、站长后台数据,也不会对任何网站或账户做改动。
你也可以直接调 DataForSEO,区别主要在接入方式:
| 直连 DataForSEO | 通过 SandBase 调 DataForSEO | |
|---|---|---|
| 鉴权 | 用 DataForSEO 的 API login 和 password 做 HTTP Basic(鉴权文档) | Authorization: Bearer $SANDBASE_API_KEY |
| 请求体 | 任务对象组成的 JSON 数组 | 单个任务对象,字段相同 |
| 返回 | DataForSEO 的 tasks[].result[] | 同样的 DataForSEO 返回体,放在 SandBase 的 outputs[0] 里 |
| LLM 那一步 | 另一家供应商、另一把 Key、另一张账单 | 同一把 Key,POST /v1/chat/completions |
| 适合 | 只用 DataForSEO、量大、想自己签合同 | 想把 SEO 数据和模型放在一把 Key、一份合同后面 |
两边的价格我没有做对比,DataForSEO 自己公开了定价。
第 1 步:拿关键词、搜索量和难度
SandBase 的 DataForSEO 参考页没有列请求字段,模型页上写的是 “0 input fields”,请求体会原样转给 DataForSEO。所以参数以 DataForSEO 的 keyword_suggestions 文档为准:keyword(必填)、location_code、language_code、limit、filters、order_by、include_seed_keyword。
截图:keyword_suggestions/live 模型页显示基础价格 Free、同步执行、0 个输入字段,所以请求字段要看 DataForSEO 自己的文档(2026-10-03 截取)。
实测:2026-10-03(UTC)
公开输入:通用种子词 “ai agent”,Google 美国(location_code 2840),英文,竞品域名用 Microsoft Learn(learn.microsoft.com)。
curl -s https://api.sandbase.ai/v1/api/dataforseo/v3/dataforseo_labs/google/keyword_suggestions/live \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"keyword": "ai agent", "location_code": 2840, "language_code": "en",
"include_seed_keyword": true, "limit": 50,
"filters": [["keyword_info.search_volume", ">=", 100]],
"order_by": ["keyword_info.search_volume,desc"]}'
截短后的返回(run id 6ed0cb41-ad1f-4bf3-a10e-bb1c046e2223,50 条中的 1 条):
{
"id": "6ed0cb41-ad1f-4bf3-a10e-bb1c046e2223",
"status": "completed",
"model": "dataforseo/v3/dataforseo_labs/google/keyword_suggestions/live",
"outputs": [{
"status_code": 20000,
"tasks": [{"status_code": 20000, "result": [{
"total_count": 615, "items_count": 50,
"items": [{
"keyword": "ai agent",
"keyword_info": {"search_volume": 49500, "competition_level": "MEDIUM", "cpc": 20.9,
"monthly_searches": [{"year": 2026, "month": 8, "search_volume": 49500}, "…"]},
"keyword_properties": {"core_keyword": "agentic ai", "keyword_difficulty": 70},
"search_intent_info": {"main_intent": "commercial"}
}]
}]}]
}]
}
抄解析代码之前,先说个坑。端点参考页写的完成态返回是 outputs[0].data,代码也优先按这个约定读。但我 2026-10-03 调的 22 次 DataForSEO,全部是把 DataForSEO 自己的返回体直接放在 outputs[0](version、status_code、tasks),没有 data 这一层。所以函数先看 outputs[0].data 里有没有 DataForSEO 返回体,没有再退回读 outputs[0],两处都找不到 tasks 就直接报错。这个不一致我已经反馈给 SandBase 团队。
上面这些业务字段(search_volume、keyword_difficulty、monthly_searches、core_keyword、main_intent)来自这几次实测和 DataForSEO 文档,属于观察到的字段,不是 SandBase 的保证。代码里一律用 .get() 读。返回里还有一个 DataForSEO 自己的 cost 字段,你的 SandBase 账户实际扣多少,以 GET /v1/tasks/<id>/cost 为准。
第 2 步:竞品已经排了哪些词
ranked_keywords 接收一个 target 域名。它的数据是 DataForSEO 的快照,官方文档说每周更新,所以这里的排名不等于当天的 Google 实时排名(后面的 SERP 那一步才是实时查的)。我只筛含种子词的关键词,按搜索量排序:
{"target": "learn.microsoft.com", "location_code": 2840, "language_code": "en", "limit": 50,
"filters": [["keyword_data.keyword", "like", "%ai agent%"]],
"order_by": ["keyword_data.keyword_info.search_volume,desc"]}
run id d38178c1-e21d-4610-9763-02c222363b14 报告 total_count 为 254,返回了 50 条。截一条:
{"keyword_data": {"keyword": "ai agent course",
"keyword_info": {"search_volume": 1000},
"keyword_properties": {"keyword_difficulty": 12}},
"ranked_serp_element": {"serp_item": {"type": "organic", "rank_group": 2, "rank_absolute": 3,
"domain": "learn.microsoft.com",
"url": "https://learn.microsoft.com/en-us/shows/ai-agents-for-beginners/"}}}
keyword_data 的字段名和第 1 步的关键词条目一致,所以一个 row_from() 就能把两边都展平。rank_group 是在自然结果里的名次,rank_absolute 则把页面上所有元素都算进去。
截图:ranked_keywords/live 参考页显示 POST /v1/api/dataforseo/v3/dataforseo_labs/google/ranked_keywords/live,完成态示例是 outputs[0].data,和实测返回的结构不一样(2026-10-03 截取)。
参考页自动生成的 curl 示例在请求体里放了 model 字段。路径里已经写明了模型,所以代码只发 DataForSEO 的字段。
定下这一步之前,我先试了个看起来最省事的办法:把之前一次测试拿到的 30 条建议词用 in 过滤丢给 ranked_keywords(run id 75cad2ec-fd7d-4eda-ace5-8274c295eff0)。结果 Microsoft Learn 只命中 1 个(“no code ai agent builder”,第 58 名)。它真正有排名的 “build ai agents”、“ai agent course” 这类词,压根不在建议列表里。做竞品差距分析,得把两张表合并,而不是取交集。
第 3 步:前几个关键词的实时 SERP
SERP 端点(DataForSEO 文档)接收 keyword、location_code、language_code 和 depth。depth: 10 时,三次调用各返回 12 到 13 个元素,其中 organic 只有 6 到 8 个,第一个元素每次都是 ai_overview。所以代码按 type == "organic" 过滤。“ai voice agent” 这次(run id a6618702-02a4-47bd-8f85-e848f2309252):
[{"type": "ai_overview", "rank_group": 1, "rank_absolute": 1},
{"type": "organic", "rank_group": 1, "rank_absolute": 2, "domain": "www.retellai.com"},
{"type": "organic", "rank_group": 2, "rank_absolute": 3, "domain": "elevenlabs.io"},
{"type": "organic", "rank_group": 3, "rank_absolute": 4, "domain": "www.reddit.com"}]
6 分钟后我把整个程序又跑了一遍。“ai agent” 在 04:10 UTC 那次的前五个自然结果是 IBM、aiagent.app、AWS、Wikipedia、Meta;04:16 那次变成了 Reddit、LinkedIn、retresco.de、airia.com、airbyte.com,一个都不重合。“what is an ai agent” 也完全不重合,“ai voice agent” 只剩 Reddit 是共同的。两次样本分不出这是 Google 本身的波动、AI Overview 版式,还是 DataForSEO 的采集方式。所以计划里 SERP 域名只用来判断页面类型,不当作排名目标清单。
第 4 步:在代码里打分
规则都在 build_table() 里,一共三条:
- 按关键词合并两张表,把竞品的
rank_group和 URL 写到对应行上。 - 去掉搜索量低于 100、没有难度值、或意图是
navigational的行(比如 “n8n ai agent” 这种品牌词)。 - 按完全相同的 12 个月搜索量序列,把近义变体归为一组。代表词优先选含种子词的,其次选最短的。
variants记录合并了几行。
然后 opportunity = volume × (100 − difficulty) / 100。说白了就是用 DataForSEO 的 0–100 难度给搜索量打个折,是个启发式分数,不是流量预测。表里保留机会分最高的 12 行,再补最多 5 行竞品有排名的词,免得竞品那部分被挤到表外。
为什么按搜索量序列分组,而不用 DataForSEO 的 core_keyword?这次数据里,core_keyword 把 “ai agent” 归到了 “agentic ai”,却把 “ai intelligent agent” 分到另一组,可这两个词的 49,500 和逐月数字完全一样。按序列分组更粗糙,但它针对的正是真问题:搜索量被重复计算。两个不相干的词也可能碰巧序列相同,不过在搜索量 ≥ 100、12 个月的条件下,我这次没遇到。
第 5 步:让 LLM 写计划
openai/gpt-6.1-sol 拿到的是 JSON 格式的表格和 SERP 域名,外加一条硬性要求:只引用给定数据里的数字,不要自己算新数字。选它,是因为我们的 Claude Sonnet 5.5 和 GPT-6.1 Sol Agent 对比测试里,它 30 个短工具任务全对,每个任务约 $0.0066。那次测试里出错的地方恰好都是对工具结果做算术,这也是本文把求和、分组、打分都留在 Python 里的原因。调用走 OpenAI 兼容的 Chat Completions 接口。
完整代码
#!/usr/bin/env python3
"""SEO keyword-research agent: DataForSEO data through SandBase, numbers in code, plan by an LLM.
Usage: python3 seo_research_agent.py "ai agent" learn.microsoft.com
Needs SANDBASE_API_KEY in the environment. Writes keyword_plan.csv and prints a plan.
"""
import csv
import json
import os
import sys
import requests
FIELDS = ["keyword", "variants", "volume", "difficulty", "intent", "opportunity",
"competitor_rank", "competitor_url", "serp_top3"]
API = "https://api.sandbase.ai"
HEADERS = {
"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}",
"Content-Type": "application/json",
}
LOCALE = {"location_code": 2840, "language_code": "en"} # United States, English
TOP_N = 12 # best rows by opportunity
COMPETITOR_EXTRA = 5 # extra rows where the competitor ranks, even below the cut
SERP_CHECKS = 3 # live SERP lookups for the top keywords
def dataforseo(path: str, params: dict) -> dict:
"""POST /v1/api/dataforseo/<path> and return the first DataForSEO task result."""
resp = requests.post(f"{API}/v1/api/dataforseo/{path}", headers=HEADERS, json=params, timeout=120)
resp.raise_for_status()
body = resp.json()
outputs = body.get("outputs") or []
if body.get("status") != "completed" or not outputs:
raise RuntimeError(f"SandBase call did not complete: {body.get('status')} {body.get('error')}")
# Documented contract: outputs[0].data. Observed on 2026-10-03: DataForSEO's body sat directly in outputs[0].
data = outputs[0].get("data")
payload = data if isinstance(data, dict) and "tasks" in data else outputs[0]
tasks = payload.get("tasks") or []
if not tasks or tasks[0].get("status_code") != 20000:
task = tasks[0] if tasks else payload
raise RuntimeError(f"DataForSEO task failed: {task.get('status_code')} {task.get('status_message')}")
results = tasks[0].get("result") or []
return results[0] if results else {}
def keyword_ideas(seed: str, limit: int = 50) -> list[dict]:
"""Keywords that contain the seed, with volume >= 100, biggest first."""
result = dataforseo("v3/dataforseo_labs/google/keyword_suggestions/live", {
"keyword": seed, **LOCALE, "include_seed_keyword": True, "limit": limit,
"filters": [["keyword_info.search_volume", ">=", 100]],
"order_by": ["keyword_info.search_volume,desc"],
})
return result.get("items") or []
def competitor_keywords(domain: str, seed: str, limit: int = 50) -> list[dict]:
"""Keywords containing the seed that the competitor domain ranks for."""
result = dataforseo("v3/dataforseo_labs/google/ranked_keywords/live", {
"target": domain, **LOCALE, "limit": limit,
"filters": [["keyword_data.keyword", "like", f"%{seed}%"]],
"order_by": ["keyword_data.keyword_info.search_volume,desc"],
})
return result.get("items") or []
def top_organic(keyword: str, n: int = 5) -> list[dict]:
"""Live Google organic results (first page) for one keyword."""
result = dataforseo("v3/serp/google/organic/live/advanced", {"keyword": keyword, **LOCALE, "depth": 10})
organic = [i for i in result.get("items") or [] if i.get("type") == "organic"]
return [{"rank": i.get("rank_group"), "domain": i.get("domain"), "url": i.get("url")} for i in organic[:n]]
def row_from(kw: dict) -> dict:
"""Flatten the keyword fields both Labs endpoints share."""
info = kw.get("keyword_info") or {}
props = kw.get("keyword_properties") or {}
# Close variants ("ai agent", "agent ai") came back with identical 12-month volume series.
series = tuple(m.get("search_volume") for m in info.get("monthly_searches") or [])
return {
"keyword": kw.get("keyword"),
"cluster": series or kw.get("keyword"),
"volume": info.get("search_volume") or 0,
"difficulty": props.get("keyword_difficulty"),
"intent": (kw.get("search_intent_info") or {}).get("main_intent"),
"competitor_rank": None,
"competitor_url": "",
}
def build_table(seed: str, ideas: list[dict], ranked: list[dict]) -> list[dict]:
"""Merge both sources, collapse close variants, score, and keep the top rows."""
rows = {}
for kw in ideas:
rows[kw.get("keyword")] = row_from(kw)
for item in ranked:
kw = item.get("keyword_data") or {}
serp_item = (item.get("ranked_serp_element") or {}).get("serp_item") or {}
row = rows.setdefault(kw.get("keyword"), row_from(kw))
row["competitor_rank"] = serp_item.get("rank_group")
row["competitor_url"] = serp_item.get("url") or ""
# One row per variant group. Representative: contains the seed, then the shortest keyword.
keep = [r for r in rows.values()
if r["volume"] >= 100 and r["difficulty"] is not None and r["intent"] != "navigational"]
keep.sort(key=lambda r: (seed not in r["keyword"], len(r["keyword"])))
best = {}
for row in keep:
winner = best.setdefault(row["cluster"], {**row, "variants": 0})
winner["variants"] += 1
if row["competitor_rank"] and not winner["competitor_rank"]:
winner["competitor_rank"], winner["competitor_url"] = row["competitor_rank"], row["competitor_url"]
# Opportunity = volume discounted by difficulty (0-100). A heuristic, not a traffic forecast.
for row in best.values():
row["opportunity"] = round(row["volume"] * (100 - row["difficulty"]) / 100)
ranked_rows = sorted(best.values(), key=lambda r: r["opportunity"], reverse=True)
# Top rows overall, plus the competitor's best keywords even if they fall below the cut.
table = ranked_rows[:TOP_N]
table += [r for r in ranked_rows if r["competitor_rank"] and r not in table][:COMPETITOR_EXTRA]
return table
def write_plan(seed: str, domain: str, table: list[dict], serps: dict) -> tuple[str, dict]:
"""Ask openai/gpt-6.1-sol for a short plan. It gets computed numbers and must not invent new ones."""
prompt = (
f"Seed topic: {seed}. Competitor: {domain}. Market: Google US, English.\n"
f"Keyword table (opportunity = volume x (100 - difficulty) / 100, computed in code):\n"
f"{json.dumps([{k: r.get(k) for k in FIELDS} for r in table], ensure_ascii=False)}\n"
f"Current top organic results for the top keywords:\n{json.dumps(serps, ensure_ascii=False)}\n\n"
"Write a content plan in at most 8 bullets: which 3-5 keywords to target first and why, "
"where the competitor already ranks, and what page type each target needs based on the SERP. "
"Quote numbers only from the data above; do not compute or estimate new ones."
)
resp = requests.post(f"{API}/v1/chat/completions", headers=HEADERS, timeout=120, json={
"model": "openai/gpt-6.1-sol",
"messages": [{"role": "user", "content": prompt}],
"max_tokens": 800,
})
resp.raise_for_status()
body = resp.json()
return body["choices"][0]["message"]["content"], body.get("usage") or {}
def main(seed: str, domain: str) -> None:
ideas = keyword_ideas(seed)
ranked = competitor_keywords(domain, seed)
print(f"{len(ideas)} keyword suggestions, {len(ranked)} competitor keywords containing '{seed}'")
table = build_table(seed, ideas, ranked)
serps = {row["keyword"]: top_organic(row["keyword"]) for row in table[:SERP_CHECKS]}
for row in table[:SERP_CHECKS]:
row["serp_top3"] = " | ".join(r["domain"] for r in serps[row["keyword"]][:3])
with open("keyword_plan.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=FIELDS, extrasaction="ignore")
writer.writeheader()
writer.writerows(table)
print(f"\n{'keyword':28} {'var':>3} {'volume':>7} {'KD':>3} {'opp':>6} {'comp':>4} intent")
for row in table:
rank = row["competitor_rank"] or "-"
print(f"{row['keyword'][:28]:28} {row['variants']:>3} {row['volume']:>7} {row['difficulty']:>3} "
f"{row['opportunity']:>6} {rank:>4} {row['intent']}")
plan, usage = write_plan(seed, domain, table, serps)
print("\nPlan (openai/gpt-6.1-sol):\n" + plan)
print(f"\nLLM usage: {usage.get('prompt_tokens')} input / {usage.get('completion_tokens')} output tokens")
if __name__ == "__main__":
if len(sys.argv) != 3:
sys.exit('usage: python3 seo_research_agent.py "<seed topic>" <competitor domain>')
main(sys.argv[1], sys.argv[2])
需要 Python 3.9 以上和 requests。运行:python3 seo_research_agent.py "ai agent" learn.microsoft.com。
实际跑出来的结果
下面是上面这段程序在 2026-10-03 04:15 UTC 的输出(6 次调用全部 completed):
50 keyword suggestions, 50 competitor keywords containing 'ai agent'
keyword var volume KD opp comp intent
ai agent 7 49500 70 14850 - commercial
ai voice agent 1 22200 38 13764 - commercial
what is an ai agent 1 18100 49 9231 - informational
open source ai agent 2 9900 20 7920 - commercial
ai code agent 3 8100 30 5670 - commercial
ai powered coding agent 1 8100 32 5508 - commercial
gemini spark ai agent 1 6600 18 5412 - informational
autonomous ai agent 1 6600 22 5148 - informational
ai agent moltbook 1 6600 43 3762 - transactional
open-source ai coding agent 2 5400 31 3726 - commercial
self-hosted ai agent 1 3600 0 3600 - commercial
no-code ai agent builder 2 3600 10 3240 58 commercial
build ai agent 4 2400 20 1920 4 commercial
how to build ai agents 1 1600 41 944 4 informational
ai agent course 2 1000 12 880 2 commercial
ai agent frameworks 1 1000 27 730 2 commercial
build an ai agent 2 1000 28 720 4 informational
LLM usage: 1957 input / 551 output tokens
计划把 “open source ai agent”、“ai voice agent”、“what is an ai agent”、“self-hosted ai agent”、“no-code ai agent builder” 列为首批目标。它指出 Microsoft Learn 的 AI Agents for Beginners 页面在 “build ai agent” 上已经排第 4、在 “ai agent course” 上排第 2,建议先别正面硬碰。有一条我挺认可:没做 SERP 检查的词,它把页面类型标成”待验证”,没有硬编一个。对于没匹配到排名的行,它说的是 Microsoft 的排名”未提供”,这个理解也是对的,因为 ranked_keywords 只返回 DataForSEO 收录到的数据。
表里有两处提醒你发布前一定要人工看一眼。“self-hosted ai agent” 的难度是 0;“ai agent moltbook” 和 “gemini spark ai agent” 看着像产品名,却被标成了信息型或交易型意图。代码抓不住这些,得靠人。
成本
| 调用 | 每次运行的次数 | 2026-10-03 计费 |
|---|---|---|
keyword_suggestions/live | 1 | $0.000000 |
ranked_keywords/live | 1 | $0.000000 |
serp/google/organic/live/advanced | 3 | 每次 $0.000000 |
openai/gpt-6.1-sol | 1 | $0.010401 |
以上是上面那次运行通过 GET /v1/tasks/<id>/cost 查到的数字。2026-10-03 当天,三个 DataForSEO 模型页的基础价格都标 Free,站点还挂着 “API Free Week” 横幅,所以 Free 只能理解为当前的标价,不代表一直免费。打算每天定时跑之前,先去模型页确认一下。
LLM 这一项是费用记录里的实际扣费,不是我推算的。模型页标的价格是提示词 272K 以下输入 $2/M、输出 $10/M,按这个算 1,957 × $2/M + 551 × $10/M = $0.0094,实际扣费 $0.010401,多出约 $0.001。同一条费用记录的 usage 里有 cache_creation_tokens: 1954,但公开的计价公式在这个提示词长度下并不对缓存写入收费,所以光凭公开价格对不上这笔差额,我已经反馈给 SandBase 团队。做预算请以费用记录为准,别只套公式。我一共完整跑了三次,LLM 那一步每次扣费在 $0.0104 到 $0.0116 之间。
局限
- 测试范围说在前面:一个种子词、一个市场(美国英文)、一个竞品、一天。搜索量和难度都是 DataForSEO 的估算,和 Google Keyword Planner、Search Console 的数字会有出入。
- 两个 Labs 调用都设了
limit: 50。keyword_suggestions报告共 615 条匹配,ranked_keywords共 254 条,这里拿到的只是头部。DataForSEO 文档里有offset翻页参数,本文没有测。 like "%ai agent%"是子串匹配,连 “openai agent” 也会匹配上,这次恰好被导航型意图过滤掉了。- 相隔 6 分钟的两次运行,SERP 域名几乎全换了。一次抓取不等于排名。
- 机会分是我自己定的启发式公式。如果你更看重意图或 CPC,就改公式。
FAQ
没有 DataForSEO 账号,能用 DataForSEO API 做关键词调研吗?
走 SandBase 可以。请求字段和 DataForSEO 一样,鉴权换成 SandBase API Key,返回的就是 DataForSEO 自己的返回体,放在 outputs[0] 里(文档写在 outputs[0].data 下,2026-10-03 实测是直接在 outputs[0])。直连 DataForSEO 的话,需要它的 API login 和 password 做 Basic 鉴权。
DataForSEO 的关键词数据就是 Google 的数据吗?
不是。这是 DataForSEO 的第三方估算。搜索量和难度适合用来给自己的关键词列表排先后,别当成 Google 官方数字。
怎么避免搜索量被重复计算?
把近义变体归到一组。这次 “ai agent” 的 7 个变体都报 49,500,逐月序列也一模一样,按序列分组后合成一行,variants = 7。
为什么不让 LLM 直接打分?
因为模型最容易在算术上出错。在我们的 GPT-6.1 Sol 和 Claude Sonnet 5.5 对比测试里,三个错误答案全是对工具结果求平均或计数。Python 算得准,还不花钱;模型只负责写计划。
这类 Agent 还能搭配哪些搜索 API?
如果 Agent 需要通用网页搜索,可以看我们的网页搜索 API 盘点和 Exa、Tavily、Firecrawl、SerpAPI 对比。DataForSEO 在这里的优势是 SEO 指标:搜索量、难度,以及按域名查已排名关键词。
下一步
申请 SandBase API Key,打开 DataForSEO 目录页,或者看 ranked_keywords 参考页。把种子词和竞品域名换成你自己的,SERP 那一列至少跑两次再信。