LLM Agent Tool Calling Benchmark 2026: 12 Models Tested
LLM agent tool calling benchmark 2026: 12 models, 17 scored tasks, 3 runs each. Accuracy, cost per task, latency and why most misses were hand arithmetic.

Task H2 gave every model the same 20 Douyin videos and asked for the IDs of the ones with more likes than the page average. Four videos qualified. Claude Sonnet 5.5 listed seven, six and five in its three runs, adding videos with as few as 337,922 likes against an average of 422,764. Eight other runs, from six different models, never delivered an answer: they spent the 1,500-token output cap writing out the arithmetic, and some were cut off halfway through the final line.
That pattern runs through this LLM agent tool calling benchmark. We gave 12 models on SandBase the same six data tools, the same system prompt and 17 scored tasks, three runs each, on 2026-10-03 (UTC). Tool use itself was clean everywhere. The misses came after the tools returned, when a model had to do bookkeeping in its head, or when it took a number from the wrong source.
Key takeaway
- Four models scored 51/51: GPT-6.1 Sol, Grok 4.7, Kimi K3 and Qwen3.8 Max Prime. MiMo v2.6 Pro scored 50/51; Claude Sonnet 5.5 was lowest at 44/51.
- All 12 models made schema-valid tool calls every time (89 to 104 calls each). Tool selection did not separate them.
- 10 of the 22 misses were output truncation at 1,500 tokens while working arithmetic out by hand. Rerunning those tasks at 4,000 tokens passed 26 of 27 runs.
- Median cost per task ran from $0.0018 (MiMo v2.6 Pro) and $0.0019 (MiniMax M3) to $0.0069 (GPT-6.1 Sol), $0.0159 (Kimi K3) and $0.0213 (Qwen3.8 Max Prime), at list price.
- Scope: one task family (Chinese social data lookups plus aggregation), N=3, default settings, one day. Not a general model ranking.
This is the expanded follow-up to our Claude Sonnet 5.5 vs GPT-6.1 Sol agent test. Same harness, same 10 short tasks (T1 to T10), plus 8 harder ones (H1 to H8) and ten more models, most of them from Chinese labs.
The 12 models and their prices
Every price below is the SandBase list price on 2026-10-03, read from the price_formula in GET https://api.sandbase.ai/v1/models/<id>, in dollars per million tokens.
| Model (SandBase id) | Input | Output |
|---|---|---|
openai/gpt-6.1-sol | $2 | $10 |
anthropic/claude-sonnet-5.5 | $2 | $10 |
x-ai/grok-4.7 | $2 | $6 |
moonshotai/kimi-k3 | $3 | $15 |
alibaba/qwen3.8-max-prime | $4 | $12 |
alibaba/qwen3.8-max-0902 | $2 | $6 |
z-ai/glm-5.3-prime | $2.80 | $8.80 |
z-ai/glm-5.3 | $1.40 | $4.40 |
deepseek/deepseek-v4-pro-0813 | $1.32 | $3.96 |
minimax/minimax-m3 | $0.30 | $1.20 |
xiaomi/mimo-v2.6-pro | $0.435 | $0.87 |
bytedance/seed-2.1-pro | $0.89 | $4.45 |
GPT-6.1 Sol and Grok 4.7 switch to a higher tier above 272K and 200K prompt tokens. No request in this test came close.

Caption: The SandBase catalog row for moonshotai/kimi-k3 shows $3 input and $15 output per million tokens, the rate behind Kimi K3’s cost figures here (captured 2026-10-03).
Kimi K3 is the most expensive per output token of the twelve, and that shows up in cost per task: only Qwen3.8 Max Prime cost more at the median. Its input was the leanest of all 12 models (median 2,870 tokens), so most of its gap to GPT-6.1 Sol is output: about 410 tokens per task at $15 per million, against 110 at $10.
Test setup: 6 tools, 18 tasks, exact-match grading
Claude Sonnet 5.5 ran through SandBase POST /v1/messages (Anthropic format). The other 11 models ran through SandBase POST /v1/chat/completions with OpenAI function tools, changing only the model field. Each turn allowed 1,500 output tokens, each task 10 turns, and temperature and reasoning settings stayed at their defaults.
The six tools wrap SandBase public data endpoints: douyin/search/user-search-v2, douyin/app-v3/user-profile, douyin/app-v3/user-post-videos, weibo/web-v2/hot-search, weibo/web-v2/realtime-search and xiaohongshu/app-v2/search-notes.
To keep the comparison fair, the harness cached each tool response by tool name and arguments and replayed the cache to every model. The T-set cache is the one from the first article, captured 02:50 to 02:58 UTC, plus three small lookups added between 07:30 and 07:35 UTC. The extra data the hard tasks needed (the Starbucks China profile and two Weibo searches) was captured around 07:50 UTC. Runs took place between 07:36 and about 08:16 UTC, all on 2026-10-03. When a model asked for something not in the cache, such as a different search keyword, the call went to the live API, and that result is part of that model’s run.
Expected answers were computed in Python from the same cache: filters, max, sum, sorted, string matching. No model output went into the answer key. The T tasks (T1 Luckin Coffee follower lookup through T10 camping maxima) are listed in full in the first article. The hard tasks:
| Task | What the agent had to do | What it tests |
|---|---|---|
| H1 | Rank Luckin Coffee, CottiCoffee and Starbucks China by profile followers and sum them | Three lookups, sort, sum of three 7-digit numbers |
| H2 | List the IDs of first-page videos with likes strictly above the page average, highest first | Mean of 20, filter, sort, ordered list |
| H3 | Sum the heat values of Weibo hot-search ranks 11 to 20 and count those above 100,000 | Adding ten large numbers |
| H4 | Total first-page Xiaohongshu likes for sunscreen and for camping; report the larger | Two sums over two lists |
| H5 | Search Weibo for hot-search topics #2 and #3 and report each first-page post count | Two chained lookups |
| H6 | In Douyin user search for camping, count accounts above 500,000 followers and give the third-largest | Count and rank |
| H7 | Find the account named exactly 库迪咖啡 (Cotti Coffee) | Excluded from scoring (see below) |
| H8 | Among videos created on 2026-10-02, find the most-shared one and total their comments | Date filter, max, sum |
The system prompt, verbatim, was the same for all 12 models:
You are a research agent with tools for public Douyin, Weibo and Xiaohongshu data. Use the tools to answer; never guess numbers. If something can't be found, say so. Finish with ONE line that starts with FINAL: followed by a compact JSON object with exactly the keys requested.
The grader parsed the JSON object after FINAL: and compared each requested key: integers after stripping commas, booleans after lowercasing, IDs and names as exact strings, and lists (H1’s ranking, H2’s IDs) element by element in order. A missing key, unparseable JSON or no FINAL: line was a miss; extra keys were ignored.
H7 is out of the score because our task was ambiguous. It asked for the account named exactly 库迪咖啡 (Cotti Coffee), and there turned out to be two small accounts with that exact nickname, each with fewer than 40 followers. The behavior was still telling. GPT-6.1 Sol and Qwen3.8 Max Prime surfaced both accounts in all three runs (Prime once as a pick plus a note). DeepSeek V4 Pro listed both every time, but as a JSON array, which our grader couldn’t parse. Grok 4.7, Kimi K3 and Seed 2.1 Pro picked the same single account in every run, and the rest varied. Our internal summary file only tallies list-valued answers as “both”, so it shows 2 for Qwen Prime and 0 for DeepSeek.
We aren’t publishing the cached responses, because they are third-party platform content about real accounts. The task list, prompt, grading rule and aggregate results are described here, which is enough to rebuild the test against your own fresh snapshot; the raw cache and harness are kept internally.
Results: accuracy, cost and latency
| Model | $/M in / out | Easy (T, 30) | Hard (H, 21) | Total (51) | Median cost / task | Median time | Truncated |
|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol | $2 / $10 | 30 | 21 | 51 | $0.0069 | 9.5 s | 0 |
| Grok 4.7 | $2 / $6 | 30 | 21 | 51 | $0.0139 | 8.7 s | 0 |
| Kimi K3 | $3 / $15 | 30 | 21 | 51 | $0.0159 | 8.5 s | 0 |
| Qwen3.8 Max Prime | $4 / $12 | 30 | 21 | 51 | $0.0213 | 8.4 s | 0 |
| MiMo v2.6 Pro | $0.435 / $0.87 | 30 | 20 | 50 | $0.0018 | 12.7 s | 1 |
| DeepSeek V4 Pro | $1.32 / $3.96 | 28 | 21 | 49 | $0.0069 | 10.2 s | 0 |
| GLM-5.3 | $1.40 / $4.40 | 29 | 20 | 49 | $0.0066 | 21.5 s | 1 |
| Qwen3.8 Max 0902 | $2 / $6 | 30 | 19 | 49 | $0.0112 | 12.3 s | 2 |
| GLM-5.3 Prime | $2.80 / $8.80 | 30 | 19 | 49 | $0.0130 | 9.8 s | 2 |
| MiniMax M3 | $0.30 / $1.20 | 28 | 20 | 48 | $0.0019 | 10.2 s | 1 |
| Seed 2.1 Pro | $0.89 / $4.45 | 29 | 19 | 48 | $0.0058 | 17.1 s | 3 |
| Claude Sonnet 5.5 | $2 / $10 | 27 | 17 | 44 | $0.0130 | 13.6 s | 0 |
Every model made between 89 and 104 tool calls across its 51 scored runs, and none of them was schema-invalid. On T9, the non-existent account, all 12 said it didn’t exist in every run.
GPT-6.1 Sol repeated its first-article result on the T set (30/30, median $0.0066 on those tasks) and kept it on the hard set. Sonnet 5.5 ran the T set fresh for this study and scored 27/30 again, with different misses.
Where models lost points
The 22 misses fall into three groups, and the two biggest have the same root.
Truncation while doing arithmetic by hand (10 runs). Qwen3.8 Max 0902 truncated twice, GLM-5.3 Prime twice, GLM-5.3 once, MiniMax M3 once, MiMo v2.6 Pro once and Seed 2.1 Pro three times. Eight of the ten were H2. The API returned finish_reason: length and there was no FINAL line. The text tails show models writing out running sums or tables of all 20 videos. In several runs the FINAL line had already started and was cut off mid-list. We reran exactly those tasks at 4,000 output tokens per turn: 26 of 27 runs passed. GLM-5.3 truncated once more, with 4,105 output tokens across the run’s two turns when its last turn hit the cap. That doesn’t make these models 51/51 at a larger budget, since we only reran the tasks that failed. It does show that the budget, not the reasoning, ended most of those runs.
Wrong answers (11 runs). Claude Sonnet 5.5 had seven, all arithmetic or counting over tool output, and every one came with a parseable FINAL line:
- T5 average likes: 422,762 and 415,962 against 422,764.
- T7 date count: 17 against 19.
- H2: included one to three videos below the average in all three runs.
- H3 heat sum: 1,822,122 against 1,832,122, off by exactly 10,000.
The other four were instruction misses. On T4, which asks for followers “per profile”, DeepSeek V4 Pro once and MiniMax M3 twice skipped the two profile calls and subtracted the follower counts shown in search results. Their difference, 1,298,178, was 117 off the profile-based 1,298,061. On T2, DeepSeek V4 Pro copied the account ID with three characters out of order. The live endpoint returned an empty video list for that ID, and the model reported, accurately, that no videos came back and answered null. It just didn’t notice the ID was its own typo. That call was schema-valid, so it doesn’t show up as an invalid tool call.
No FINAL line (1 run). GLM-5.3 on T5 wrote the correct average, 422,764, in prose and stopped without the FINAL line.
Two lessons transfer to any agent, whichever model you pick:
- Compute aggregates in code. Most misses happened while a model added or compared 10 to 20 numbers in its head. Have the tool return the sum, mean, ranking or filtered list, or give the agent a calculator tool. The cost is a few lines of Python.
- Name the source of every number. “Per profile” wasn’t enough for two models when search results already showed a follower count. If the number must come from a specific call, say so in the tool description, or don’t return the competing number at all.
Cost: same accuracy at very different prices
Cost per task here is the token usage each API reported, multiplied by the list price above. It is not a billing record, and we ran without prompt caching. The whole study (675 runs, including the 4,000-token reruns) used about $7.54 of LLM tokens at list price; a short pilot before it ($0.29, T1/T5/T9 only) is excluded from every table. The six data endpoints showed base_price 0 (Free) in GET /v1/models/<id> on 2026-10-03.

Caption: MiniMax M3 is listed at $0.30 input and $1.20 output per million tokens on SandBase, one of the two cheapest rates in this test (captured 2026-10-03).

Caption: Xiaomi MiMo v2.6 Pro is listed at $0.435 input and $0.87 output per million tokens, the lowest output price of the 12 models (captured 2026-10-03).
At the medians, 1,000 tasks a day like these would cost about $1.80 on MiMo v2.6 Pro, $1.90 on MiniMax M3, $6.90 on GPT-6.1 Sol, $15.90 on Kimi K3 and $21.30 on Qwen3.8 Max Prime. That’s an extrapolation from 51 short runs per model, but the spread is wide enough that it survives a lot of noise. MiMo and MiniMax came in at roughly a quarter of GPT-6.1 Sol’s median cost per task.
The price sheet didn’t predict the order. Grok 4.7 lists below GPT-6.1 Sol on output ($6 vs $10) yet cost twice as much per task, because its median input was 5,718 tokens against 2,887. We didn’t isolate where Grok’s extra input comes from. Output differed even more: GPT-6.1 Sol’s median was 110 output tokens per task, while the other 11 ranged from 279 (Grok 4.7) to 620 (MiniMax M3). GPT-6.1 Sol gave terse answers. For models that report reasoning tokens, those are included in the completion count as reported in usage.
Latency medians ran from 8.4 s (Qwen3.8 Max Prime) to 21.5 s (GLM-5.3). Claude Sonnet 5.5 had the one big outlier, a 259 s H3 run; the first article’s pilot also saw a 259 s Sonnet run. Times include the SandBase gateway and a cache lookup of a few milliseconds per tool call, from one machine.
How to choose
| If you need | Start with | But |
|---|---|---|
| Accuracy first on short tool tasks, at moderate cost | GPT-6.1 Sol: 51/51 at $0.0069 per task | Its outputs are terse; ask for explanations if you need them |
| Accuracy first, a second or third option | Kimi K3, Grok 4.7 or Qwen3.8 Max Prime: all 51/51 | 2x to 3x GPT-6.1 Sol’s cost per task here |
| Lowest cost | MiMo v2.6 Pro (50/51) or MiniMax M3 (48/51) at about $0.002 per task | Give them a larger output budget and do arithmetic in code; MiniMax also missed “per profile” twice |
| Another option at about GPT-6.1 Sol’s cost | DeepSeek V4 Pro: same $0.0069 median | Two instruction misses on the easy set (source, ID typo) |
| Claude-specific features | Claude Sonnet 5.5 | Lowest score here (44/51); all misses were hand arithmetic, so move that into tools first |
This isn’t a general ranking. It covers one task family, Chinese social data lookups plus aggregation, with N=3 and default settings. A model that truncated here might do fine with max_tokens raised or reasoning effort lowered, and none of this tells you about coding or long documents.
How to run the same kind of test on SandBase
One SandBase API key reaches all 12 models and the data tools. Eleven of them use the same OpenAI-compatible Chat Completions contract, so switching models means changing one string. SandBase lists these data operations on its catalog pages as GET /apis/v1/<vendor>/..., while this test uses the Model API route from each endpoint’s API reference, POST https://api.sandbase.ai/v1/api/<vendor>/<path>; don’t mix the two.

Caption: The SandBase Chat Completions guide documents POST /v1/chat/completions and the tool_calls / tool_call_id loop that the 11 OpenAI-compatible models in this test used (captured 2026-10-03).
The sample below applies both lessons. The tool computes the like statistics in Python and returns only aggregates, and the loop raises on finish_reason: length instead of failing silently. The payload field names (aweme_list, statistics.digg_count) are observed in our 2026-10-03 calls, not documented guarantees, so they’re read with .get().
import os
import json
import requests
BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
# Any OpenAI-compatible model on SandBase; only the model id changes.
MODELS = ["openai/gpt-6.1-sol", "minimax/minimax-m3", "xiaomi/mimo-v2.6-pro"]
def call_data_api(model: str, params: dict) -> dict:
"""Call a SandBase data endpoint and return outputs[0].data, failing loudly otherwise."""
resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
resp.raise_for_status()
body = resp.json()
outputs = body.get("outputs") or []
if body.get("status") != "completed" or not outputs:
raise RuntimeError(f"SandBase call did not complete: {body.get('status')} {body.get('error')}")
return outputs[0].get("data") or {}
def douyin_video_stats(sec_user_id: str) -> dict:
"""First page of an account's videos, with the aggregates computed here, not by the model."""
data = call_data_api("douyin/app-v3/user-post-videos",
{"sec_user_id": sec_user_id, "max_cursor": 0, "count": 20})
likes = [(item.get("statistics") or {}).get("digg_count") or 0 # observed field names
for item in data.get("aweme_list") or []]
if not likes:
return {"videos": 0, "found": False}
mean = sum(likes) / len(likes)
return {"videos": len(likes), "likes_sum": sum(likes), "likes_mean_floor": sum(likes) // len(likes),
"likes_max": max(likes), "above_mean": sum(x > mean for x in likes)}
TOOLS = [{"type": "function", "function": {
"name": "douyin_video_stats",
"description": "Like statistics for the first page (up to 20) of a Douyin account's videos.",
"parameters": {"type": "object", "properties": {"sec_user_id": {"type": "string"}},
"required": ["sec_user_id"]}}}]
SYSTEM = "Use the tools to answer; never guess numbers. Finish with one line: FINAL: {json}."
def run(model: str, task: str, max_tokens: int = 4000) -> str:
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": task}]
for _ in range(10):
resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=240,
json={"model": model, "max_tokens": max_tokens,
"tools": TOOLS, "messages": messages})
resp.raise_for_status()
body = resp.json()
choice = body["choices"][0]
msg = choice["message"]
messages.append({k: v for k, v in msg.items() if k in ("role", "content", "tool_calls")})
if choice.get("finish_reason") == "length":
raise RuntimeError(f"{model}: output cut at max_tokens={max_tokens}")
calls = msg.get("tool_calls") or []
if not calls:
return msg.get("content") or ""
for call in calls:
args = json.loads(call["function"].get("arguments") or "{}")
messages.append({"role": "tool", "tool_call_id": call["id"],
"content": json.dumps(douyin_video_stats(**args))})
raise RuntimeError("turn limit reached")
if __name__ == "__main__":
task = ("For the Douyin account with sec_user_id "
"MS4wLjABAAAA8U_l6rBzmy7bcy6xOJel4v0RzoR_wfAubGPeJimN__4, what is the average like count "
"on the first page, rounded down? Keys: videos, average_likes.")
for model in MODELS:
print(model, run(model, task).strip().splitlines()[-1])
Tested on 2026-10-03 (UTC). We ran the program above verbatim from 13:59 to 14:01 UTC. Input: the public People’s Daily Douyin account used in T2, T5 and T7. Each model’s tool call sent POST https://api.sandbase.ai/v1/api/douyin/app-v3/user-post-videos with {"sec_user_id": "MS4wLjABAAAA8U_l6rBzmy7bcy6xOJel4v0RzoR_wfAubGPeJimN__4", "max_cursor": 0, "count": 20}. Trimmed response from the GPT-6.1 Sol run’s call:
{"id": "14d98563-da7d-45b4-900a-aa1c4e3c6b92", "status": "completed",
"model": "douyin/app-v3/user-post-videos",
"outputs": [{"data": {"aweme_list": [
{"aweme_id": "7692418769125199138", "statistics": {"digg_count": 9522}},
{"aweme_id": "7692417622113111323", "statistics": {"digg_count": 3843}}],
"has_more": 1}}]}
First 2 of 20 items in aweme_list shown; inside each item only aweme_id and statistics.digg_count are kept, and all other keys are omitted. The envelope (id, status, model, outputs[0].data) is documented; the fields inside data are names we observed in this response, not documented ones.
The program printed:
openai/gpt-6.1-sol FINAL: {"videos":20,"average_likes":403206}
minimax/minimax-m3 FINAL: {"videos": 20, "average_likes": 403251}
xiaomi/mimo-v2.6-pro FINAL: {"videos": 20, "average_likes": 403304}
All three models returned the likes_mean_floor their own tool call produced. The three numbers differ because each run made a live call and the like counts were still climbing, which is exactly why the benchmark replayed a cache instead. GET /v1/tasks/<id>/cost billed the two LLM calls per model at $0.001608 for GPT-6.1 Sol, $0.000461 for MiniMax M3 and $0.000349 for MiMo v2.6 Pro, and each data call $0.000000. That endpoint and GET /v1/models/<id> need the same Authorization: Bearer key as the calls; the public model pages show prices without one.
The request fields are in the Douyin user-post-videos API reference. Get a SandBase API key to run the same loop with your own model list.
Scope of the data: these are public, read-only lookups that need a SandBase API key. SandBase isn’t an official Douyin, Weibo or Xiaohongshu partner, and none of this touches private accounts, DMs, owner analytics or account actions. If you want a fuller agent built on the same pattern, the product research agent tutorial computes its stats in Python and has GPT-6.1 Sol write only the summary. For background on one of the stronger models here, see our Kimi K3 overview.
FAQ
Which LLM is best for agent tool calling in 2026?
On this test, GPT-6.1 Sol, Grok 4.7, Kimi K3 and Qwen3.8 Max Prime all scored 51/51, and GPT-6.1 Sol was the cheapest of the four at $0.0069 per task. That’s one task family with three runs per task, so treat it as a shortlist for lookup-and-aggregate agents, not a universal ranking.
Are Chinese LLMs good at agent tool calling?
Yes, on these tasks. Kimi K3 and Qwen3.8 Max Prime were perfect, MiMo v2.6 Pro missed one run, and GLM-5.3, GLM-5.3 Prime, Qwen3.8 Max 0902 and DeepSeek V4 Pro scored 49/51. Every model, Chinese or not, made zero schema-invalid tool calls.
Kimi K3 vs GPT-6.1 Sol: which should I use?
Both scored 51/51. GPT-6.1 Sol’s median cost was $0.0069 per task against $0.0159 for Kimi K3, mostly because Kimi wrote about 410 output tokens per task at $15 per million while GPT-6.1 Sol wrote 110 at $10. Kimi K3 was slightly faster at the median (8.5 s vs 9.5 s).
What is the cheapest model that still works for tool-calling agents?
MiMo v2.6 Pro scored 50/51 at a median $0.0018 per task, and MiniMax M3 scored 48/51 at $0.0019. Both lost an H2 run to the 1,500-token cap, and those reruns passed at 4,000 tokens; MiniMax also took T4’s follower counts from search results twice. Raise the output budget and keep arithmetic in code.
Why did Claude Sonnet 5.5 score lowest?
All seven of its misses were arithmetic or counting over tool output (two averages, a date count, three over-long filtered lists, a sum off by 10,000). It never truncated and never made an invalid call. Moving the arithmetic into the tool removes the operation it got wrong.
Limitations
This is one task family: Chinese social data lookups plus aggregation, with 17 scored tasks and three runs each. We used default temperature and reasoning settings, a 1,500-token output cap per turn, no prompt caching, and one machine for timing. Some misses were truncations that a larger budget fixed, so the scores partly measure verbosity at that cap. H7 was excluded because our own task was ambiguous. Twenty-two misses across 612 scored runs is too few to rank the models in the 48 to 50 range against each other with confidence. Models update often, so rerun the tasks that look like your workload before you commit.