Blog/Model Comparison/

Claude Sonnet 5.5 vs GPT-6.1 Sol: Agent Test

Claude Sonnet 5.5 vs GPT-6.1 Sol on 10 tool-calling agent tasks, 3 runs each: accuracy, cost per task, latency and token overhead at the same $2/$10 price.

Claude Sonnet 5.5 vs GPT-6.1 Sol tool-calling agent test cover; editorial artwork, not a test result

Task T5 gave both models the same 20 Douyin videos and asked for the average like count. GPT-6.1 Sol returned 422,764 in all three runs, which matches the value our code computed. Claude Sonnet 5.5 returned 422,764 once, then 416,992, then 434,643.

That was the clearest difference in this Claude Sonnet 5.5 vs GPT-6.1 Sol test. If you’re picking a default model for a tool-calling agent and both list at $2 per million input tokens and $10 per million output tokens, the price sheet won’t settle it. So we ran 10 agent tasks against the same six data tools, three times per model, through SandBase on 2026-10-03 (UTC), and graded every answer by exact match.

Key takeaway

  • GPT-6.1 Sol answered 30 of 30 runs correctly. Claude Sonnet 5.5 answered 27 of 30. All three misses were arithmetic or counting over tool output (two averages, one date count).
  • Both models picked the same tools with the same number of calls on every task. Tool selection did not separate them.
  • GPT-6.1 Sol cost about half as much per task (median $0.0066 vs $0.0123) and finished faster (median 6.7 s vs 11.1 s). Most of the gap came from tokens: Sonnet 5.5 used about 1.7x the input and 3x the output.
  • Scope: 10 tasks, 3 runs each, Chinese social data tools, default settings, one time window. This is evidence for short structured agent tasks, not a general ranking.

Why compare Claude Sonnet 5.5 and GPT-6.1 Sol now

Three launches landed within three days. Anthropic released Claude Sonnet 5.5 on September 28, 2026, describing it as a faster, lower-cost complement to Claude Opus 5.5 that is strongest at well-scoped everyday tasks. Its list price stays at $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5.

OpenAI showed GPT-6.1 Sol at DevDay on September 29 (recap, TechCrunch coverage). OpenAI’s model page pitches near-Astra performance at a lower cost, and its pricing page lists gpt-6.1-sol at $2 input and $10 output, one fifth of GPT-6 Astra’s $10 and $50.

Google announced Gemini 4 Argon on September 30 with the same $2/$10 introductory price, but it is rolling out to a limited group through Google’s Fairwind program and is not on SandBase. We didn’t test it.

So two generally available models share a price tag. The question for an agent builder is narrower than “which model is smarter”: which one gets a short tool-calling job right, and what does each correct answer cost?

SandBase language model catalog filtered to sonnet-5.5, listing Anthropic Claude Sonnet 5.5 with 1M context, $2 input and $10 output per million tokens

Caption: The SandBase model catalog lists anthropic/claude-sonnet-5.5 at $2 input and $10 output per million tokens, the rates our cost figures use (captured 2026-10-03).

The catalog row above matches the model_card returned by GET /v1/models/anthropic/claude-sonnet-5.5, which we checked on 2026-10-02 and again on 2026-10-03. You can open the Claude Sonnet 5.5 model page and the GPT-6.1 Sol model page on SandBase to run either model.

Test setup: 10 tasks, 6 tools, exact-match grading

Each model got the same system prompt, the same six tool definitions and the same task text. Sonnet 5.5 ran through SandBase POST /v1/messages (Anthropic Messages format, tools as input_schema). GPT-6.1 Sol ran through SandBase POST /v1/chat/completions (OpenAI function tools). Each turn allowed up to 1,500 output tokens, each task up to 10 turns, and temperature and reasoning settings stayed at their defaults.

The six tools wrap SandBase public data endpoints:

ToolSandBase endpoint
douyin_user_searchdouyin/search/user-search-v2
douyin_user_profiledouyin/app-v3/user-profile
douyin_user_videosdouyin/app-v3/user-post-videos
weibo_hot_searchweibo/web-v2/hot-search
weibo_searchweibo/web-v2/realtime-search
xhs_search_notesxiaohongshu/app-v2/search-notes

Live social data changes by the minute, which would make a fair comparison impossible. So the harness called each endpoint once, between 02:50 and 02:58 UTC on 2026-10-03, stored the trimmed response keyed by tool name and arguments, and replayed that cache to both models. The expected answers came from the same cache, computed by plain Python: filter, max, sum(...) // len(...), string matching. Neither model’s output played any part in the answer key.

TaskWhat the agent had to doWhat it tests
T1Find the account named exactly 瑞幸咖啡 (Luckin Coffee) and read its follower count from the profileSearch, pick the right account among look-alikes, chain to profile
T2Find the non-pinned video with the most likes on People’s Daily’s first pageFilter plus max
T3Count Weibo hot-search topics (top 30) that contain the name of table tennis player Wang ChuqinSubstring counting
T4Compare followers of Luckin Coffee and CottiCoffee and give the differenceTwo searches, two profiles, subtraction
T5Average likes across all 20 first-page videos, rounded downArithmetic over 20 numbers
T6Top Xiaohongshu note by likes for sunscreenMax over a list
T7Count first-page videos created on or after 2026-10-01Date filtering
T8Take the #1 hot-search topic, search Weibo for it, report the post countChained calls
T9Look up a non-existent account and report found=falseSaying “not found” instead of guessing
T10For camping, report the top Xiaohongshu note’s likes and the top Douyin account’s followersTwo tools, two maxima

Every answer had to end with one FINAL: line holding a JSON object with exactly the requested keys. Grading was strict: every key had to match the code-computed value. Commas in numbers and "true" vs true were normalized; a number off by one was a miss.

The system prompt, verbatim, was the same for both models:

You are a research agent with tools for public Douyin, Weibo and Xiaohongshu data. Use the tools to answer; never guess numbers. If something can't be found, say so. Finish with ONE line that starts with FINAL: followed by a compact JSON object with exactly the keys requested.

The grader parsed the JSON object after FINAL:, then compared each requested key with the expected value: integers after stripping commas, booleans after lowercasing, IDs and names as exact strings. A missing key (including answers nested under an extra wrapper) or unparseable JSON counted as a miss; extra keys were ignored. We are not publishing the cached responses themselves, because they are third-party platform content about real accounts. The task table above, this prompt and the grading rule are enough to rebuild the test against your own fresh snapshot, which is what we’d recommend anyway, since your numbers will be different.

One honest caveat on T2: none of the 20 videos captured for that account was pinned, so the “non-pinned” filter wasn’t actually exercised. T2 ended up testing max-finding only.

Results: accuracy, cost and latency

Metric (30 runs per model)Claude Sonnet 5.5GPT-6.1 Sol
Correct answers27 / 3030 / 30
Tool calls / invalid calls48 / 048 / 0
Median cost per task$0.0123$0.0066
Mean cost per task$0.0128$0.0063
Total for 30 runs$0.385$0.189
Median input tokens per task4,9302,838
Median output tokens per task24277
Median wall time11.1 s6.7 s
Slowest run27.0 s19.1 s

Both models made exactly 48 tool calls, all with valid arguments, and the per-task tool sequence was identical: two calls for T1, four for T4, one for T5 and so on. Neither model skipped a lookup or invented a number when the tool returned nothing. On T9 both said the account didn’t exist in all six runs.

TaskSonnet 5.5 correctGPT-6.1 Sol correctSonnet median timeGPT median timeSonnet median costGPT median cost
T13/33/317.2 s9.4 s$0.0157$0.0069
T23/33/310.9 s6.5 s$0.0123$0.0066
T33/33/310.3 s5.3 s$0.0086$0.0035
T43/33/317.3 s11.7 s$0.0239$0.0117
T51/33/311.7 s9.6 s$0.0129$0.0081
T63/33/39.2 s5.5 s$0.0083$0.0037
T72/33/311.2 s6.7 s$0.0125$0.0065
T83/33/314.5 s8.5 s$0.0144$0.0070
T93/33/39.8 s5.6 s$0.0082$0.0034
T103/33/39.4 s6.1 s$0.0109$0.0054

GPT-6.1 Sol was cheaper and faster on every one of the 10 tasks. The latency includes the SandBase gateway and, for tool turns, a cache lookup that takes milliseconds, so it reflects model time plus routing rather than data-API time.

Where Sonnet 5.5 lost points

All three Sonnet 5.5 misses were in the two tasks that ask the model to do bookkeeping over a 20-item list.

T5, the average. The 20 like counts in the cached page sum to 8,455,283, so the floor of the average is 422,764. Sonnet 5.5 got that in run 1. Run 2 answered 416,992, which implies a total about 115,000 too low. Run 3 answered 434,643, which implies a total about 238,000 too high. The list contained values from 4,676 up to 2,780,636, the kind of mixed magnitudes where a dropped or doubled item is easy to miss in mental arithmetic. GPT-6.1 Sol returned 422,764 in all three runs.

T7, the date count. 19 of the 20 videos were created on October 1 or 2; the last one in the list was dated September 30. Sonnet 5.5 answered 19 twice and 18 once.

The pattern points to a fix that works for any model: an LLM shouldn’t add up 20 numbers when your code can. If the agent needs an average or a count, have the tool return it (a like_stats field with count, sum and mean) or expose a small calculator tool. We didn’t rerun the benchmark with that change, so we can’t show the improvement here, but it removes the exact operation that failed.

The pilot run, before the scored run, had the same shape. Sonnet 5.5 missed T5 once (409,542) and had two very slow runs, 225 s and 259 s. GPT-6.1 Sol missed T10 because the prompt then said “top note”: it reported 7 likes, while the note with the most likes had 1,006. We rewrote T10 to say “the note with the most likes” for both models before the formal run. The pilot used one run per task, earlier prompt wording, and is excluded from every number in the tables above.

Cost and latency: where the 2x comes from

Both models bill the same rates, so the cost gap is a token gap. Two things drive it.

First, fixed overhead per request. With our system prompt, the six tool definitions and a trivial user message (“Reply with OK.”), Sonnet 5.5 through /v1/messages reported 1,036 input tokens, and GPT-6.1 Sol through /v1/chat/completions reported 336 (rechecked on 2026-10-03). An agent pays that on every turn because the full context is resent. A three-turn task like T1 carries roughly 2,100 extra input tokens for Sonnet before any tool output is counted. That accounts for about 70% of the 2,992-token input gap we saw on T1. The rest comes from tool results and conversation history as counted by each provider’s own tokenizer; we didn’t isolate that part.

Second, output. Sonnet 5.5 wrote a median 242 output tokens per task against 77 for GPT-6.1 Sol, for the same answers and the same tool calls. At $10 per million output tokens, that difference adds up. GPT-6.1 Sol’s output count includes its reasoning tokens (the usage object reports them separately), and it still came out lower.

SandBase language model catalog filtered to gpt-6.1-sol, listing OpenAI GPT-6.1 Sol with 1.1M context, $2 input and $10 output per million tokens

Caption: The same catalog lists openai/gpt-6.1-sol at $2 input and $10 output per million tokens, so the cost difference in this test is token usage, not price (captured 2026-10-03).

To make that concrete: at 1,000 tasks a day with this task mix, the medians work out to about $12.30 a day for Sonnet 5.5 and $6.60 for GPT-6.1 Sol, or roughly $370 and $200 over 30 days. That’s an extrapolation from 30 short runs per model, without prompt caching. Both providers discount cached input, and a long, stable system prompt plus tool list is exactly what caching helps with, so measure with caching on before you budget.

The whole experiment, pilot plus formal run, used $0.77 of LLM tokens ($0.57 formal, $0.19 pilot). The data endpoints were marked Free in the SandBase catalog when we checked on 2026-10-02.

Which should you choose?

Choose GPT-6.1 Sol whenChoose Claude Sonnet 5.5 when
Your agent runs short, structured tool-calling jobs: look up, filter, compare, reportThe work is documents, slides, spreadsheets or bug fixing, which Anthropic names as Sonnet 5.5’s strengths (we didn’t test these)
Cost per task matters at volume; here it was about halfYour stack is already built on the Anthropic Messages API and you want to stay on one protocol
Latency is user-facing; median wall time was about 40% lower hereYou need the long-horizon work or image understanding Anthropic highlights for Sonnet 5.5 (vendor claims, untested here)
You can’t move aggregation into code yetYou’ve moved arithmetic into tools, which removes the failure mode we saw

For short, structured tool-calling tasks like these, we’d start with GPT-6.1 Sol as the default and keep Sonnet 5.5 as a candidate for the document- and code-heavy work this test didn’t cover. Whichever you pick, compute aggregates in code. That change is cheap and protects you against the one class of error we observed.

If you’re designing the tools themselves, the sync-only data API design post explains why one blocking call per tool keeps agent turns short, and the Douyin competitor monitor agent shows a full agent built on the same Douyin endpoints.

A one-tool example on SandBase

One SandBase API key reaches both models and the data tools. SandBase exposes more than one surface for these operations: the catalog pages list them as GET /apis/v1/<vendor>/..., while this test uses the Model API route from each endpoint’s API reference. Don’t mix the two. The tool executor calls the data API at POST https://api.sandbase.ai/v1/api/<model> and reads only the documented outputs[0].data. The payload field names below (aweme_list, statistics.digg_count, is_top, create_time) come from our 2026-10-03 calls. They’re observed fields, not documented guarantees, so read them defensively. The code below is a one-tool slice of the harness (the douyin_user_videos tool behind T2, T5 and T7), not the full six-tool benchmark.

SandBase API reference for douyin/app-v3/user-post-videos showing the POST route, the count, max_cursor, sec_user_id and sort_type fields, and the outputs data envelope

Caption: The user-post-videos reference documents POST /v1/api/douyin/app-v3/user-post-videos, a count of at most 20 and max_cursor 0 for the first page, which is the request our T2, T5 and T7 tool sends (captured 2026-10-03).

import os
import json
import datetime
import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}


def call_data_api(model: str, params: dict) -> dict:
    """Call a SandBase data endpoint and return outputs[0].data, failing loudly otherwise."""
    resp = requests.post(f"{BASE}/api/{model}", headers=HEADERS, json=params, timeout=120)
    resp.raise_for_status()
    body = resp.json()
    outputs = body.get("outputs") or []
    if body.get("status") != "completed" or not outputs:
        raise RuntimeError(f"SandBase call did not complete: {body.get('status')} {body.get('error')}")
    return outputs[0].get("data") or {}


def douyin_user_videos(sec_user_id: str) -> dict:
    data = call_data_api("douyin/app-v3/user-post-videos",
                         {"sec_user_id": sec_user_id, "max_cursor": 0, "count": 20})
    videos = []
    for item in data.get("aweme_list") or []:  # observed field names, read defensively
        stats = item.get("statistics") or {}
        created = datetime.datetime.fromtimestamp(item.get("create_time") or 0, datetime.timezone.utc)
        videos.append({"aweme_id": item.get("aweme_id"), "create_time": created.strftime("%Y-%m-%d"),
                       "pinned": bool(item.get("is_top")), "likes": stats.get("digg_count")})
    likes = [v["likes"] or 0 for v in videos]
    # Do the arithmetic here so the model never has to.
    like_stats = {"count": len(likes), "sum": sum(likes), "mean_floor": sum(likes) // len(likes) if likes else None}
    return {"videos": videos, "like_stats": like_stats}


SCHEMA = {"type": "object", "properties": {"sec_user_id": {"type": "string"}}, "required": ["sec_user_id"]}
DESC = "Get the first page (up to 20) of a Douyin account's videos by sec_user_id, plus like_stats."
SYSTEM = "Use the tools to answer; never guess numbers. Finish with one line: FINAL: {json}."


def run_sonnet(task: str) -> str:
    tools = [{"name": "douyin_user_videos", "description": DESC, "input_schema": SCHEMA}]
    messages = [{"role": "user", "content": task}]
    for _ in range(10):
        resp = requests.post(f"{BASE}/messages", headers={**HEADERS, "anthropic-version": "2023-06-01"},
                             json={"model": "anthropic/claude-sonnet-5.5", "max_tokens": 1500,
                                   "system": SYSTEM, "tools": tools, "messages": messages}, timeout=240)
        resp.raise_for_status()
        content = resp.json().get("content") or []
        messages.append({"role": "assistant", "content": content})
        uses = [c for c in content if c.get("type") == "tool_use"]
        if not uses:
            return "".join(c.get("text", "") for c in content if c.get("type") == "text")
        messages.append({"role": "user", "content": [
            {"type": "tool_result", "tool_use_id": u["id"],
             "content": json.dumps(douyin_user_videos(**u["input"]), ensure_ascii=False)} for u in uses]})
    raise RuntimeError("turn limit reached")


def run_gpt(task: str) -> str:
    tools = [{"type": "function", "function": {"name": "douyin_user_videos", "description": DESC, "parameters": SCHEMA}}]
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": task}]
    for _ in range(10):
        resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS,
                             json={"model": "openai/gpt-6.1-sol", "max_tokens": 1500,
                                   "tools": tools, "messages": messages}, timeout=240)
        resp.raise_for_status()
        msg = resp.json()["choices"][0]["message"]
        messages.append({k: v for k, v in msg.items() if k in ("role", "content", "tool_calls")})
        calls = msg.get("tool_calls") or []
        if not calls:
            return msg.get("content") or ""
        for call in calls:
            args = json.loads(call["function"].get("arguments") or "{}")
            messages.append({"role": "tool", "tool_call_id": call["id"],
                             "content": json.dumps(douyin_user_videos(**args), ensure_ascii=False)})
    raise RuntimeError("turn limit reached")

Tested on 2026-10-03 (UTC). At 14:07 UTC we loaded the code above unchanged and called run_sonnet and run_gpt once each with the T5-style question “For the Douyin account with sec_user_id MS4wLjABAAAA8U_l6rBzmy7bcy6xOJel4v0RzoR_wfAubGPeJimN__4, what is the average like count on the first page, rounded down? Keys: videos, average_likes.”. Input: the public People’s Daily Douyin account used in T2, T5 and T7. Each tool call sent POST https://api.sandbase.ai/v1/api/douyin/app-v3/user-post-videos with {"sec_user_id": "MS4wLjABAAAA8U_l6rBzmy7bcy6xOJel4v0RzoR_wfAubGPeJimN__4", "max_cursor": 0, "count": 20}. Trimmed response from the Sonnet run’s call:

{"id": "2372254c-8d24-4f1c-a7f3-4c777d5df7a3", "status": "completed",
 "model": "douyin/app-v3/user-post-videos",
 "outputs": [{"data": {"aweme_list": [
   {"aweme_id": "7692418769125199138", "statistics": {"digg_count": 10344}},
   {"aweme_id": "7692417622113111323", "statistics": {"digg_count": 4090}}],
  "has_more": 1}}]}

First 2 of 20 items in aweme_list shown; inside each item only aweme_id and statistics.digg_count are kept, and all other keys (including create_time and is_top) are omitted. The envelope is documented; the fields inside data are observed, not documented.

Sonnet 5.5 answered FINAL: {"videos": 20, "average_likes": 404603} and GPT-6.1 Sol answered FINAL: {"videos":20,"average_likes":404670}. Each matched the like_stats.mean_floor of its own tool call; the two differ because each made a live call a few seconds apart while likes were still coming in. GET /v1/tasks/<id>/cost billed Sonnet’s two turns $0.006404 in total and GPT-6.1 Sol’s two turns $0.003716, and each data call $0.000000. For POST /v1/messages, the msg_... id in the body didn’t resolve in that endpoint (404), but the x-task-id response header did. Both cost lookups and GET /v1/models/<id> need the same Authorization: Bearer key; the public model pages show prices without one.

The request fields are in the Douyin user-post-videos API reference. Get a SandBase API key to run both loops on your own tasks.

Our 12-model tool-calling benchmark reran these same ten tasks and documents the method, task list, grading rule and per-model aggregates; it does not publish the raw cached responses or the full harness either.

Scope of the data: these are public, read-only lookups that need a SandBase API key. SandBase isn’t an official Douyin, Weibo or Xiaohongshu partner, and none of this touches private accounts, DMs, owner analytics or account actions.

Two protocol notes. The tool schema is the same JSON Schema on both sides; only the wrapper differs (input_schema for Messages, function.parameters for Chat Completions). And if you call OpenAI directly rather than through SandBase, OpenAI’s model page says to use the Responses API for GPT-6.1 Sol tool calling, with Chat Completions supported without tools. Through SandBase /v1/chat/completions, tool calls worked in all 30 of our runs on 2026-10-03, but check the current docs before you depend on that path. For a longer agent loop with the OpenAI SDK, see build a social monitor agent with the OpenAI SDK.

FAQ

Is GPT-6.1 Sol better than Claude Sonnet 5.5 for agents?

On our 10 short tool-calling tasks, yes: 30 of 30 correct against 27 of 30, at about half the cost per task. That’s one task set with three runs each, so treat it as a strong signal for lookup-and-report agents rather than a verdict on every agent workload.

Do Claude Sonnet 5.5 and GPT-6.1 Sol cost the same?

The list prices match: $2 per million input tokens and $10 per million output tokens on both the vendor pages and SandBase. Cost per task didn’t match. Sonnet 5.5 used a median 4,930 input and 242 output tokens per task against 2,838 and 77. The median cost per task, computed from each run’s own usage × list price and then taking the median of those 30 values, was $0.0123 against $0.0066 (the median of costs is not the cost of the median token counts).

Why did Sonnet 5.5 use more input tokens?

Mostly fixed overhead. The same system prompt and six tool definitions counted as 1,036 input tokens through /v1/messages and 336 through /v1/chat/completions, and agents resend that on every turn. Prompt caching would shrink this for both models; we ran without it.

Which model picks tools more reliably?

We couldn’t separate them. Both made 48 tool calls with zero invalid arguments, and both chose the same tools in the same order on every task. The differences showed up after the tools returned, when a model had to add or count.

Can I switch between the two without rewriting my agent?

The tool logic and JSON Schemas carry over. The request and response wrappers differ between the Messages and Chat Completions formats, as the code above shows, so keep the protocol layer thin and swap it per model.

Limitations

This is a small test. Ten tasks, three runs each, one tool set built on Chinese social platforms, one 2026-10-03 (UTC) data snapshot, and default temperature and reasoning effort. We used no prompt caching and no extended thinking. Latency was measured from one machine through the SandBase gateway, so your numbers will differ by region and load. Three misses out of 30 is too few to estimate Sonnet 5.5’s error rate with any precision. And both models are days old; updates on either side can change these results, so rerun the tasks that look like your workload before you commit.