Claude Sonnet 5.5 coding benchmark results are impressive when the test measures more than a repository patch or a one-line programming puzzle. In a new eight-task KingBench 3 evaluation, the model produced simulations, games, 3D objects, a computational solver, and a locally fine-tuned model workflow, earning 71.5 out of 80—or 89.38%—on the reviewer’s comparative scale.
That headline score is useful, but it is not the most interesting takeaway. The real story is that Sonnet 5.5 appears strong at the difficult middle ground of agentic coding: taking an ambiguous visual or interactive request, creating project files, executing commands, testing what it built, and delivering something a person can inspect in a browser. For founders, creators, and developers, that is closer to everyday AI-assisted product work than a narrow code-completion benchmark.
What the Claude Sonnet 5.5 coding benchmark tested
The original video evaluates Claude Sonnet 5.5 through OpenCode on eight KingBench 3 prompts, each worth 10 points. The run used fresh folders, text-only prompts, high reasoning effort, and a generous 128,000-token output ceiling. The reviewer treated Anthropic’s higher-tier Opus 5.5 result as the quality reference, then scored each completed artifact for both functional behavior and presentation.
That setup matters. This was not an official Anthropic evaluation, nor was it a statistically rigorous leaderboard with hundreds of repeated trials. It was a hands-on editorial test of eight outputs. The scores therefore reflect one reviewer’s judgment about quality, visual finish, and whether the project actually worked under interaction.
Still, it is a useful format because the task mix is broad:
- A multi-elevator browser simulation.
- An interactive Three.js contact lens case.
- A slider-controlled folding table in Three.js.
- An SVG panda eating a burger.
- A bow-and-arrow browser game.
- A constrained 60-point counting problem.
- A local Gemma 2B LoRA fine-tuning workflow and web app.
- A functional dual-time-zone 3D wristwatch.
The benchmark asks an AI coding agent to move between visual design, JavaScript interaction, geometry, algorithmic reasoning, debugging, shell workflows, machine-learning tooling, and lightweight product design. That breadth is precisely why the result is more interesting than a simple “can it write a function?” test.
OpenCode is relevant here because it gives models an agentic environment rather than a blank chat box: it is an open-source coding agent available in a terminal interface, desktop application, and IDE extension, with the ability to work with files and tools. (opencode.ai) The model was not merely asked to emit code; it could participate in the create-run-inspect-revise loop that defines practical AI coding work.
The final score: 71.5 out of 80, with context
Sonnet 5.5 received these individual scores in the test:
- Elevator simulation: 8/10
- Contact lens case: 9.5/10
- Folding table: 8/10
- Panda SVG: 9/10
- Archery game: 8.5/10
- Counting problem: 10/10
- Local fine-tuning workflow: 9/10
- 3D wristwatch: 9.5/10
That adds up to 71.5/80, or 89.38%. In the same reviewer’s earlier comparison, Opus 5.5 received 75/80, or 93.75%, leaving Sonnet behind by 4.38 percentage points. The gap is not trivial, but it is also not a blowout. The more affordable model was close enough that the choice will often depend on task scope, iteration volume, and the cost of a human reviewer rather than on raw capability alone.
Anthropic positions Sonnet 5.5 as the quicker, lower-cost companion to Opus 5.5, with Sonnet aimed at well-scoped everyday tasks, bug fixing, and polished work products while Opus is positioned for work requiring more careful judgment. (anthropic.com) The KingBench result broadly supports that framing: Sonnet looked particularly convincing when the request was concrete, demonstrable, and bounded by visible acceptance criteria.
The score should not be interpreted as “Sonnet will complete 89.38% of your software project.” A benchmark percentage is a summary of this particular task set and this reviewer’s scoring choices. It is better read as evidence that Sonnet can deliver many credible first versions of interactive applications, rather than as a guarantee of engineering reliability in a production codebase.
Why interactive coding tasks are a tougher test than code generation
Modern model comparisons often rely on coding benchmarks that evaluate bug fixes, test passing, or algorithmic solutions. Those measures are valuable, especially for repository-level engineering. But interactive builds introduce other failure modes that ordinary unit tests do not capture.
A browser game can technically launch while still feeling broken. A 3D model can include all named components while looking implausible. A dashboard can have every requested field but bury the user’s primary action. A training script can finish without proving that the deployed app is actually using the trained model.
Interactive tasks test several layers at once:
Requirement interpretation
The agent must translate natural-language details into an implementation plan. “Each elevator carries one person” is not merely styling; it changes queue management, scheduling, occupancy state, and movement logic. “Caps open independently” requires separate interaction states rather than a single animation toggle.
Implementation and tool use
The model must write files, use libraries correctly, run commands, inspect errors, and revise the result. This is where coding agents differ from traditional chat-based code generation: the agent can validate assumptions against execution rather than leaving all debugging to the user.
Product behavior
The output must work when people click, drag, retry, refresh, alter settings, or create edge-case conditions. A model that produces a perfect screenshot but fails under normal use has not solved the task.
Visual coherence
For creator-facing and customer-facing work, functional correctness is only half the job. Geometry intersections, odd spacing, weak hierarchy, unreadable controls, and unfinished surfaces all lower the value of an otherwise working prototype.
This is why the video’s approach—interacting with the finished work rather than only reading source code—is valuable. It looks at the artifact a stakeholder would actually experience.
Where Sonnet 5.5 looked strongest: visible, bounded product work
The two visual highlights were the contact lens case and the wristwatch, both rated 9.5/10. These outcomes point to a notable strength: Sonnet 5.5 can apparently combine visual detail with functional interaction when the object and the requested behavior are well defined.
The contact lens case was more than a static 3D render
The contact lens case included recognizable blue and red caps, clear left/right labeling, ridged cap details, shadows, and separate clickable interactions. The caps reportedly unscrewed, lifted, turned over, and exposed hollow wells beneath them. Importantly, the left and right sides could open and close independently.
That is a meaningful distinction. Generating a recognizable 3D object is increasingly common. Building a small, understandable interaction model around that object is harder because it requires state management, animation sequencing, object hierarchy, and a design that remains legible as the user rotates the camera.
For marketers and product teams, this category of output has immediate uses: interactive concept demos, landing-page prototypes, product visualizations, educational explainers, and pre-development design experiments. It does not eliminate the need for a professional 3D artist or frontend engineer on a polished commercial experience. But it can compress the distance between an idea and a useful prototype.
The wristwatch showed broader feature integration
The watch was not only a 3D asset. It included a strap, case, crown, buckle, case back, moving hands, primary and secondary time zones, date and day options, night mode, color controls, and a zone-swap behavior. The reviewer specifically checked non-integer time-zone offsets and found the output aligned with current time.
That blends several traditionally separate tasks: Three.js presentation, date/time logic, user controls, responsive state updates, and product-level embellishment. The lesson is not that every agent-generated watch should be trusted unreviewed. The lesson is that a well-scoped prompt can now result in a substantially richer prototype than a wireframe or static mockup.
Anthropic’s current platform documentation lists Sonnet 5.5 with a 1 million-token context window, up to 128,000 output tokens, adaptive thinking, image-input support, and a June 2026 knowledge cutoff. (platform.claude.com) Those specifications help explain why a model can sustain large project context and lengthy implementation traces, although context size alone does not ensure good architecture or good taste.
The most important signal may be the local fine-tuning task
The local fine-tuning task is arguably the most commercially relevant result in the set. Sonnet 5.5 was asked to create a panda-fact dataset, fine-tune Gemma 2B with LoRA, preserve the adapter, and build a local web interface that generates a new fact when refreshed.
According to the review, it created a 119-item dataset, completed the training flow, saved the adapter, built a reproducibility-oriented experiment, and connected inference to the local application. The reviewer also checked that the application was using Gemma with the trained adapter rather than faking the output through static text.
This is an excellent example of why “it wrote the code” is not an adequate evaluation standard. A convincing AI agent workflow has to handle:
- Data creation and formatting.
- Environment and dependency setup.
- Training configuration.
- Artifact saving and loading.
- Inference integration.
- A usable interface.
- Some evidence that the entire chain actually executes.
The project was not flawless. The review found questionable facts and dates in the generated dataset, and the final evaluation reused examples from the training data. Those are important warnings. A model can automate the mechanics of a fine-tuning pipeline while still failing at data quality and evaluation design—the exact areas where careless teams can convince themselves a model works better than it does.
For builders, the practical conclusion is straightforward: let an agent scaffold the pipeline, but maintain human ownership of data provenance, labeling policy, train-validation separation, evaluation metrics, safety constraints, and deployment monitoring. AI can reduce setup time; it cannot make weak source material trustworthy.
The counting problem demonstrates a different kind of agentic value
The 60-grid-point counting task received the only perfect score. It required the model to count valid orders through a constrained grid that wrapped horizontally, visit every point once, begin on the top row, finish on the bottom row, and obey permitted moves. The expected answer was 20,460.
Sonnet 5.5 reportedly derived the correct result, wrote a solver to verify the count, encountered portability trouble on macOS, corrected the issue, and completed the computation. This is a strong example of an agent using code not merely to generate an answer but to test its own reasoning.
That pattern matters far beyond math puzzles. In production engineering, the higher-value behavior is often:
- Interpret the requirements.
- Form a hypothesis or implementation plan.
- Create a reproducible test, script, or check.
- Run it.
- Investigate failures.
- Deliver a result with evidence.
This is also why execution-aware coding evaluations are useful. LiveCodeBench, for example, describes its focus as broader code capabilities that include self-repair, code execution, and test-output prediction—not just static code generation. (livecodebench.github.io) A capable agent is not simply eloquent about a solution; it can seek feedback from the environment.
Where the benchmark exposed limitations
A strong score should not erase the defects. In fact, the defects are more useful than the successes because they show where a human reviewer should concentrate effort.
Visual edge cases still need inspection
The elevator simulation successfully delivered people across multiple floors, including runs where additional passengers were spawned while the app was active. However, at higher volumes, people overlapped and counters temporarily missed passengers walking toward the elevator. The logic worked, but the presentation degraded under load.
The folding table behaved smoothly across slider positions and could reverse mid-motion, but closer examination found small leg intersections and an imperfect hinge area. This is a classic generative-3D problem: broad motion can look right while physical plausibility falls apart in the details.
Controls can fail outside the happy path
The archery game included wind, arrow drop, moving targets, timing, collision handling, replay, and a leaderboard. Automated checks reportedly passed for the core game logic. Yet holding the mouse after an Escape-key cancellation could restart the draw interaction.
This is a small bug, but it illustrates a large principle. AI agents often optimize for the primary interaction path named in the prompt. Edge inputs, cancellations, partial state transitions, accessibility requirements, and stress conditions may receive less deliberate treatment unless they are explicitly specified and tested.
Generated content requires fact checking
The local panda-fact dataset was sufficient to prove the technical workflow, but the reviewer identified factual issues. This is unsurprising: an LLM that generates training data can pass its own uncertain knowledge into a fine-tuned model, then make those statements harder to detect because the output appears tailored and consistent.
The correct response is not to avoid agent-generated data categorically. It is to treat it as a draft corpus requiring validation, especially for customer-facing, educational, medical, financial, legal, or brand-sensitive content.
Sonnet 5.5 versus Opus 5.5: choose by review burden, not prestige
The immediate comparison is with Opus 5.5, which scored higher in the reviewer’s historical KingBench 3 testing. But the more useful decision framework is not “which model wins?” It is “which model gives my team the lowest total cost for this type of work?”
Anthropic lists Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, versus $4 per million input tokens and $20 per million output tokens for Opus 5.5. Cached reads for Sonnet are listed at $0.20 per million tokens, and the Batch API offers a 50% discount on input and output pricing. (platform.claude.com) In simple token-rate terms, Sonnet is half the listed input and output price of Opus.
But token pricing is only one component of cost. A more complete view is:
Total task cost = model spend + human review time + rework + risk of failure.
If Sonnet gets you to a usable prototype in one or two passes, its lower cost and quicker iteration can make it the obvious choice. If a task involves unclear requirements, high-stakes judgment, complicated architecture, legal or security risk, or a painful downstream failure mode, paying more for a stronger model—or simply assigning more human engineering time—may be the better trade.
A practical split might look like this:
| Use Sonnet 5.5 first | Consider Opus 5.5 or intensive review |
|---|---|
| UI prototypes and interactive demos | Ambiguous product requirements |
| Routine bug fixes with good tests | Security-sensitive changes |
| One-off data scripts | Major refactors across large systems |
| Marketing tools and internal utilities | Compliance-heavy workflows |
| Clearly scoped feature implementation | Architecture decisions with long-term consequences |
| Fast variations on existing patterns | Tasks where silent errors are expensive |
The KingBench 3 output supports using Sonnet as a high-throughput builder for bounded tasks. It does not establish that Sonnet should autonomously own your production repository.
What this means for founders, creators, and marketing teams
The arrival of capable coding agents changes the economics of experimentation first. A founder can validate an interaction concept without waiting for a full sprint. A marketer can prototype a calculator, quiz, demo, or interactive content experience. A creator can build a utility around an audience need rather than outsourcing every first draft.
The best use cases share three characteristics:
The desired outcome is visible
An interactive calculator, an internal dashboard, a microsite, a game, or a demo has a concrete result that a human can quickly inspect. That reduces the chance that a plausible-sounding but incorrect implementation survives unnoticed.
Success criteria can be written down
“Build a price calculator” is vague. “Build a mobile-friendly calculator with three inputs, a visible formula, validation for invalid values, downloadable results, and tests for five sample scenarios” gives an agent a much better target.
Someone owns the final review
AI agents reduce the amount of manual implementation required. They do not remove the need for product judgment. Someone still has to decide whether the interaction makes sense, whether the copy is accurate, whether the data is appropriate, whether the UI meets accessibility expectations, and whether the solution is maintainable.
This is also where teams should resist the temptation to measure output only by velocity. A fast prototype that cannot be safely extended may be less valuable than a slower one with readable components, predictable state, documented assumptions, and tests around critical behavior.
A better way to evaluate AI coding agents in your own workflow
The video’s task-by-task inspection provides a useful template. Rather than asking whether a model is “the best coder,” build a short evaluation suite that mirrors the work your team actually does.
Use five to 10 representative tasks, then test each candidate model on the same environment and constraints. Include at least one task from each meaningful category in your workflow: frontend interface work, backend logic, debugging, data handling, documentation, automation, and an intentionally awkward edge case.
Score the artifact, not just the response
A practical rubric can include:
- Functional correctness: Does it meet the acceptance criteria?
- Edge-case behavior: What happens with invalid input, cancellation, retries, or scale?
- Visual and UX quality: Is the interface coherent and usable?
- Code quality: Is it understandable, modular, and reasonably maintainable?
- Test evidence: Did the agent run meaningful checks and interpret them correctly?
- Security and privacy: Did it expose secrets, weaken controls, or make unsafe assumptions?
- Time and cost: How many turns, tokens, and review hours did it require?
A weighted scoring system is usually better than a single subjective impression. For example, an internal campaign landing page may weight speed and visual polish heavily. A billing integration should heavily weight tests, security, error handling, and maintainability.
Add adversarial checks deliberately
Models often perform well on the obvious path. Ask what happens when a user enters an empty field, clicks rapidly, changes state midway through an animation, refreshes during an operation, uses a keyboard only, switches time zones, uploads malformed data, or performs an action twice.
The elevator crowding and archery cancellation bug in this evaluation are reminders that these details are not optional polish. They are where a prototype becomes either a dependable product or a fragile demo.
Separate model quality from harness quality
An agent’s result depends on more than the underlying model. It depends on the coding harness, tool permissions, prompts, environment setup, package availability, test suite, context provided, reasoning settings, and how much iteration is allowed.
That does not invalidate the Sonnet 5.5 result. It simply means that your own results may differ. Documenting the harness and evaluation settings is essential before using a benchmark score to make a vendor or workflow decision.
Why this result is more useful than a leaderboard placement
Leaderboards remain helpful for narrowing choices, but they rarely answer the most practical questions: Can the model build a credible interactive demo? Can it repair itself when the first command fails? Does it know when to test? Does it create something that a designer, customer, or stakeholder can understand without reading source code?
The strongest outputs in this Sonnet 5.5 evaluation were compelling because they connected code with a visible result: a lens case that opened, a table that folded, a game that could be played, a watch that tracked time, and a local ML app that actually generated outputs. Those artifacts are easier to audit and more aligned with the way many small teams work.
At the same time, the test reinforces a less glamorous truth: agentic coding is already good enough to move the bottleneck from implementation toward specification and review. When a model can build the first 80% of an interactive project quickly, the competitive advantage goes to teams that define the remaining 20% clearly—quality standards, brand taste, accurate data, edge-case handling, security boundaries, and shipping discipline.
The bottom line on Claude Sonnet 5.5
This Claude Sonnet 5.5 coding benchmark is a strong endorsement for using the model on scoped, iterative, tool-enabled development work. An 89.38% score across eight visually and technically varied tasks suggests the model can do far more than autocomplete functions: it can assemble complete prototypes, debug execution problems, and connect multiple technical layers into working artifacts.
But “working artifact” is the key phrase, not “finished product.” The score came from an editorial, single-run evaluation, and the individual tasks still exposed familiar AI-agent weaknesses: degraded presentation under scale, minor interaction bugs, geometry imperfections, and unverified facts in generated training data.
For most teams, Sonnet 5.5 looks best suited to fast prototyping, implementation of clear tickets, internal tools, visual experiments, and repeatable coding workflows with human oversight. Use it to accelerate the build-test loop. Do not use any benchmark result as permission to skip validation.
FAQ
What score did Claude Sonnet 5.5 receive in the KingBench 3 test?
Claude Sonnet 5.5 scored 71.5 out of 80, or 89.38%, in the reviewed eight-task KingBench 3 run. The score was based on the evaluator’s functional and presentation assessment, not an official Anthropic benchmark.
Did Claude Sonnet 5.5 beat Opus 5.5?
No. In the reviewer’s comparison, Opus 5.5 scored 75 out of 80, or 93.75%, compared with Sonnet 5.5’s 89.38%. Sonnet was close enough, however, that its lower token pricing may make it a better fit for many well-scoped tasks.
What kinds of coding tasks did Sonnet 5.5 perform best on?
The strongest results were interactive 3D projects—the contact lens case and wristwatch—plus a constrained counting problem and a local LoRA fine-tuning workflow. These tasks combined code generation with execution, UI behavior, and visible outputs.
Is Claude Sonnet 5.5 good enough to build production software alone?
It can accelerate production work, but this test does not justify unsupervised deployment. The evaluation found minor bugs, visual issues, and factual problems in generated data. Production software still needs code review, testing, security checks, accessibility review, and ownership by experienced people.
How much does Claude Sonnet 5.5 cost through the API?
Anthropic currently lists Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, with lower pricing for cache reads and discounted Batch API usage. Pricing can change, so verify current rates before budgeting. (platform.claude.com)