Mistral Large 4 is arriving with enormous headline numbers, but this Mistral Large 4 coding benchmark points to a more useful question for builders: can Le Chonk finish a real task in a fresh folder without someone rescuing it? In an eight-task KingBench evaluation, the answer was mixed—excellent on several visual projects and one end-to-end fine-tuning workflow, but inconsistent when execution, file paths, or multi-step delivery became unforgiving.

The original hands-on test, published on YouTube and run through OpenCode with OpenRouter, awarded Mistral Large 4 a first-pass total of 42 out of 80, or 5.25 out of 10. That score is not a universal ranking of the model. It is a practical reliability signal from one toolchain: eight fixed tasks, fresh project folders, actual browser checks, and grading based on working output rather than a persuasive agent summary.

That distinction matters. A frontier model can write elegant code, reason about a solution at length, and still leave its user with a blank browser tab, a missing file, or an app that fails at the final integration point. For founders, creators, and developers using AI coding agents, the gap between “the model knows how” and “the model delivered” is where real project cost lives.

What the Mistral Large 4 coding benchmark actually tested

The benchmark used eight KingBench-style tasks designed to test different parts of an AI coding workflow. Rather than relying on a single repository-repair benchmark or a text-only coding question, the prompts covered browser simulations, 3D interfaces, SVG illustration, game mechanics, a calculation task, local model fine-tuning, and a full watch interface.

The evaluator connected the preview model to OpenCode through OpenRouter. Each prompt began in a fresh folder, which is important because it limits the chance that a previous project, cached artifact, or manual setup quietly helps the result. For browser projects, the output was opened in Chromium and evaluated as a user would experience it: did the page render, did the requested interaction work, and did the console show errors?

The eight tasks were:

  1. A three-elevator passenger simulation with queues, destinations, and animations.
  2. A one-file interactive 3D contact lens case with independent lids.
  3. A continuously folding 3D table.
  4. An SVG illustration of a panda eating a burger.
  5. A bow-and-arrow target game with timing and a leaderboard.
  6. A counting problem with a specific correct answer.
  7. A local Gemma fine-tuning workflow plus a working panda-fact web app.
  8. A 3D wristwatch with live time, moving hands, and selectable time zones.

This is a useful mix because it reveals that “coding ability” is not one skill. A model may be excellent at generating a self-contained HTML canvas experiment yet less dependable when it must manage dependencies, choose correct local paths, produce a final answer after reasoning, or validate its work against a browser error.

Why first-pass scoring is the right lens for agents

The benchmark’s strictest rule was also its most valuable: a plan does not count as a delivered project. If the agent spent a session reasoning but never created the file, it received no credit. If it wrote a watch interface that looked ambitious in the source but rendered as a blank page, it received no credit on the first pass.

That may sound harsh, but it matches how AI agents are increasingly used. The appeal of an agent is not merely that it can suggest code in a chat window. The appeal is delegation: assign a scoped task, return later, and receive something testable. A system that regularly requires a user to identify the final import error or tell it to write the file is still useful, but it is not truly hands-off.

Mistral Large 4: big specs, ambitious positioning

Mistral launched Mistral Large 4 in public preview on October 6, 2026. The company calls the model “le Chonk,” a playful reference to its scale. Mistral describes it as a natively multimodal granular mixture-of-experts model with 1.05 trillion total parameters, a 1.6 billion-parameter vision encoder, and approximately 52 billion active parameters in its product documentation. Its launch announcement describes 49 billion active parameters, so builders should treat the active-parameter figure as a release-detail discrepancy worth checking against the final model card and weights release. (mistral.ai)

Mistral says the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its European data centers. It positions Large 4 for coding, agents, multimodal work, cybersecurity, finance, law, and visual grounding, while promising open weights later in October 2026. (mistral.ai)

For API users, the current public preview is not a trivial experiment. Mistral’s documentation lists a one-million-token context window, structured outputs, function calling, document Q&A, agent features, and built-in tools. OpenRouter lists a 524,288-token context window for the routed offering, along with a maximum 262,144-token output, which shows why developers should verify the limits of the specific provider endpoint they are using rather than assuming a headline specification applies everywhere. (docs.mistral.ai)

Pricing is also aggressive at the preview level. Mistral’s model page shows $0.68 per million input tokens and $2.09 per million output tokens, while OpenRouter identifies those as a temporary 50% discount from higher posted pricing. That can make Large 4 attractive for lengthy implementation passes, but token price is only one part of the economics. A cheap run that fails near the end and needs two more runs plus developer debugging is not necessarily cheaper than a more expensive model that completes the task cleanly. (docs.mistral.ai)

Where Le Chonk performed best: visual prototypes with clear boundaries

The clearest pattern in the test is that Mistral Large 4 looked strongest when the request was visually concrete, reasonably self-contained, and easy to validate in one browser page. Those are the kinds of tasks where an AI agent can create visible value quickly: interactive mockups, demo concepts, small product experiments, data visualizations, and marketing microsites.

Elevator simulation: a nearly complete success

The first task asked for a three-elevator simulation, with a single passenger allowed per car, passengers spawning on floors, destination information on hover, animations, queues, and a mechanism for people left behind to catch the next available elevator.

Le Chonk produced a clean building view with colored elevators, controls for passenger count, passenger destinations, door animations, and a functioning queue. When the test spawned five people and let the simulation run, the waiting and riding counts returned to zero without browser errors—a meaningful functional validation, not just a visual inspection.

The evaluator noted two minor issues. Resetting during an active trip could briefly carry an old exit animation into the reset scene, and the waiting count could decrement when a passenger was assigned to an elevator rather than when they physically entered it. Those are exactly the kind of edge cases that distinguish a convincing demo from production interaction design. Still, the project scored 9.5 out of 10.

For product teams, the lesson is encouraging. If you need a stakeholder demo for a scheduling concept, logistics simulation, onboarding interaction, or animated product explainer, this kind of model output may be highly usable as a first artifact. The work is contained enough to review in a browser, and visual defects are easy to spot before they become customer-facing defects.

Folding table: strong 3D implementation, modest polish gap

The 3D folding-table task also performed well, earning 8.5 out of 10. The model generated a Three.js scene with lighting, materials, shadows, orbit controls, and a slider that moved the object from a compact folded state to an unfolded tabletop-and-legs configuration.

A continuously controlled 3D movement is more meaningful than a pair of static renders. It requires the agent to coordinate geometry, transforms, UI state, and rendering. The project rendered without page errors in both folded and unfolded states, which validates the central request.

The remaining flaws were visual rather than foundational: the feet could lift from the ground during the animation, and the model was somewhat small in the viewport. Those defects matter if the output is a final product configurator, but they are entirely normal in a prototype. A designer or front-end developer could use the generated scene as a starting point rather than recreating the structure from scratch.

Panda SVG: proof that small creative assets can be finished

The panda-eating-a-burger SVG earned 8 out of 10. It rendered as a valid, recognizable visual asset: panda, paws, burger, and open mouth were present. The weaker point was semantic specificity—the burger was held at the mouth, but the “eating” action was not as clear as the prompt requested.

That result may sound less important than the 3D builds, but it shows an often-overlooked AI workflow. Marketing and creator teams frequently need small, original illustrative assets for landing pages, slides, tutorials, social posts, or prototype UIs. A model that can reliably create a usable SVG offers a faster starting point than searching stock libraries for a close-but-not-quite fit.

The caveat is that “usable” does not equal “brand ready.” Teams should still review paths, colors, accessibility labels, file weight, consistency with a design system, and whether the image communicates the exact action or emotion required.

The failures were not just coding failures—they were workflow failures

The most important warning from the benchmark is not that Mistral Large 4 made bugs. Every coding model makes bugs. The larger issue is that several failures occurred because the system did not reliably finish the basic workflow.

Contact lens case: capability appeared only after changing the run mode

On the 3D contact-lens case prompt, the default OpenCode session spent its turn reasoning and then ended without creating a file. The first-pass result was correctly scored 0 out of 10, because there was nothing to run.

In a separate retry using a supported no-reasoning configuration, the model did create the requested case. The independent left and right caps opened on click, although they rose too far and partly left the frame. That retry shows the model may have had the capability to solve the task, but it does not erase the first-pass reliability problem.

This distinction should influence how teams configure coding agents. Reasoning modes are not simply “more intelligence” settings. They affect token budgets, tool behavior, completion behavior, latency, and the likelihood that an agent transitions from planning into action. If a model frequently consumes its turn in internal deliberation, a team needs guardrails that require artifact creation before completion.

Useful implementation guardrails include:

  • Require an agent to list the files it created or modified in its final response.
  • Add a completion gate that checks whether required files exist.
  • Tell the agent to run the project or a smoke test before declaring success.
  • Set explicit task milestones for longer jobs: scaffold, implement, test, repair, summarize.
  • Capture browser console errors and return them automatically to the agent for one bounded repair pass.

The counting task: a simple test of whether the agent closes the loop

The counting problem had a known correct answer of 20,460. Yet the default session again ended without a usable final answer, receiving 0 out of 10.

This is revealing precisely because the underlying task was not an elaborate dependency-management challenge. It demonstrates that an agent can fail at the last mile even when the job is primarily reasoning. For a human collaborator, an incomplete line of thought may be recoverable. For an autonomous workflow, it is a failed task.

Builders should therefore track more than accuracy. A useful internal evaluation scorecard has at least four dimensions:

  1. Task correctness: Is the answer, implementation, or decision right?
  2. Artifact completion: Did the agent create the requested deliverable?
  3. Execution validity: Does the deliverable run, compile, render, or pass tests?
  4. Recovery behavior: Can the system identify and repair its own failure with bounded feedback?

A model can rank well on the first dimension while still creating operational friction on the other three.

The standout result: local Gemma fine-tuning end to end

The benchmark’s strongest outcome was the seventh task, which asked the agent to generate a panda-fact dataset, fine-tune a local Gemma 2B model, and build a web page that displayed a new fact on each refresh.

There was an initial environment blocker: the OpenCode setup did not have permission to access the cached model. The evaluator changed that permission and reran the identical prompt in a new folder, treating the access issue as a setup constraint rather than a model error. With the environment available, Large 4 completed the workflow.

According to the test, it generated 510 training examples, trained a LoRA adapter for 400 steps, reduced validation loss from roughly 7.97 to 0.31, identified and repaired a server-threading problem during its own end-to-end test, and produced a local browser app. Multiple refreshes returned different panda facts without browser errors. The task received a full 10 out of 10 in the benchmark’s scoring.

Why this result matters more than a polished one-file demo

This project involved many failure points:

  • creating suitable training data;
  • selecting and running local fine-tuning code;
  • accessing model files;
  • managing a training process;
  • serving a resulting model or adapter;
  • building a web interface;
  • fixing a runtime issue; and
  • checking that refreshed output changed as requested.

That is much closer to a small AI product workflow than an isolated HTML exercise. It suggests that Mistral Large 4 can be valuable for assisted experimentation with fine-tuning, internal knowledge demos, local proof-of-concepts, and model-backed applications—provided the environment is correctly prepared and the developer keeps verification in the loop.

It also highlights a second-order lesson: environment quality can overwhelm model quality. File permissions, cached checkpoints, package versions, GPU availability, server ports, and local paths are all part of the effective system. When a team says a model “failed,” it should separate failures caused by the model’s plan from failures caused by the harness it was allowed to operate in.

The 3D watch explains why browser testing must be mandatory

The final task asked for a fully featured 3D wristwatch with current time, smoothly moving hands, and two time zones. The generated code was ambitious. It included a local approach for Three.js modules and a fairly detailed watch page.

But the first browser check revealed a blank render area and placeholder dashes for both time displays. The cause was an import map that referenced orbit and environment modules at paths that did not exist. The benchmark scored the first attempt 0 out of 10, which is appropriate: an application that never loads has not completed its job.

When the exact browser error was provided to the model, it repaired the path problem. The revised result rendered the watch, advanced the clock, showed a sweeping second hand, and allowed a second time zone to change. However, it still had visual and state-quality issues: the bracelet dominated the frame, parts of the case crossed the dial, and changing the second zone did not immediately update the printed label on the watch face. The repaired output was judged 6 out of 10, but kept outside the first-pass total.

This is one of the benchmark’s most actionable findings. AI coding systems should not be allowed to self-certify from source code alone. A minimal web-agent test loop should include:

  1. Start the local server.
  2. Open the target URL in a real browser or browser automation environment.
  3. Capture console errors, network failures, and a screenshot.
  4. Run one or more core interactions.
  5. Return concise failure evidence to the agent.
  6. Limit the repair loop to a defined number of attempts.

That workflow turns vague “it didn’t work” debugging into machine-readable feedback. It also helps teams measure a model’s true recovery capability rather than its ability to produce plausible code on the first draft.

What the 5.25/10 score does—and does not—mean

A score of 5.25 out of 10 sounds mediocre beside the model’s trillion-parameter scale and Mistral’s frontier claims. But it would be a mistake to interpret it as a simple statement that the model is weak at coding.

The individual results show a high variance profile:

TaskFirst-pass scoreMain takeaway
Elevator simulation9.5/10Strong interactive visual prototype with minor edge cases
Contact lens case0/10No artifact in the default run; retry proved underlying capability
Folding table8.5/10Working 3D interaction with modest animation/polish flaws
Panda SVG8/10Finished visual asset, slightly weak action semantics
Bow-and-arrow game6/10Systems existed, but playability and responsive layout needed work
Counting problem0/10Failed to produce a final usable answer
Gemma fine-tuning app10/10Excellent end-to-end completion after environment access was fixed
3D wristwatch0/10Initial project failed to render because of missing module paths

The bow-and-arrow game, which scored 6 out of 10, belongs in the middle of this picture. It included targets, wind, projectile physics, hit tracking, and a browser-stored leaderboard, so the agent did build meaningful systems. Yet targets were too small at a typical desktop size, positions were fixed in a way that could push a distant target off-screen on mobile, and the aiming guide did not precisely match projectile drag. It was technically substantial but not ready to be called fun or polished.

That is the larger pattern: Le Chonk often generated enough structure to prove competence, while reliability and user-centered refinement remained uneven. For a developer supervising the work, that can still be very productive. For someone hoping to delegate a task and return to production-ready output, it is not enough yet.

How this compares with Mistral’s official claims

Mistral’s launch materials position Large 4 as its largest and most capable model to date, with strong results across coding, agentic behavior, multimodal understanding, cybersecurity, finance, and law. The company says the model is competitive with leading open-source models globally and stronger than other US- or Europe-developed open-weight models in its comparisons. (mistral.ai)

The benchmark does not necessarily contradict those claims. It tests a narrower but highly practical issue: how a preview model behaves inside a particular agent framework, through a third-party routing layer, on specific prompts, with real browser checks. Official benchmark results typically measure standardized tasks under carefully selected configurations. A hands-on agent test measures the compound system of model, reasoning mode, tools, provider behavior, permissions, shell environment, and validation process.

That difference is why both types of evidence matter. Vendor benchmarks can indicate capability ceilings. Independent hands-on work reveals operational floors. A buyer deciding whether to use Large 4 for code review, prototype generation, local experimentation, security research, or autonomous implementation needs both.

Mistral has also said the current release is a public preview and that weights are planned by the end of October 2026; the associated Hugging Face page lists an expected October 31 release date. That makes today’s agent behavior a moving target, not a final verdict on the eventual open-weights release. (mistral.ai)

Practical guidance for founders, marketers, and developers

The most productive way to use Mistral Large 4 today is not to ask whether it is “the best coding model.” Instead, identify whether your task resembles its demonstrated strengths and build safeguards around its demonstrated weaknesses.

Use it for contained prototype work

Le Chonk appears promising for tasks with visible acceptance criteria and constrained file scope. Examples include interactive landing-page sections, data-storytelling demos, web-based calculators, simple internal dashboards, animated product concepts, SVG assets, and 3D proof-of-concepts.

These tasks benefit from fast generation, and a human reviewer can inspect the output in minutes. A flawed animation, an off-center object, or a mobile layout issue is usually cheaper to fix than an unnoticed backend logic error.

Do not make it your unsupervised default for long workflows

For multi-step builds involving dependencies, local environments, file-system assumptions, model training, or a deployment pipeline, use the model as a capable operator with checks—not as an unattended employee.

A practical operating model is:

  • Give the agent a clear definition of done.
  • Require it to create and list artifacts.
  • Run automated lint, test, build, and browser checks.
  • Feed errors back once or twice with a fixed retry budget.
  • Review diffs before merging.
  • Preserve logs so failures become evaluation data, not anecdotes.

For teams building email-triggered product flows, the same principle applies to outbound systems: validate the critical end-to-end behavior, not just the generated implementation. That includes API response handling, retries, logging, and recipient validation before a workflow reaches users.

Treat reasoning settings as a product decision

The contact-lens task suggests that reasoning configuration may materially affect whether the agent produces an artifact. Mistral’s API and third-party gateways can expose different controls, context limits, pricing, and routing behavior, so teams should test the exact configuration they plan to deploy.

Make a small internal evaluation set from your real work. Include one short front-end request, one bug fix, one multi-file feature, one data task, one tool-use task, and one workflow requiring a browser or API verification. Run each prompt several times. Measure success rate, not just the best output.

The winning model for your business is often not the one with the most impressive launch benchmark. It is the one that produces acceptable artifacts repeatedly in your stack, at a total cost your team can justify.

Community reaction is still limited, so implementation evidence matters more

No substantive top-comment reaction was supplied with the original benchmark material, which is unsurprising for a newly released preview model. Early discussion around a launch often focuses on parameter counts, pricing, open-weight promises, and vendor benchmark charts before enough developers have run the model through their own toolchains.

That makes the original video especially useful as early counterweight. It does not claim scientific finality. Instead, it documents the outputs on screen, separates first attempts from repaired retries, and explains where environmental configuration affected the result. That methodology gives builders something more concrete than launch enthusiasm: examples of where to expect leverage and where to expect supervision.

As more independent evaluations appear, the most valuable comparisons will not be generic leaderboard snapshots. Watch for tests that hold prompts, agent harnesses, tool permissions, retry rules, and validation steps constant across models. Those comparisons will better reveal whether Large 4’s inconsistency is specific to preview behavior, the OpenCode setup, its reasoning mode, or a broader pattern in long-horizon agent work.

Final verdict: capable enough to test, not reliable enough to trust blindly

The Mistral Large 4 coding benchmark delivers a balanced early verdict. Le Chonk can make genuinely impressive things: a working elevator simulation, a controlled folding 3D table, a polished SVG asset, and—most notably—a complete local Gemma fine-tuning workflow connected to a functioning web app.

At the same time, it missed the most basic definition of completion on multiple first attempts. A contact-lens task produced no file. A counting task produced no final answer. A 3D watch generated a blank page because dependencies were referenced incorrectly. These are not exotic edge cases for coding agents; they are everyday reasons why autonomous workflows disappoint users.

For builders, the right conclusion is neither hype nor dismissal. Use Mistral Large 4 for bounded visual prototypes and supervised technical experiments. Put browser checks, file checks, tests, and repair loops around it for larger work. And judge it by the metric that ultimately matters: how often you can assign a task and come back to something that actually works.

FAQ

What score did Mistral Large 4 receive in the benchmark?

The first-pass KingBench run scored Mistral Large 4 at 42 out of 80, equivalent to 5.25 out of 10. The total excluded separate retries for the contact-lens project and 3D watch, while including the fine-tuning rerun because its initial blocker was a permissions setting in the local environment.

Is Mistral Large 4 good for coding?

It appears strong for constrained visual coding tasks and can complete sophisticated workflows when the environment is correctly configured. However, the benchmark found inconsistent first-pass delivery on some tasks, so it is better treated as a supervised coding partner than a fully autonomous implementation agent.

Why did the 3D watch score zero at first?

The initial watch page failed to render because its import map referenced Three.js-related module paths that were not present. After the browser error was supplied, the model repaired the project, but the repaired version was scored separately from the first attempt.

Can Mistral Large 4 fine-tune a local model?

In the tested workflow, it successfully generated a dataset, fine-tuned a local Gemma 2B model using a LoRA adapter, fixed a server issue, and produced a web app that returned different panda facts on refresh. Your results will depend heavily on local permissions, hardware, dependencies, and the agent environment.

Are Mistral Large 4’s open weights available now?

Mistral released the model as a public preview on October 6, 2026 and says weights are expected by the end of October. The model’s Hugging Face page lists an expected release date of October 31, 2026, so developers should confirm availability and license terms before planning a self-hosted deployment. (mistral.ai)