> ## Content Index
> Fetch the complete content index at: https://corti.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# When Using Open Models Make Sense
- URL: https://corti.com/when-using-open-models-make-sense/
- Published: 2026-10-02T07:19:13.000Z
- Updated: 2026-10-02T07:19:13.000Z
- Author: Sascha Corti

Open-weight models make sense when the task is bounded and checkable, when the data should not leave your network, or when you need an endpoint nobody else can reprice, rate-limit or retire. They do not make sense as a blanket replacement for the frontier.

This post sets out where I draw the line, using my own setup as the worked example: GLM-5.3-Flash served from a two-node NVIDIA DGX Spark cluster, driving my Hermes and OpenClaw agents and the chat I use for research and everyday questions.

## The gap is about four months, and it is not the point

The best open-weight models trail the closed frontier by four months on average. [Epoch AI](https://epoch.ai/data-insights/open-closed-eci-gap?ref=corti.com) measured that gap on its Capabilities Index for the period since January 2026\. It equals 8 index points, about the distance between GPT-5 and GPT-5.5.

Four months is short. The models I can download today match what I was paying frontier prices for earlier this year, and I did real work with those. So the useful question is not "which model is best?" but "which model is good enough for this task?". That framing comes from Every's guide [Getting Started With Open Models](https://every.to/guides/getting-started-with-open-models?ref=corti.com), which prompted this post.

[GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash?ref=corti.com) is a good example of what "good enough" now looks like. Z.ai released it on 26 August 2026 under the MIT licence. It has 320B total parameters with 18B active per token, a 1M-token context window, and accepts text, image and video input.

| Benchmark           | GLM-5.3-Flash | Claude Opus 4.8 |
| ------------------- | ------------- | --------------- |
| Terminal Bench 2.1  | 84.3          | 85.0            |
| DeepSWE v1.1        | 63.4          | 58.0            |
| NL2Repo             | 56.3          | 69.7            |
| Toolathlon Verified | 78.4          | 76.2            |
| HLE with tools      | 55.3          | 57.9            |

Scores are Z.ai's own, as published on the [Ollama model page](https://ollama.com/library/glm-5.3-flash?ref=corti.com). Read them with two caveats. They are vendor-reported, and Opus 4.8 is no longer Anthropic's newest Opus: [Opus 5 and Opus 5.5](https://platform.claude.com/docs/en/about-claude/pricing?ref=corti.com) have shipped since. That is the four-month lag in miniature.

## Four reasons to run open models

Open weights buy four things: lower cost per task, control over where data goes, independence from a vendor's decisions, and control over the model itself. Only the first is about money, and it is the weakest argument for running your own hardware.

### 1\. Cost per task

Open models are cheap because anyone can host them, so hosts compete on price. GLM-5.3-Flash is offered at $0.15 per million input tokens and $0.50 per million output tokens on [Ollama's cloud](https://ollama.com/library/glm-5.3-flash?ref=corti.com). Claude Opus 4.8, the model it is benchmarked against, [lists at $5 and $25](https://platform.claude.com/docs/en/about-claude/pricing?ref=corti.com). That is roughly 33 times cheaper on input and 50 times on output.

Per-token price is not cost per task. A weaker model that needs three attempts, or a human to repair its output, can cost more than a strong model that gets it right once. Measure the full loop: tokens, retries and your own review time.

Pricing models are also shifting under heavy users. [GitHub Copilot moved every plan to usage-based billing](https://stackfutures.com/blog/github-copilot-usage-based-billing-june-2026/?ref=corti.com) on 1 June 2026, because agent sessions consume far more tokens than autocomplete did. Agents are exactly the workload where a cheap or fixed-cost model matters.

### 2\. Data control

A model is a file of numbers. Running it on hardware I own means prompts, documents and agent memory never leave my network. For research notes, private repositories and anything an agent reads from my inbox or file system, that is a property no hosted service can match.

The model's origin does not change this. GLM comes from a Chinese lab, but the weights run offline on my machines and send nothing anywhere. If you do not want to run hardware, a hosted open model on a provider and region you trust is the middle ground.

### 3\. Independence

A closed model is a service someone else operates. They can raise the price, change rate limits, alter behaviour between versions, restrict a capability or retire the model. Weights on my disk do none of that. The checkpoint I validated last month is byte-for-byte the checkpoint I run today.

This matters most for agents. An agent's prompts, tool descriptions and skills are tuned against one model's behaviour. A silent upstream change breaks that tuning without any error message.

### 4\. Control over the model

With the weights I choose the quantisation, the context length, the sampling parameters, the reasoning effort and the serving engine. I can fine-tune, and I can inspect every request and response. GLM-5.3-Flash's MIT licence allows all of it, commercially, without asking anyone.

## When to stay on a frontier model

Use the strongest model you can get when the task is vague, long, hard to verify or costly to get wrong. Those are the conditions where the four-month gap shows.

- **Ambiguous requests.** Frontier models are better at working out what you meant from an underspecified prompt.
- **Long unsupervised runs.** Errors compound over many steps. A small per-step reliability gap becomes a large end-to-end gap.
- **Work you cannot check cheaply.** If verifying the output takes as long as doing the task, a wrong answer from a cheaper model saves nothing.
- **High stakes.** A one-off production refactor, a legal or financial judgement, anything that sends, spends, publishes or deletes before a human reviews it.

The measured gap probably understates the real one. Epoch AI raises two caveats, as [summarised by AlphaSignal](https://alphasignal.ai/news/epoch-finds-open-weight-ai-models-trail-closed-frontier-by-four-months?ref=corti.com): open models tend to do worse on private benchmarks than on public ones, and closed labs do not always release their most capable systems.

There is also a local-specific cost: quantisation. The model I can fit on two desktops is not the model that produced the published scores. More on that below.

## My setup: GLM-5.3-Flash on two DGX Sparks

I serve GLM-5.3-Flash from two NVIDIA DGX Spark units clustered over their ConnectX-7 ports. Each Spark has 128 GB of unified memory, so the pair gives the model 256 GB to live in.

| Component       | Specification                                                 |
| --------------- | ------------------------------------------------------------- |
| Chip            | NVIDIA GB10 Grace Blackwell, 20-core Arm CPU                  |
| Memory per node | 128 GB LPDDR5x unified, 273 GB/s bandwidth                    |
| Interconnect    | ConnectX-7 NIC at 200 Gbps                                    |
| Power supply    | 240 W per node                                                |
| Size and weight | 150 × 150 × 50.5 mm, 1.2 kg                                   |
| List price      | $4,699 per node (launched at $3,999, raised 23 February 2026) |

Hardware figures are from [NVIDIA's specification page](https://www.nvidia.com/en-us/products/workstations/dgx-spark/?ref=corti.com); the price history is from [GPUSmith's review](https://gpusmith.com/articles/en/nvidia-dgx-spark-review-specs-performance?ref=corti.com).

### Why this model fits

GLM-5.3-Flash is a mixture-of-experts model: 320B parameters on disk, but only 18B active for any one token. Memory capacity limits which model loads. Memory bandwidth and active parameters limit how fast it generates. A large sparse model suits a machine with a lot of slow-ish memory.

The full-precision weights do not fit. The BF16 source is 598.5 GiB. [LibertAI's NVFP4 build](https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4?ref=corti.com) quantises only the routed-expert weights, which are 97% of the parameters, to 4 bits and leaves attention, the vision tower and embeddings in BF16\. That brings the checkpoint to about 181 GiB, or roughly 194 GB of the 256 GB available. The rest goes to the KV cache, per-request state and the operating system.

### What throughput to expect

Published two-Spark recipes land around 30 tokens per second for a single stream. [LimeChain's vLLM recipe](https://github.com/LimeChain/glm-5.3-flash-2xdgx-spark-vllm?ref=corti.com)measured 29.74 tok/s with one request and 101.74 tok/s aggregate across eight concurrent requests, with a 262,144-token context configured. Those are their numbers on their configuration, not a benchmark of mine.

That throughput settles the cost question. One stream running flat out for 30 days produces about 77 million output tokens. At the hosted price of $0.50 per million, that is under $40 of inference from $9,398 of hardware at list price. Self-hosting GLM-5.3-Flash does not pay for itself against hosted GLM-5.3-Flash. I run it locally for reasons two to four above, not reason one.

## What I run on it

Three clients share the one endpoint: two agent harnesses and a chat app. All three speak the OpenAI-compatible API, so swapping the model behind them is a configuration change.

| Client                                                                                              | What it is                                                                                        | Why a local model suits it                                                     |
| --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| [Hermes Agent](https://hermes-agent.nousresearch.com/docs/getting-started/quickstart?ref=corti.com) | Nous Research's MIT-licensed agent with persistent memory and self-created skills                 | Always on, token-hungry, holds long-lived personal context                     |
| [OpenClaw](https://docs.openclaw.ai/gateway/local-models?ref=corti.com)                             | Open-source, self-hosted agent gateway that connects a model to tools and messaging channels      | Reads files, mail and web pages that I would rather not send to a third party  |
| [Cumbersome](https://folding-sky.com/cumbersome?ref=corti.com)                                      | Native API client for iPhone, iPad and Mac that connects directly to a provider or a local server | Research and everyday questions, with the exact model and token counts visible |

### Agents

Agents are the best fit for a local model. They run in the background, so 30 tokens per second is acceptable. They burn tokens on every tool call, so metered pricing hurts most here. And they carry the most private context of anything I run.

Two configuration details matter more than the model choice:

- **Context length.** Hermes Agent requires at least 64,000 tokens of context and rejects smaller windows at startup. On memory-tight hardware, context length and concurrency compete for the same memory, so decide how many agents run in parallel before you set it.
- **Tool calling.** A model that answers a plain prompt can still fail an agent turn. OpenClaw's documentation says so directly: a successful text-only probe does not prove the model can complete a multi-step agent task. Test real tasks before making a local model the default.

### Chat

For research and question-answering I talk to the same endpoint through Cumbersome. It sends requests straight from the device to my server with no vendor in the middle, keeps API keys in the Apple Keychain, and shows which model answered and how many tokens it used.

GLM-5.3-Flash always reasons, and its `reasoning_effort` defaults to `max`. The [model card](https://huggingface.co/zai-org/GLM-5.3-Flash?ref=corti.com) accepts `low`, `high` and `max`, and recommends passing `clear_thinking=true` for chat. On local hardware, effort level is the main lever on how long an answer takes.

## What self-hosting costs you

Self-hosting trades a token bill for an operations burden. With a five-week-old architecture on desktop hardware, that burden is real. Four costs stand out.

### New models break serving stacks

GLM-5.3-Flash uses a new architecture, `glm5_next`, that combines sparse and linear attention. According to [LibertAI's model card](https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4?ref=corti.com), support was not in vLLM's main branch at release and shipped in per-model container images instead. Getting it running on the Spark's GB10 chip needed two independent fixes, and either fault alone left the model producing garbage.

The failure modes are silent, which is the dangerous part:

- **Degenerate output without an error.** A checkpoint missing an activation scale made vLLM's NVFP4 mixture-of-experts path multiply every expert by zero. The model emitted one token repeatedly, with no error and no warning.
- **Wrong tool-call parser.** With the `glm` parser instead of `glm47`, requests succeed but return empty content and no tool calls. An agent sees a model that does nothing.
- **Wrong reasoning parser.** The correct parser name differs between SGLang and vLLM. The wrong one can discard the whole reply, which looks like a hung model.
- **Context versus concurrency.** The linear-attention layers keep per-request state. On two Sparks under SGLang, LibertAI found that eight concurrent requests at 131,072 tokens failed to start, while two at 65,536 worked.

None of these exist on a hosted API. Someone else already debugged them.

### Quantisation is an unmeasured discount

The benchmark scores above belong to the full-precision model. The 4-bit build that fits two Sparks is a different artefact. LibertAI reports a per-expert round-trip cosine of about 0.99665 against the source weights, which measures weight fidelity, not task performance. LimeChain states that its benchmark does not establish model-quality equivalence. Nobody has published the agentic scores of the quantised model, so run your own evaluation on your own tasks.

### Agents on local models need more guarding, not less

OpenClaw's documentation is blunt: local models do not provide hosted providers' safety filters, so tool permissions and prompt-injection defences must match the model and the task. The harness itself is also attack surface. [Adversa AI's review](https://adversa.ai/blog/openclaw-security-101-vulnerabilities-hardening-2026/?ref=corti.com) documents CVE-2026-25253, a CVSS 8.8 flaw patched in version 2026.1.29, and the ClawHavoc campaign, in which 341 of 2,857 audited ClawHub skills were malicious.

Running the model at home removes one data flow. It does not remove indirect prompt injection from a web page or an email the agent reads. Sandbox the agent, keep tool permissions narrow, pin and review skills, and keep the gateway off the public internet.

### Open licences can change

Open is a property of one release, not of a vendor. GLM-5.3-Flash is MIT. Its larger sibling, [GLM-5.3](https://huggingface.co/zai-org/GLM-5.3?ref=corti.com), ships under a custom `glm-5.3` licence. Check the licence on every new checkpoint. The weights already on your disk keep the terms they shipped with.

## A decision table

Start from the task, not the model. Each row is a property of the work; the more rows that land in the middle column, the stronger the case for an open model.

| Task property | Points to an open model                     | Points to a frontier model                     |
| ------------- | ------------------------------------------- | ---------------------------------------------- |
| Definition    | Clear inputs and a describable output       | Vague or shifting requirements                 |
| Verification  | Cheap to check, cheap to rerun              | Checking costs as much as doing                |
| Stakes        | A wrong result is caught before it matters  | Sends, spends, publishes or deletes unreviewed |
| Frequency     | Recurring, high volume, background          | One-off                                        |
| Horizon       | One bounded pass, or short tool loops       | Long unsupervised multi-step runs              |
| Data          | Private, regulated or simply yours          | Already public or low sensitivity              |
| Latency       | Background work tolerates 30 tok/s          | Interactive work at scale                      |
| Dependency    | Must keep working if a vendor changes terms | A vendor change is an inconvenience            |

A second decision follows the first: hosted or self-hosted. Hosted open models win on cost and effort. Self-hosting wins only when data control, independence or the ability to tune the stack is worth an operations burden, and when you have the hardware already or want it for other reasons.

It is not all or nothing. My agents and research chat run on GLM-5.3-Flash at home. Work that is hard to define, hard to check or expensive to get wrong still goes to a frontier model. The skill worth building is knowing which task is which.

## Sources

- [Getting Started With Open Models](https://every.to/guides/getting-started-with-open-models?ref=corti.com), Every
- [Open models lag state-of-the-art closed models by 4 months](https://epoch.ai/data-insights/open-closed-eci-gap?ref=corti.com), Epoch AI, 29 May 2026
- [Epoch Finds Open-Weight AI Models Trail Closed Frontier by Four Months](https://alphasignal.ai/news/epoch-finds-open-weight-ai-models-trail-closed-frontier-by-four-months?ref=corti.com), AlphaSignal
- [zai-org/GLM-5.3-Flash model card](https://huggingface.co/zai-org/GLM-5.3-Flash?ref=corti.com), Hugging Face
- [zai-org/GLM-5.3 model card](https://huggingface.co/zai-org/GLM-5.3?ref=corti.com), Hugging Face
- [glm-5.3-flash](https://ollama.com/library/glm-5.3-flash?ref=corti.com), Ollama: hosted pricing and Z.ai's benchmark table
- [LibertAIDAI/GLM-5.3-Flash-NVFP4 model card](https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4?ref=corti.com), Hugging Face
- [GLM-5.3 Flash NVFP4 on 2× NVIDIA DGX Spark](https://github.com/LimeChain/glm-5.3-flash-2xdgx-spark-vllm?ref=corti.com), LimeChain
- [NVIDIA DGX Spark specifications](https://www.nvidia.com/en-us/products/workstations/dgx-spark/?ref=corti.com), NVIDIA
- [NVIDIA DGX Spark Review](https://gpusmith.com/articles/en/nvidia-dgx-spark-review-specs-performance?ref=corti.com), GPUSmith
- [Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing?ref=corti.com), Anthropic
- [GitHub Copilot usage-based billing](https://stackfutures.com/blog/github-copilot-usage-based-billing-june-2026/?ref=corti.com), StackFutures
- [Hermes Agent Quickstart](https://hermes-agent.nousresearch.com/docs/getting-started/quickstart?ref=corti.com), Nous Research
- [OpenClaw local models](https://docs.openclaw.ai/gateway/local-models?ref=corti.com), OpenClaw
- [OpenClaw security 101](https://adversa.ai/blog/openclaw-security-101-vulnerabilities-hardening-2026/?ref=corti.com), Adversa AI
- [Cumbersome](https://folding-sky.com/cumbersome?ref=corti.com), Folding Sky