
- I asked five small local models, all 4-bit, running on my M2 Max, to build the same Apple-style todo app through the same agent harness.
- Qwen 3.8 27B produced the most polished app, but its winning session actually ran FP16, not 4-bit. Among genuine 4-bit runs, Qwen 3.6 35B and Ornith run 3 are the ones to beat.
- My first results were poisoned by config, not by the models: a presence penalty of 1.5 on Ornith, a corrupted speculative-decoding path on Qwen, a fantasy 1M context limit on Nemotron.
- The same model went from flat and generic to 17 of 17 tests passing with clean lint between two runs. The only things that changed were a skill file and an AGENTS.md.
- Small models are good at copying gold-standard reference files and bad at opt-in behaviors: loading skills, running lint, cleaning up scaffold.
- Thinking budgets didn't help these models build. Measured: a 1,024 token thinking budget produced planning prose and zero code.
- I packaged everything that worked into a repo of skill packs, enforcement gates and a project template, so the next run starts at roughly 70% craft instead of zero.
The experiment
The goal was simple. I wanted to know what small local models can actually build in 2026, on my Mac.
This goal as you will later see turned out to be "how can I improve the output quality of my 4 bit models without training". I decided to, first of all, test what can you actually build, secondly, can I fix the mistakes the models make.
I ran the same prompt through every small model I had, one at a time, through the same coding agent harness (zcode), served locally (oMLX and MTPLX) on my M2 Max with 64 GB of memory.
The prompt:
I want to build a web app which will be a todo app, you can add items to the todo
list and then delete them, each item should have the option to add subitems, no
login so far needed, we need an option to save the todo lists and remove them as
well, this needs to be visually looking like an apple product, nice rich and
polished ui. prepare a full implementation plan, you can check internet for
similiar examples if needed
The idea was - a typical prompt, a bit vague, but not a voxel pagoda, not a threejs animation. I want to build an actual app.
A quick word on the models, because the names hide the details. Most of these are Mixture of Experts models (MoE): the full model is large, but only a small slice of it, a few billion parameters, fires for each token. I picked this class on purpose, because I run a Mac.
Generating each token means reading the active weights from memory, and on Apple silicon memory bandwidth is the bottleneck. A dense 27B reads all 27 billion weights for every token it emits. An A3B MoE only wakes the ~3 billion it needs, roughly a ninth of the bytes per token, so it decodes that much faster on the same machine. On my M2 Max that's the difference between the dense Qwen 3.8 crawling at around 20 tokens per second and the MoE models shipping code at around 60 or more. The sleeping parameters aren't wasted either: total size is where knowledge lives, active size is what you ‘pay’ per token. An A3B model knows what a 35B model knows and runs like a 3B one. For a 64 GB Mac that's the sweet spot, and 4-bit quantization (each weight stored with 4 bits instead of 16, roughly a quarter of the memory) is what makes these sizes fit at all.
The fleet:
| Model | Size | Active params | Runs |
|---|---|---|---|
| Qwen 3.8 27B | dense 27B | 27B | 1 (plan + build) |
| Ornith 1.5 35B-A3B | MoE | ~3B | 3 |
| Gemma 4 26B-A4B | MoE | ~4B | 1 |
| Nemotron 3.5 Lightning 30B-A3B | MoE + Mamba2 hybrid | ~3B | 3 attempts |
| GLM-4.7 Flash | MoE 30B-A3B | ~3B | 2 attempts |
| Qwen 3.6 35B-A3B (comparator) | MoE | ~3B | 1 |
Afterwards I ran a full forensic analysis of every session: the session database, raw model logs, the serving configs, then fresh builds, tests and lint on every app folder. The full report and the per-run notes were used to compile this article. Everything below comes from that evidence.
Two honest caveats before the findings, because they change how you read them:
- The winning Qwen app was not built by the 4-bit model. The planning session ran Qwen 3.8 at 4-bit, but the 62-call build session ran on my FP16 serving path. So the polished winner does not prove "4-bit Qwen did this". It proves dense 27B at full precision did this. This is exactly the kind of detail that gets lost in most benchmarks, and it was sitting right there in the folder name. I didn’t plan this I was loading different models and actually ran the session with that model, decided not to re run it, maybe will do later for completeness.
- This is not a controlled benchmark, and it shouldn't read as one. One run per cell (three for Ornith and Nemotron), models sample outputs so they're not deterministic, and my serving config changed across the week. Most importantly, the models did not run identical settings: different precision paths, thinking on for some and off for others, sampling tuned per model, sometimes diverging from the vendor's own recommendation. That was deliberate. I was trying to get the best out of each model, not to hold them under laboratory conditions. Read this as field notes from real runs with the drift flagged where it matters, not as a leaderboard.
First: config bugs poisoned my early results
Before you compare models, you have to trust the plumbing. I found three bugs that had nothing to do with model quality, and each one had been quietly wrecking a run.
| Bug | Where | What it did |
|---|---|---|
presence_penalty: 1.5 | Ornith serving config | Penalizes every token that already appeared. Code repeats brackets, keywords and JSX tags constantly. The model was being punished for writing valid code. Output came out flat and mechanical. |
| Fantasy context limit | Nemotron config | zcode was told the model had 1M context. It never compacted, prompts grew, and generation stalled for up to 24 minutes. |
The presence penalty one is worth sitting with. Code is repetition. A presence penalty pushes the model away from repeating tokens, which sounds harmless for chat and is poison for const, {, </div>. Ornith run 1 was scored on a broken config, and its output looked exactly like you'd expect: broad functionality, zero taste.
After fixing these, re-runs looked completely different. First lesson from the whole project: a raw model A versus model B comparison on a local stack means nothing until you've verified the serving config.
What each model actually delivered
The merged scorecard. "Undefined tokens" means CSS variables the app used but never defined, which renders as invisible borders and broken styles. Aria counts are accessibility attributes in the source. Numbers measured from the final code trees and session logs. I judged every app against the same weighted rubric: functional completeness and UI craft 25% each, interaction and motion 15%, then accessibility, architecture, verification and agent process.
| Run | Build | Tests | Undefined CSS tokens | Aria | Verdict |
|---|---|---|---|---|---|
| Qwen 3.8 27B (FP16) | yes | 3 fail (stale tests) | 0 | 30 | best craft, best chrome |
| Ornith run 1 (4-bit) | yes | none | 8 | 15 | functional, generic |
| Ornith run 2 (4-bit) | yes | 13/13 | 0 | 35 | great app, 746 calls |
| Ornith run 3 (4-bit) | yes | 17/17 | 2 | 17 | best 4-bit discipline |
| Qwen 3.6 35B (4-bit) | yes | 18/18 | 0 | 26 | best all-round 4-bit |
| Gemma 4 26B (4-bit) | yes | none | 9 | 0 | fast prototype, thin |
| Nemotron (4-bit) x3 | no | none | 0 | 4 | failed all three |
| GLM-4.7 Flash (4-bit + 6-bit) | no | none | 0 | 0 | failed completely |
Qwen 3.8 27B: the reference bar, with an asterisk
This was the only dense model in the test, and it showed. It planned in one session (asked me two good questions, chose a two-pane layout, wrote a real plan), then built in a second session: 62 model calls, 27 files, about 1,760 lines of source.
The craft is what separates it from everything else. A checkbox that animates a spring and draws its checkmark with an SVG path. Frosted sidebar materials. A centered 680px reading column inside real app chrome instead of a floating web page. 30 aria attributes, keyboard handlers, three tiers of shadow mapped to elevation. It reads like an authored design system, not styled divs.

Its one real defect: it shipped a test suite with 3 failing assertions. I checked them, the failures are stale test expectations (a helper comparing an item to itself, a progress count expecting 3 on a four-node tree), not app bugs. But that's the same disease in nicer clothes: it wrote tests and never ran them green. I've since retired the dense 27B from daily use. At roughly 20 tok/s on my machine it's too slow to live in, even though it's the quality bar.
Ornith 1.5 35B: three runs, three configs, one clear trend
This is where the story gets interesting, because the model stayed the same and everything around it changed.
Run 1 (4-bit, thinking on, presence penalty bug): the most feature-rich app of the early runs. Search across lists, bulk actions, JSON export and import. Also: 6 motion elements, 13 aria attributes, 8 undefined CSS tokens, scaffold leftovers, no tests. Functional, flat, generic. This is the "poisoned config" baseline.
Run 2 (4-bit fast profile, thinking off, penalty fixed, plus my new apple-ui-craft skill and AGENTS.md): the quality jumped. Custom scrollbars, pathLength checkmark, 35 aria attributes, 13 passing tests, real design tokens. Components copied from the skill's reference files came out byte-identical.
But the session itself was a disaster. The model invented an 85-minute browser-audit rabbit hole: a subagent made 213 tool calls and burned 12.9M tokens chasing one click bug. In total, run 2 spanned ten related sessions and 746 model calls, with 3 hours 17 minutes of wall time of which only 43 minutes was actual model generation. Along the way it created a tailwind.config.ts, which silently breaks Tailwind v4's CSS-first setup, and spent 11 minutes debugging the CSS crisis it had caused itself.
Run 3 (4-bit main profile, thinking on with a 1,024 token budget, hardened skill, AGENTS.md): 43 minutes total, 102 calls, zero subagents, 17 of 17 tests passing, clean lint, clean repo. Empty states came out near-verbatim from the skill's gold-standard reference. The Apple easing curve showed up 5 times.
Remaining defects: 2 undefined CSS tokens (the unchecked checkbox ring was literally invisible), and it never landed the plan.md I asked for, because plan mode blocked the write and it never followed up. Small stuff, but it's the difference between a good run and a great one.
The trend across the three runs is the single biggest finding of this whole experiment: the model didn't change. The config, the skill and the guardrails did.
Gemma 4 26B: the fastest run and the emptiest
27 minutes, the fastest of the fleet. The core functionality is complete: nested subitems, multiple lists, clean persistence middleware. And it's the only run that shipped a frosted sidebar without being told, which still makes me smile, because three "smarter" runs missed it.
Everything else is the cautionary tale. It never loaded the craft skill that was sitting right there in its environment. Zero tests. Zero aria attributes in the entire app. No dark mode. 9 undefined CSS tokens. Scaffold leftovers and a package name of my-app. It failed lint on a real bug: a conditional useState call, which can corrupt React's hook order across renders. 8 of its 12 edit operations failed on string matching and it needed 2 nudges from me to continue.
The guidance was there and went unused. That's not a knowledge problem, it's an opt-in problem, and it repeats across models. Which is why my next iteration stopped asking nicely.
Nemotron 3.5 Lightning: three failures, three different autopsies
This one never produced a buildable app, and each of the three failures died differently.
- Run 1 failed on infrastructure: the 1M fantasy context limit which I’ve missed, generation stalls up to 24 minutes, a fragmented shared cache, and a rogue subagent that scaffolded a duplicate project in the wrong place.
- Run 2 failed on a quantization artifact: the model dropped a single structural token in generated JSX (whileTap {{ ... }} missing its =), and when it rewrote the file, it reproduced the same dropped token. It cannot heal its own quantization artifact. Then it fell into full text degeneration: "Writing EmptyState.tsx:" streamed 70+ times until I stopped it.
- Run 3 was the clean experiment: fixed sampling, cleared cache, 64K cap. Below roughly 45K of context it produced real, clean code. Past 50K it started looping, phrases repeating 4 times, then 7, and I stopped it at 73K.
That last run is the diagnosis. Nemotron's architecture is a Mamba2 hybrid: only 6 of its 52 layers are attention, the rest is recurrent state. At 4-bit, quantization noise compounds in that state during decoding, and with almost no attention layers to re-read its own transcript, it drifts into repetition loops on long sessions. NVIDIA's own docs flag long-context retrieval as this architecture's weak spot [4].
Same task, same 4-bit width, standard transformer MoE (Ornith): finished in 43 minutes with zero loops. So this is not "4-bit is bad". It's architecture x quantization x session length.
One fairness note: my local settings had thinking off and temperature 0.6, while NVIDIA recommends thinking on at 1.0 [4]. Their failure here is well-evidenced, but "this model can never code" is not a claim I can make from this setup.
GLM-4.7 Flash: failure before and during implementation
The dedicated 4-bit run folder is empty. The folder wasn't thin or broken, it was empty. The second attempt accumulated 241 model calls across five sessions and 100 tool errors, and the surviving 6-bit codebase (429 lines) fails at the toolchain foundation: missing React and Vite resolution, JSX not enabled, import configuration conflicts.
Honest boundary: I removed both GLM models locally after this, so I can't reconstruct the full 4-bit server config anymore. And GLM's own model card reports strong SWE-bench numbers under a very different harness, precision and sampling regime [5]. What I can say is narrower and still true: both of my local GLM attempts failed to produce a buildable app.
Qwen 3.6 35B: the quiet winner of the 4-bit class
I added this run late as a comparator, and it won the head-to-head against Ornith's best disciplined run. Thinking off, the fast variant, roughly 35 minutes, 41.5K output tokens (Ornith run 3 used 89K to get where it got). 18 of 18 tests passing, zero undefined tokens, a real plan.md in the repo root, a 5-color list picker, 26 aria attributes. At around 82 tok/s it's the best local coder I have right now.
Its two misses are small and telling: it defined the frosted .material class and then never applied it to the sidebar (the third consecutive model to do that), and it promised debounced storage writes in its plan and shipped direct writes.
The exact run settings
If you're going to reproduce any of this, the settings are half the story. One mechanism matters more than any single value: the zcode harness sends no sampling or thinking parameters at all. Every setting below came from the server-side profiles, which means config drift between runs is a first-class experimental variable, not a footnote.
| Run | Precision | Thinking | Sampling (server logs) | Outcome |
|---|---|---|---|---|
| Qwen 3.8 plan | 4-bit | off | temp 1.0, top-p 0.95 | plan only |
| Qwen 3.8 build | FP16 | on | server profile | best app |
| Ornith run 1 | 4-bit | on, no budget | presence penalty 1.5 (bug) | flat, generic |
| Ornith run 2 | 4-bit fast | off | top-p 0.95, top-k 20 | great app, 746 calls |
| Ornith run 3 | 4-bit main | on, 1,024 budget | top-p 0.95, top-k 20 | best 4-bit process |
| Gemma 4 26B | 4-bit | off | temp 0.6 (vendor says 1.0) | fast, thin |
| Nemotron x3 | 4-bit | off | temp 1.0, then 0.6, then 0.35 | failed all three |
| GLM-4.7 Flash | 4-bit + 6-bit | off | temp 1.0, then 0.6 | no buildable app |
| Qwen 3.6 | 4-bit fast | off | temp 0.7, top-p 0.95 | best all-round 4-bit |
One more trap when reading that table: "4-bit" is too coarse a label by itself. These runs used different quantizers (plain MLX 4-bit, OptiQ mixed builds that keep sensitive layers like attention at 8 bits, oQ4), and per Hugging Face's quantization docs the accuracy you keep varies by method and task [7]. Same bit width, different recipes.
Two things jump out of that table. First, my Gemma and Nemotron settings diverged from their own model cards' recommendations [3][4], so part of their underperformance is on me, not them. Second, thinking on did not beat thinking off for these small models, which deserves its own paragraph.
The thinking budget finding. Same model, same task, measured: thinking off produced well-structured code in 8.5 seconds. Thinking on with no budget took 19.6 seconds and wrote planning prose first. Thinking on with a 1,024 token budget took 21.6 seconds and spent the entire budget on planning prose. Zero code. A ~3B-active model doesn't plan better with more thinking, it drafts everything twice. Budgeted thinking stayed useful for Ornith's plan phase, and thinking off was the right choice for every build session.
Qwen 3.8 is the counterweight. The dense 27B's build session ran with thinking on and recorded 47,225 reasoning tokens, and that session produced the best app of the experiment. Its chain of thought is strong enough to be worth the tokens. Its planning session ran thinking off and still produced a strong plan, so thinking isn't what made the plan good. Would the build have been even better thinking off? I don't know, I never ran that cell, and I'm not going to dress up a guess as a finding. What the data does support: in the 3B-active MoE class more thinking measurably hurt, and in the dense 27B it at least didn't. Until I run the missing cell, my working rule is to let model size decide. Small models think short or not at all, big models can afford to think.
Why I added a skill and an AGENTS.md
Run 1 knew how to implement todos. It did not know how to make something that feels like an Apple product, and no amount of "polished UI" in the prompt fixed that.
So after run 1 I built the apple-ui-craft skill: hard rules for the design language (typography scale, one accent color, 8pt spacing, elevation tiers), a motion language with real numbers (spring stiffness 300 to 400, stagger 40 to 120ms), and, the part that actually worked, gold-standard reference files copied from the Qwen app. A full token system, the animated checkbox, the confirm sheet, the theme hook, the recursive row.
AGENTS.md carried the process side: cheap verification (build, test, lint), no browser automation unless asked, no rabbit holes, read before you edit, clean as you go.
The results, Ornith run 1 versus runs 2 and 3:
| What changed | Run 1 | Runs 2 and 3 |
|---|---|---|
| Motion elements | 6 | 10 to 11 |
| Aria attributes | 13 | 17 to 35 |
| Tests | 0 | 13 and 17, passing |
| Undefined CSS tokens | 8 | 0 (run 2), 2 (run 3) |
| Scaffold leftovers | 3 | 0 |
| Subagent rabbit holes | some | 0 (run 3) |
And the sharpest pattern I found in this entire experiment: whatever exists as a reference file gets copied correctly. Whatever is only prose in the skill gets dropped. The empty states were a reference file, and run 3 reproduced them near-verbatim. The frosted sidebar was a prose rule, and three different models across three runs dropped it. That pattern is why the packs repo now ships a Sidebar.example.tsx.
Gemma is the control group for the other direction: the skill was available and it never opened it, and its output looks exactly like run 1's config-poisoned output.
The mistakes small models make
Every failure mode I watched, collected in one place. None of these are hypothetical, each one happened in a session log.
- They stop at visible. UI on screen feels like done. Build passes feels like done. Tests written feels like done. Closing the loop, build plus tests plus lint all green, was the exception, not the rule.
- They opt out of available help. Gemma had a craft skill and never loaded it. Qwen 3.8 wrote tests and never ran them.
- They invent rabbit holes. Run 2's 85-minute browser audit, 213 subagent calls, 12.9M tokens, for one click bug. Left alone, "improve until perfect" eats your evening.
- They repeat failed strategies. Run 3 retried the identical failing edit 4 times. Nemotron rewrote the same store file 4 times and reproduced its own dropped token each time.
- They drop single structural tokens. A missing = in whileTap broke a whole build, and the rewrite carried the same missing =.
- They collapse on long context. Nemotron was clean below ~45K of context and looping past 50K, on a 4-bit recurrent architecture. Long agent sessions are quant noise compounding.
- They hallucinate packages. @types/zustand doesn't exist. npm said so four times.
- They leave undefined CSS tokens. 8, then 2, then 9 across three runs. Unresolved variables silently render as nothing, so the app looks broken in ways the model never sees.
- They leave scaffolding behind. hero.png, vite.svg, template READMEs, a package still named scaffold. Every model that started from a bare npm create vite shipped leftovers.
- They break frameworks they don't know. A tailwind.config.ts on Tailwind v4 silently kills all utility generation, and the model then debugged the wrong thing for 11 minutes.
- They drift between plan and implementation. Debounced writes promised, direct writes shipped. plan.md requested twice, never landed.
- They misuse tools. Six consecutive failed heredoc file writes instead of the Write tool. Write-before-read failures in a row.
What I built to skip those mistakes
Watching the same ten failure modes across models is what pushed me to stop relying on prompting. If a model only follows rules sometimes, rules aren't enough. The behaviors either happen mechanically or they don't happen.
So I built local-llm-packs, a repo that moves capability out of the model's weights and into the environment. It's built directly from these run analyses, and it works with zcode, pi and OMP. Four layers:
Knowledge: skill packs with reference files. Five packs (app shell craft, data viz, forms and CRUD, landing pages, API-backed apps), each with a SKILL.md plus gold-standard example files. The lesson from the todo runs, reference files get copied, prose gets dropped, is baked into the format. The sidebar example exists because three models dropped the prose version of that rule.
Process: enforcement gates. Five hooks that apply to every model uniformly:
- A skill-load gate: writes are blocked until the matching pack is loaded. (Gemma's failure, made impossible.)
- A verify gate: the turn can't end while build, test, lint or the token check is red, capped at 3 blocks so a weak model can't ping-pong forever. (The "stops at visible" failure, made impossible.)
- A hygiene gate: scaffold leftovers, template READMEs, missing plan files get flagged. (The leftovers disease.)
- An edit-failure redirect: after 2 failed string-match edits, it's told to rewrite the whole file. (Gemma failed 8 of 12 edits that way.)
- A scaffold redirect: bare npm create vite gets intercepted in favor of a house scaffold that starts from tokens, theme and a verify script.
Starting state: a pre-scaffolded template. Every run starts at roughly 70% craft instead of zero, because the token system, theme hook, verify script and app shell already exist.
Measurement: a battery. Five canonical tasks (the original todo prompt, a finance dashboard, a contacts manager, a landing page, an API-backed notes app) with automated scoring that re-runs verify and craft metrics per run, plus historical baselines from these very runs so there's something to compare against.
There's also a toggle that strips the whole package back to raw mode, so I can A/B test what a model does with and without it in a clean single-variable experiment. Honest status: the repo ships with a test suite for the gates, but the live validation pass, watching the gates fire on a real small-model run, is still on my list.
One more lever I haven't wired up yet is review. The small model drafts, deterministic tools verify, and a stronger model reviews only the failing diffs and the architecture, then the small model applies bounded fixes. The run analyses kept pointing at this hybrid loop: review tokens are a fraction of writing tokens, so you buy most of the quality for a fraction of the cost. It's on the packs roadmap, not enforced yet.
I haven’t tested this on any model yet. I was inspired to build something like this to go further down the path of adding more equipment to the small models, universal assets they could reuse to improve the quality of their outputs.
I will test this and think if I could have a universal reusable asset pack to improve the quality of outputs for my small models.
What small models like these are good for
Based on what I actually watched them do, not on vibes:
- Extending an existing, well-structured codebase. This is where they shine. Given tokens, references and conventions, they produce clean work.
- Copying and adapting reference patterns. Byte-identical, every time. This is the most reliable thing a small model does.
- Fast iteration. 82 to 85 tok/s means a change lands in seconds, not minutes. The edit loop feels different at that speed.
- Bounded tasks with hard verification: small bug fixes with a reproducer, unit tests for stable interfaces, single-module refactors, types from schemas, docs and comments.
- Everything privacy-sensitive and offline. Personal documents, notes, local RAG over your files, email drafts, classification and tagging, structured transformations. Runs on your machine, nobody can reprice it.
- First-pass work that a human or a bigger model reviews: local code review for obvious errors, drafts, summaries.
Where they fail usually
- Greenfield, multi-file builds where taste is implicit. Without references, you get run 1: functional, generic. The craft doesn't come from the weights.
- Closing the loop unprompted. They will not reliably run build, test and lint and fix what's red. Gate it.
- Long autonomous agent sessions. Especially at 4-bit on recurrent architectures. Nemotron was clean below 45K context and looping past 50K. The advertised 256K window is not your working window.
- Precision editing. String-match edits are near-unusable at the 4B-active tier. Rewrite-the-file workflows or edit-failure gates fix this.
- Consistency across many files. Types created three files apart didn't stay compatible. The model doesn't hold the whole project in its head.
- Anything security-sensitive or high-stakes without review. They produce plausible local patches that can violate project-level invariants.
The one-line version: small 4-bit models are capability-rich and discipline-poor. Every fix that worked in this experiment added discipline from the outside: references, gates, templates, verification. None of them added intelligence to the weights.
I went looking for a bigger MoE that fits. Haven’t found one yet.
The obvious train of thought for me was: if I know what 3B-active delivers at 4bit, wouldn't 9B-active or 13B-active be better, as long as it still fits 64 GB? I asked Hermes to scan the open-weights market for exactly that. Short answer: nothing available right now both fits my Mac at 4-bit or better and clearly beats the Qwen 3.6 / Ornith class for agentic coding.
What I checked, with published numbers where they exist (all publisher-reported or community-measured, none of these are my runs):
| Candidate | Total / active | Fit on 64 GB | Why it didn't make the cut |
|---|---|---|---|
| Qwen3-Next-80B-A3B | 80B / 3B | 4-bit only, ~45 GB | more total knowledge, same active compute, slower, published coding scores below this class |
| Nemotron-Labs-3-Puzzle-75B-A9B | 75B / 9.3B | mixed 4/6-bit MLX, ~45 GB | the most interesting one: community-tested on an identical M2 Max at ~14 tok/s, about 6x slower than Ornith, runtime support still in open PRs |
| Hunyuan-A13B | 80B / 13B | MLX Q4, ~45 GB | older generation, slower, published coding profile not better |
| Ling 3.0 Flash | 124B / 5.1B | full model ~70 GB, doesn't fit | the fitting community build is one-shot expert-pruned with no recovery training |
| Laguna S 2.1 | 118B / 8B | Q4 ~58 GB, no headroom | the one real Ornith competitor on paper, only fits at 3-bit |
| GPT-OSS-120B, Llama 4 Scout, Qwen3.5-122B | 109B to 122B | don't fit usefully at Q4 | the memory wall |
For scale: the models I actually run fit comfortably, Ornith 1.5 even at Q8 (about 38 GB) while still decoding at 80 to 85 tokens per second, and everything that adds active parameters subtracts speed and headroom. None of the candidates adds enough quality on paper to pay that bill. Which is exactly why my next levers are the environment (the packs) and precision (6-bit), not size. The moment a 9B-active class model ships that fits at 4-bit with 2026-grade post-training, I'll run it through this same todo test.
What's next
- 6-bit runs of the same fleet. Everything in this experiment was 4-bit. The one 6-bit attempt (GLM) failed for its own reasons, so higher precision on this task is still an open question, and it's the next thing I'll test. If 6-bit removes the dropped-token and repetition artifacts, that alone might be worth the extra memory.
- Testing my local pack repo, with gates on. Same five tasks, the packs installed, scorecards per model, compared against the pre-package baselines from these runs.
- The raw-mode A/B. Same model, same prompt, packs toggled off. That's the clean measurement of how much the environment is worth.
Hope you found this useful and you learned a bit from my experiment. I am far from being an expert, I love experiments and I gave myself a task to improve usefulness of models like these.
More tests to come and better tests to come as I will improve future runs based on my current experiences and mistakes
Sources
Web sources, all checked September 1, 2026:
- Qwen3.8-27B model card: https://huggingface.co/Qwen/Qwen3.8-27B
- Ornith-1.5-35B-A3B model card: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B
- Gemma-4-26B-A4B-it model card: https://huggingface.co/google/gemma-4-26B-A4B-it
- NVIDIA Nemotron 3.5 Lightning 30B-A3B model card: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
- GLM-4.7-Flash model card: https://huggingface.co/zai-org/GLM-4.7-Flash
- Apple WWDC 2025, Explore large language models on Apple silicon with MLX: https://developer.apple.com/videos/play/wwdc2025/298
- Hugging Face, Selecting a quantization method: https://huggingface.co/docs/transformers/en/quantization/selecting
- MLX LM documentation: https://github.com/ml-explore/mlx-lm
- NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B model card: https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
- Community mixed 4/6-bit MLX build of Puzzle with M2 Max measurements: https://huggingface.co/tekosML/Nemotron-Labs-3-Puzzle-75B-A9B-MLX-4bit-experts-6bit
- Tencent Hunyuan-A13B-Instruct model card: https://huggingface.co/tencent/Hunyuan-A13B-Instruct
- Qwen3-Next-80B-A3B-Instruct: https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
Candidate numbers for models I did not run come from the publisher cards above plus community pages (benchlm.ai, atomic.chat) catalogued in my research session; nothing is my own measurement except where the text says I measured it.
