I Benchmarked 27 LLM Configurations Across Two Machines and a Zero-Retention Cloud Tier
A frozen 42-task suite across a 128GB Strix Halo system, a Blackwell GPU workstation, and LM Studio Bionic’s cloud showed why model routing should be based on measured domain performance — not size, vendor tier, or price.Twenty-seven open-weight model configurations. One frozen 42-task suite. Two…
A frozen 42-task suite across a 128GB Strix Halo system, a Blackwell GPU workstation, and LM Studio Bionic’s cloud showed why model routing should be based on measured domain performance — not size, vendor tier, or price.Twenty-seven open-weight model configurations. One frozen 42-task suite. Two machines built on opposing hardware philosophies, plus a metered cloud tier, all scored by the same judge panel under the same rules.I built this harness to make an actual routing decision, which model handles which kind of work, on which machine, before a task ships, not to produce another leaderboard screenshot that stops mattering the week the next model drops.Most public benchmarks compare models in isolation from the hardware or price tier they’d actually run on. That’s a different question from the one that decides a real routing table: given what’s sitting on my desk and what I can call over the network, what’s the highest-quality option that still finishes the task?Answering it meant scoring local, on-prem, and cloud configurations against the identical suite, under the identical rubric, in the same run.The harnessThe suite: 42 tasks across seven weighted domains. Every model configuration, local or cloud, runs the identical frozen prompts against the identical rubric, benchmark v1.2.2.Scoring is judge-based, not string-matched. A four-model panel — Gemini 3.5 Flash Lite, GLM-5.3-Flash, and Xiaomi MiMo v2.5 — scores blind, with DeepSeek V4.1 Flash as tie-breaker. This routes through Openrouter, so the judges sit on infrastructure decoupled from whatever is serving the model under test. A result only counts as final, and eligible for cross-model comparison, at 90% task coverage, 90% completion, and at least medium judge confidence.Fall short on any of the three, and the result is marked provisional and excluded from the headline numbers, whatever the raw score looks like. On this run, 9 of 14 quality-tested Z13 configurations and 6 of 12 P8 configurations landed below that bar, more than a third of everything tested, on hardware I fully control.Where this fitsThe benchmark below covers three leaves of a larger routing setup I run day-to-day.Bionic is the privacy-preserving path — it talks to the Z13 and P8 directly or escalates to the Bionic ZDR cloud tier this piece benchmarks. Hermes is the automation and orchestration layer alongside it: cron jobs, delegated tasks, a kanban board, dispatching work to whichever model provider fits — the same two local machines, or one of three external providers.User-initiated work splits into Bionic’s privacy-preserving local path and Hermes’ automation/orchestration path; both ultimately resolve to a model provider -local (Z13, P8) or cloud (OpenAI, Anthropic, Openrouter)The routing that matters for this piece happens inside Bionic, not Hermes. Bionic defaults every job to the Z13 or the P8 — whichever has capacity and fits the task — and only escalates to the cloud when neither local machine can handle it, through the ZDR-gated tier this piece benchmarks.Hermes runs a separate lane: cron jobs, delegated background tasks, and kanban-tracked work that isn’t privacy-sensitive by design, and it’s free to call OpenAI, Anthropic, or OpenRouter directly, no ZDR gate required.That split is the actual privacy boundary in the stack, not a policy on paper, but which layer a job enters through. Work routed through Bionic never leaves local hardware unless the ZDR guarantee is in force; work routed through Hermes was never treated as sensitive in the first place.The numbers in this piece are about the hardware and the ZDR tier specifically.Keeping it honestTwo things are worth highlighting.First, a control check on the cloud transport. Bionic sessions can’t be hit with a direct API call, so calibration answers move through a batch bridge; and a bridge is exactly the kind of extra hop that can quietly bias a benchmark without anyone noticing.I re-ran GPT-OSS-120B through the bridge as a control and compared it against a direct run: 76.1 through the bridge, 78.4 direct, a −2.3-point offset, well inside the ±3 point tolerance I’d set beforehand. That check is what lets the rest of this piece treat cloud-versus-local comparisons as measured rather than reported.Second, credit where it’s due: the cloud tier here runs entirely through LM Studio’s Bionic pool, under their zero-data-retention (ZDR) policy. Pushing 42 tasks’ worth of prompts and model output through a third party is a real data-handling question the moment any of that content is remotely sensitive, and a verified no-retention guarantee is what made a cloud escalation tier safe to include at all, without turning the benchmark itself into a disclosure risk.I’m deliberately leaving dollar figures out of this piece. The harness tracks per-task cost down to fractions of a cent, but publishing a specific vendor’s rate card isn’t the point, and the quality findings stand on their own.Z13: Strix Halo APU, 128GB unified memoryThe Z13 is a Strix Halo APU mini-PC/laptop with 128GB of unified memory; no discrete GPU, everything sharing one memory pool, which changes which model sizes and quantizations are practical to run. Fourteen configurations went through quality testing.The default I settled on is Qwen3-Coder-30B-A3B. It’s the fastest model on the machine and the eighth-highest quality of the fourteen, with a combined speed-and-quality score of 86.0. This is nineteen points clear of the next entry once both are blended. Its weak spot is financial analysis, where it scores 48.0 against a top domain score of 89.5.The quality ceiling on this machine belongs to Muse-Glimmer-30B, at either quantization -83.9 at Q5_K_M, 83.4 at Q6_K_XL, but it’s the slowest strong-quality option on the fleet. That makes it a deliberate tradeoff for the domains where getting it right matters more than getting it fast: it takes the domain win on financial analysis and research, and places second on legal/risk assessment.GLM-4.7-Flash sits at the floor of the tested field on both speed and quality. For these use cases, the config to skip.The quantization delta on Muse-Glimmer deserves a closer look: Q5_K_M beats Q6_K_XL by 1.9 points on financial analysis despite being the smaller-footprint quant — the opposite of what bit-width alone would predict.Domain-level testing catches that. An aggregate score buries it.14 locally-hosted configurations × 7 benchmark domains · judge-scored, 0–100 · ★ = top score in that domain · sorted by overall qualityP8: CUDA workstation, Blackwell-generation GPUsThe P8 is a CUDA workstation on Blackwell-generation GPUs, discrete VRAM, no unified-memory ceiling, the machine built for pushing larger active-parameter counts. Twelve configurations went through quality testing here.Qwen3.5–122B-A10B posts the best quality score among eligible configs, 84.6, and the best long-context score in the entire set, 96.2; a mixture-of-experts model at roughly 10B active parameters that earns the size.For latency-sensitive work, Qwen3-Coder-Next is the faster eligible default: 100% completion, final status, and the strongest research/evidence score on the machine.Qwen3-VL-235B-A22B (Instruct) puts up the standout figure of the whole assessment, 95.4 on document fidelity, but its throughput collapses at long context, so it reads as a document-review specialist rather than a generalist.Two configurations are flagged as do-not-deploy-as-configured. Qwen3.5–397B-A17B (Thinking) truncated on 47.6% of tasks and came back blank on most of those. Qwen3.8–27B (Instruct) posted the lowest quality score on the machine, 46.3.Neither failure is subtle, and both only surface when you run the full suite instead of spot-checking a handful of prompts.Two more things worth flagging for anyone reading a benchmark at face value. Two P8 configurations, Qwen3.5–397B-A17B (Instruct) and Qwen3.5–122B-A10B, changed completion behavior between passes, jumping from roughly 58–60% to 98–100% with no change on my end. That’s a serving-configuration difference I still need to pin down.The P8 quality run also swapped three of the four Openrouter judges from the prior pass — a panel change, not just a model change — so this run is treated as a rebaseline rather than claimed as clean comparability with earlier P8 numbers.12 locally-hosted configurations × 7 benchmark domains · judge-scored, 0–100 · ★ = top score in that domain · sorted by overall qualityLaid side by side, the two machines make the trade-off visible in one picture: the Z13 clusters toward the low-speed, high-quality corner whenever it’s not running its fast default, while the P8 spreads out more evenly across the speed axis.Bubble size adds a third axis, the quantized model footprint, and it shows something a flat scatter can’t: parameter count and quality don’t move together on either machine. Some of the largest bubbles sit mid-field. Some of the smallest hold their own against configs many times their size.12 configs per machine · bubble area = total parameters (quantized footprint) · Speed Index is field-relative, not cross-machine absolute · judge-scored quality, 0–100The cloud tier: quality clusters, cost doesn’tFour Bionic ZDR cloud configurations cleared calibration, all final, all 100% completion. The spread here isn’t in quality; all four land within 3.5 points of each other. It’s in cost, which spans roughly a 21x range across the tier.Consider that for a second. Whatever the most expensive model in this pool is charging for, on this evidence it isn’t a proportional quality gain. The cheapest model in the tier, GLM-5.3-Flash, scores 92.77. The most expensive, Kimi K3, scores 92.46; lower, at over twenty times the run cost. DeepSeek V4.1 Flash takes the top overall score, 93.00.It’s neither the cheap option nor the expensive one. It’s the second-cheapest. Again, recall this is for my 42 business use cases; your mileage may vary.A similar version of the same pattern shows up inside a single model family, where price and lineage should track each other most closely. GLM-5.3 and GLM-5.3-Flash share an architecture; GLM-5.3 costs roughly nine and a half times more to run the identical battery. On overall quality, it scores three points below its own cheaper sibling.On document fidelity, extraction and structuring work, it collapses to 77.4 against 92.6 for the Flash model it’s supposed to supersede. Fifteen points, same lab, same family, in the wrong direction.Don’t route by vendor tier instead of by measured domain performance, that becomes the failure mode: the bigger, pricier sibling isn’t a strict upgrade, and averaging across domains hides exactly where it loses.4 calibrated cloud configurations × 7 benchmark domains · judge-scored, 0–100 · ★ = top score in that domain · sorted by overall qualityIf you’re building your own harness• Score by domain, not just in aggregate. Every regression in this dataset that mattered — the Muse-Glimmer quantization flip, the GLM-5.3 document-fidelity collapse — is invisible in the overall number but obvious in the domain breakdown.• Set your eligibility bar before you see the scores, not after. Coverage, completion, and confidence gates only do their job if a model that clears 89% doesn’t quietly get rounded up because the number is inconvenient. On this run, more than a third of the local configurations landed in “provisional” and stayed there.• Run a control before you trust a bridge. Any indirection between your harness and the model under test — a proxy, a batch transport, a retention-compliant relay — is a place bias can enter unmeasured.Rerun something you already trust through the same path and check the delta against a tolerance you set in advance.• Price and quality are two separate axes — measure both, but don’t assume either predicts the other. The cheapest option in a tier and the best-scoring option in the same tier weren’t the same model, and the most expensive option was neither.• Don’t average across a judge-panel swap. If your panel composition changes between runs, treat the new run as a rebaseline, not a continuation of the same trend line, even when the harness and task suite are otherwise identical.What’s nextNone of this is a final verdict on any of these models.Pricing on the cloud tier resets monthly, and the next open-weight release could reshuffle the local rankings. What should hold up longer is the method: score by domain, gate on completion and confidence before you compare, control your transport, and keep price and quality as two separate measurements until the data says otherwise.I’ll rerun the suite the next time enough of the roster turns over, and I’ll publish what changes.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!I Benchmarked 27 LLM Configurations Across Two Machines and a Zero-Retention Cloud Tier was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI