AI Efficiency Toolbox
Back to videos

September 29, 2026 · 5:35 · Evidence through September 24

14 local AI setups on a real app

Signal stood out on one Swift parser task. Here are the outcomes, recorded settings, independent checks, and files behind the video.

Watch on YouTube: I Tested 14 Local AI Setups on a Real App — Signal Stood Out

What the test showed

7 / 14

local setups passed the reported task criteria

Signal Q6: 11m21s, 91/91 tests, Debug build passed. The frozen patch passed independent evaluation without evaluator repairs. This was the fastest passing local observation.

The task added U+2011 non-breaking hyphen support to verse ranges, normalized accepted separators, and preserved visible Markdown text and decoded URL references. Four authored regression tests accompanied two production-line changes.

Self-written tests were insufficient: older Qwopus 3.6 Q4 passed its own tests using the wrong Unicode character. Fixed independent acceptance checks caught the mistake.

Compare the recorded runs

Bar length shows recorded duration, not quality or a controlled speed ranking. A fast failed run did not complete the task. Captured spans and termination times are labeled separately; missing times have no bar. Select a setup for its settings and caveats.

Showing 14 of 14 setups · 7 reported passes

Selected setup

Signal 3.8 27B · Q6

AP-Q6_K GGUF · medium · Pass

Runtime
Standalone llama.cpp 2.37.0 / Hermes
Recorded duration
11m 21s · recorded run
Settings
118,016 context · q8_0 K/V · native MTP max 4 · backend draft sampling off · BF16 vision projector loaded · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 · same harness/runtime settings as Signal IQ3_S
Evaluation
Tests: 91 · failures: 0 · fixed acceptance failures: 0 · build: Pass

What to keep in mind

Fastest passing local observation: 11m21s, 28.7% less time than Signal IQ3_S and 42.5% less than GSQ-RCO. 24 tool calls versus 36 and 22 respectively. Simpler two-line production patch; four authored tests with all-dash full Markdown links and direct raw-variant URL round-trips. Recovered from four URL-encoding assertion failures; corrected own 89-test suite passed, then independent 91/91 and build passed without repair. File search instead of graph, production-before-test ordering, /tmp diagnostics and piped exit-code handling remain workflow limitations. Unrelated TypeScript check lacked tsc and was disclosed. One run per configuration; cache, machine state and work trajectory can affect timing.

Comparison scope and caveats

One real Swift parser task, with different models, settings, dates and runtimes. Reported pass status and timing coverage are separate. The complete evidence entries and downloadable results remain below.

All 14 local setups

Each entry is a complete setup, including runtime and reasoning tier. Open an entry for its settings and evaluation. Unknown values remain unknown.

Qwen3.8 27B

Q6_K GGUF · medium · 2026-08-24

Pass52m 40s
Settings and evidence

Correct two-file patch. Repeated planning and test launches delayed completion; first write took nearly 35 minutes.

Runtime: LM Studio / llama.cpp

Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens

Independent combined suite: 92/92 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Grug v1.1 27B

Q6_K GGUF · medium · 2026-08-24

Fail5m 22s
Settings and evidence

Fast but missed U+2011; inserted U+200B removal and repeated ASCII hyphen. Wrote tests but did not run them.

Runtime: LM Studio / llama.cpp

Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens

Independent combined suite: 78/90 passed. Fixed acceptance assertion failures: 8. Debug build: Pass.

Kiwen1.1 27B

Q6_K GGUF · medium · 2026-08-25

Pass88m 43s
Settings and evidence

Eventually correct after a long regex repair loop. 68 tool calls and 88.7 minutes; also expanded support to U+2010 and wrote scratch files outside its worktree.

Runtime: LM Studio / llama.cpp

Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens

Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Grug v1.1 27B

Q6_K GGUF · xhigh · 2026-08-25

Fail4m 23s
Settings and evidence

Verified native xhigh did not repair correctness. Recognized U+2011 but failed normalization; did not run its useful regression test.

Runtime: LM Studio / llama.cpp

Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens

Independent combined suite: 69/88 passed. Fixed acceptance assertion failures: 12. Debug build: Pass.

Qwen3.8 27B

4-bit MLX · medium · 2026-08-26

No patch69m 15s
Settings and evidence

Diagnosed the intended fix but never applied it. Two Metal out-of-memory retries, unsupported structured requests, and long non-tool responses. Only wrote a scratch harness.

Runtime: mlx_vlm.server

Settings: 262,144 loaded context · MTP max 3 · KV 8-bit / group 64, starts at token 5,000

Independent combined suite: 71/87 passed. Fixed acceptance assertion failures: 16. Debug build: Pass.

Qwopus3.6 35B A3B

oQ4 MLX · medium · 2026-08-26

Fail33m 42s
Settings and evidence

Its own 92 tests passed, but production code and tests shared the same wrong character: U+200C instead of U+2011. Independent checks caught it. Two Xcode diagnostic timeouts also increased wall time.

Runtime: mlx_vlm.server

Settings: 118,016 Hermes context / 262,144 model context · extracted native MTP max 3 · KV 8-bit / group 64

Independent combined suite: 86/94 passed. Fixed acceptance assertion failures: 8. Debug build: Pass.

Muse Glimmer 30B

6-bit MLX · Not verified · 2026-08-28

No patch47m 27s captured span
Settings and evidence

Retained usage and session evidence show 16 model calls without a usable patch or final response. The later status audit confirmed a clean benchmark worktree. DFlash was not exercised. Full independent Xcode results were not recovered.

Runtime: MLX / Hermes

Settings: Base MLX path; no DFlash drafter. Exact effective context and rendered reasoning tier not verified in retained summary.

Independent combined suite: Not recovered. Fixed acceptance assertion failures: Not recovered. Debug build: Not recorded.

Qwen3.8 + DFlash2 · 118K

6-bit MLX · medium · 2026-08-30

Pass27m 15s
Settings and evidence

Passed with 24 tool calls, no exact repeated calls, and no recorded context compaction. 48.3% less wall time than the historical Q6 control. About 12.88 GiB observed swap and seven memory-guard cache sheds remained. DFlash draft acceptance: 56.50%. This end-to-end speed difference is not an isolated on/off measurement of speculation.

Runtime: mlx-dspark / DFlash2

Settings: 118,016 context · DreamFoundries/Qwen3.8-27B-6bit + incoai/Qwen3.8-27B-DFlash2 · medium

Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Qwen3.8 GSQ-RCO 27B

IQ3_S GGUF · medium · 2026-09-13

Pass19m 44s
Settings and evidence

Correct two-line parser change and four authored tests. Recovered from an obsolete Xcode path and a faulty URL assertion; final independent suite passed. Test coverage has minor gaps described in the run report. Previously fastest passing local configuration; Signal later completed this task in 15m56s.

Runtime: LM Studio / llama.cpp 2.37.0

Settings: 118,016 effective context · one slot · native MTP, max 3 draft tokens

Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Signal 3.8 27B

AP-IQ3_S GGUF · medium · 2026-09-16

Pass15m 56s
Settings and evidence

Earlier passing local observation: 15m56s, 19.3% less time and 47.7% fewer output tokens than GSQ-RCO. More tool calls (36 vs 22), so speed did not mean fewer actions. Four authored tests; recovered from an invalid Xcode selector and two URL-encoding assertion failures. Explicit Unicode inspection and exact Markdown links were strengths. Used file search instead of graph; wrote a /tmp scratch file despite the scope constraint; edited production before tests. Independent 91/91 and build passed without repair. Different endpoint wrapper, q8 KV and MTP max 4 mean this is a complete-configuration comparison, not an isolated fine-tune effect.

Runtime: Standalone llama.cpp 2.37.0 / Hermes

Settings: 118,016 context · q8_0 K/V · native MTP max 4 · backend draft sampling off · vision projector loaded · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 · Xcode 27 beta 6

Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Signal 3.8 27B · Q6

AP-Q6_K GGUF · medium · 2026-09-17

Pass11m 21s
Settings and evidence

Fastest passing local observation: 11m21s, 28.7% less time than Signal IQ3_S and 42.5% less than GSQ-RCO. 24 tool calls versus 36 and 22 respectively. Simpler two-line production patch; four authored tests with all-dash full Markdown links and direct raw-variant URL round-trips. Recovered from four URL-encoding assertion failures; corrected own 89-test suite passed, then independent 91/91 and build passed without repair. File search instead of graph, production-before-test ordering, /tmp diagnostics and piped exit-code handling remain workflow limitations. Unrelated TypeScript check lacked tsc and was disclosed. One run per configuration; cache, machine state and work trajectory can affect timing.

Runtime: Standalone llama.cpp 2.37.0 / Hermes

Settings: 118,016 context · q8_0 K/V · native MTP max 4 · backend draft sampling off · BF16 vision projector loaded · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 · same harness/runtime settings as Signal IQ3_S

Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Qwopus3.8 27B Flash V2 · Q6

MLX affine 6bit · group 64 · medium · 2026-09-22

Pass17m 03s
Settings and evidence

Passed in 17m03s: 91 independent tests and Debug build, without external patch repair. Two production lines and four authored tests. 16 tools and 5,619 output tokens, fewer than Signal Q6, but 50.2% more elapsed time. Used file search instead of graph and production-before-test ordering; recovered unaided from stale Xcode path and test-without-building before a build. Authored Markdown and URL checks are narrower than Signal Q6; fixed independent acceptance tests passed. Runtime, KV representation, no MTP and released Xcode differ from Signal. Single observation, not an isolated model-quality comparison. Vision passed a basic image smoke test only.

Runtime: mlx-vlm 0.6.15 / Hermes

Settings: 118,016 context · MLX 8-bit KV group 64 from token 0 · prompt cache on and observed · no MTP drafter · vision included and smoke-tested · temperature .6 / top-p .95 / top-k 20 / min-p .05 / presence penalty 1.0 (64-token window) · verified sampler correction · Xcode 27.0 released

Independent combined suite: 91/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

ThinkingCap Qwen3.8 27B

Q6_K GGUF · medium · 2026-09-24

Agent failure60m 15s to termination
Settings and evidence

Incomplete: stopped by the controller after 60m 15s with no final agent response. Independent evaluation of the unmodified patch passed both fixed acceptance methods (eight dash/spacing cases) and the Debug build. The combined suite passed 90 of 91 tests; the sole failure was a model-authored Markdown test incorrectly expecting colon percent-encoding (%3A). The one-hour cutoff was chosen around minute 48, not preregistered or uniformly applied to earlier runs. Elapsed time at termination is not a successful completion time.

Runtime: LM Studio bundled llama.cpp 2.37.0, commit 8172e65; standalone server

Settings: Hermes 3aee290899e478c5fdfb6a241ef62758a49829b3; 118,016 context; q8_0 K/V; native MTP max 4; backend draft sampling disabled; temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0. Released Xcode 27; iPhone 17 / iOS 26.5. Same historical baseline, prompt and fixed evaluator. Vision projector loaded and image smoke test passed; coding benchmark was text-only

Independent combined suite: 90/91 passed. Fixed acceptance assertion failures: 0. Debug build: Pass.

Ornith 1.5 35B-A3B

Not recovered · Not recovered · Date not recovered

Reported failNot recovered
Settings and evidence

The user-provided historical comparison lists Ornith 1.5 35B-A3B as failed. Detailed benchmark logs, timing, quantization, context, and test counts were not recovered. This is a reported result, not an independently revalidated measurement. It is not conflated with a separately found Ornith 9B download.

Runtime: Not recovered

Settings: Configuration named in the supplied screenshot; detailed runtime settings unavailable

Independent combined suite: Not recovered. Fixed acceptance assertion failures: Not recovered. Debug build: Not recorded.

The 14 include the historical user-reported Ornith result, explicitly unverified. Gemini (14m07s) and Luna/max (6m50s) passed as hosted references outside this local count. The separate DFlash Sharp diagnostic failed the agent task despite external test passes. Bonsai was cancelled, incomplete, and unscored.

Signal Q6 recorded settings

Weights
Signal 3.8 27B AP-Q6_K GGUF
Runtime
Standalone llama.cpp 2.37.0 / Hermes
Context capacity
118,016 tokens
KV precision
q8_0 K/V
Native MTP
Max 4; backend draft sampling off
Sampler
Temperature .6; top-p .95; top-k 20; min-p .05; presence penalty 1.0

These are observed run settings, not a tested installation recipe. Weight quantization, KV precision, context capacity, prompt-prefix reuse, and MTP are separate controls. No isolated MTP speedup was measured.

Take the findings with you

Curated data and settings, with checksums. No private app source, patches, raw logs, credentials, or model weights are included.

How far these results travel

One bounded parser task on Jason’s MacBook Pro M5 Pro with 48 GB memory. This is not a general model ranking. The combined test counts include baseline tests, model-authored tests, and two fixed acceptance methods; 91/91 does not mean 91 independent coding tasks.

Agent wall time excludes model loading, warm-up, baseline preparation, and independent evaluation. Runtime, harness, reasoning tier, cache state, Xcode version, and work trajectory differ. A single observation per setup does not isolate quantization, fine-tuning, or context size as a cause.

ThinkingCap was stopped after 60m15s without a final answer. Both fixed acceptance methods and the build passed, but one authored test failed. Its termination time is not a completion time; the cutoff was chosen during the run, not preregistered.

Supplementary DFlash 65K and 118K runs both passed. The 65K run reported 95% context occupancy; the 118K run recorded zero compactions. Both showed memory pressure and cache trimming. This supports a capacity-headroom observation, not a correctness advantage.

Evidence identity

The curated export derives from the immutable September 24 film ledger. All 66 indexed source-file hashes matched the retained archive when prepared. Raw receipts remain private.

Ledger SHA-256: b1fc4463de731a470b42f189f9b51863c036ac38a34f43c57d954e920c2dfded

Fixed evaluator SHA-256: f8c54b591bed62ee5fd7eec5e81e7c9e88b6481511b726369f9f31e768fc7f64