AI Efficiency Toolbox
← LLM Findings

Setup guide · version 2026-10-01

Terse-Coder 27B Q6_K · native MTP

A reproducible candidate and cautious setup path. This guide does not establish that it will succeed on every device.

Before you download

  1. Inspect existing hardware, free RAM/VRAM, disk, models and runtime. Confirm the selected hardware and memory; never assume separate GPUs form one memory pool.
  2. Check the model license, exact pinned GGUF filename, revision and checksum below. Check the report's historical-byte verification notes. Keep existing models and configs.
  3. Use a current compatible llama.cpp build, or LM Studio with an engine supporting this architecture and template. Follow its official model-download guide; do not silently substitute another quant or latest revision.
  4. Begin with a modest context and one session. Inspect real allocation before raising limits. Weights, cache, buffers and other apps all need memory. Recorded long-context/MTP settings are not guaranteed on other versions.
  5. Keep any local endpoint bound to loopback. Consult your local agent’s provider documentation. Pi/Hermes need compatible provider and tool parsing; chat assistants may only be able to explain manual steps.
  6. Run a bounded scratch task with measurable acceptance checks. Record quant/hash/runtime/template/context, elapsed time, memory, corrections and failures. Stop on memory pressure or repeated malformed tool calls. Installation and configuration changes follow your agent’s own approvals.
Settings, artifact identity & publisher claim

Medium reasoning; context 118016; q8_0 K/V; native draft-mtp max 4, backend draft sampling disabled; temperature 0.6, top-p 0.95, top-k 20, min-p 0.05, presence penalty 1.0; response cap 8192; parallel 1; no vision projector. Hermes 3aee290899e478c5fdfb6a241ef62758a49829b3. Author template Shockem/froggeric-terse-coder at c557b014b6675a87683b45d0800a68842311ffe2 (rendered anti-rumination/medium-effort use verified). Qualify MTP/tool parsing on the exact runner before relying on these settings.

Settings source · Apache-2.0

Qwen3.8-27b-Terse-Coder.Q6_K.gguf
Revision: 7602eadf6a90853492e4e8a377c2a2f531bf21e2
SHA-256: 3164a8d8cc7f5c669b7a6de57ce9e5db126d73b8d39a0f424a2b9c649f25df07

Actual local model bytes rehashed and size checked against pinned download: SHA256 3164a8d8cc7f5c669b7a6de57ce9e5db126d73b8d39a0f424a2b9c649f25df07. Author template separately pinned; runner path/alias and rendered template proof verified. This exact first-party run does not establish equivalence with other quants/runtimes.

Publisher/author claim: Shockem identifies this pinned artifact as Qwen3.8-27b-Terse-Coder Q6_K and licenses the repository Apache-2.0. No publisher speed/quality claim is used for ordering. Source

Claim is attributed, not independently confirmed; it does not determine ordering. No comparable vendor/community thinking-reduction percentage is available.

27B total parameters · dense. Full stored weights determine memory.

AI Efficiency Toolbox · 2026-10-02 · Personally tested lead; community coverage absent. One independently checked first-party FaithVoice parser run passed on Apple M5 Pro 48GiB. Exact bytes/settings are documented; no cross-variant quality transfer, isolated MTP speedup or release qualification established. Keep this separate from the historical 14-setup study and community observations.

  • 2026-10-02: add verified latest native-MTP run; earlier no-MTP pass retained as context, not another independent community report.
Copy an agent setup handoff

Includes your selected hardware/task and the exact artifact. Your agent must check resources and its own permissions; without computer tools it can guide manual setup.

Evidence and limits

One independently checked first-party FaithVoice parser run passed on Apple M5 Pro 48GiB. Exact bytes/settings are documented; no cross-variant quality transfer, isolated MTP speedup or release qualification established. Keep this separate from the historical 14-setup study and community observations.

Actual local model bytes rehashed and size checked against pinned download: SHA256 3164a8d8cc7f5c669b7a6de57ce9e5db126d73b8d39a0f424a2b9c649f25df07. Author template separately pinned; runner path/alias and rendered template proof verified. This exact first-party run does not establish equivalence with other quants/runtimes.

Publisher claims and community reports are attributed evidence, not independent verification of your installation. No automatic install or device access is performed by this website.