Edition 2026-10-11 latest · digest built 2026-10-11T12:10:12+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud

Probabilistic MTP Lands in llama.cpp, Plus a Parallel-Slot Re-Prefill Trap and a 322M Tool-Call Guard

Today's strongest signals are all about making local and Claude-centric stacks measurably better: llama.cpp merged probabilistic MTP for a free decode speedup, a 322M local model now vets every agent tool call in ~90ms, and a hands-on bench shows standard evals can't distinguish your quants while agentic tasks can. Alongside those, a local 9B beat 14 hosted decision models on triage, GLM-5.3-Flash runs at 61 tok/s on a single Mac Studio, and two open decision-model releases ship with GGUF and MLX builds. There's also a sycophancy paper worth folding into your evals and a quiet Anthropic credit program worth checking.

Local Inference Gets Faster and Cheaper

The most immediately usable item is llama.cpp's merged probabilistic MTP path, which lifts decode throughput on prose by roughly 14% with no hardware change. A companion tip from the same ecosystem: if you serve a coding agent through llama-server with --parallel 2, a concurrent background summarization request can silently re-prefill a 119K-token prompt, turning a 2-second cache hit into a 207-second stall. On the model side, GLM-5.3-Flash is now demonstrably runnable on a single high-RAM Mac, with a 320B GGUF at ~61 tok/s and 262K context on an M5 Ultra, and a 4-bit MLX build on an M3 Ultra.

Guardrails and Evals Are Where the Real Work Is

laya-guard is the standout safety tool: a ~322M local model that judges every tool call a coding agent makes in about 90ms on CPU, with no API dependency. On the measurement side, a careful bench shows five Qwen3.8-27B quants scoring within ±2 points on GSM8K, MMLU-Pro and IFEval while diverging sharply on hidden-test agentic coding tasks — a strong argument for building your own private bench before trusting a quant. A separate test found a local Qwen3.5-9B matching or beating 14 hosted closed-set 'decision models' on 56 real triage tasks, and a new sycophancy paper shows LLMs shift diagnoses under user pushback, which is a failure mode worth adding to your eval suite.

Models, Training Recipes, and Tooling

Two open decision-model releases landed: Jebadiah v2.1 (27B and 9B) returns a probability per allowed label straight from the logits, with GGUF and MLX quants that mostly preserve the full-weight pick, and the broader decision-model comparison above suggests this interface is worth trying before paying for a routing API. For training, there's a dense-to-sparse MoE upcycling write-up done on an 8GB RTX 4060, and a self-hosted SFT/RL distillation pipeline that trains a Qwen3.8 on reasoning traces from a larger teacher in about 12 hours. On tooling, tensorViz now inspects Hugging Face checkpoints and captures a forward pass to show the real model graph, and Anthropic's Team-plan linkage reportedly unlocks $500/month in API credits — worth checking if your team already pays for Claude.

Today's findings

  1. #1 llama.cpp Merges Probabilistic MTP for ~14% Faster Decodetool

    llama.cpp's probabilistic multi-token prediction path is now merged, giving roughly +14% decode throughput on prose with tuned draft settings.

    llama.cpp · inference
    Probabilistic MTP is merged
    +14%
    decode throughput on prose with tuned draft settings
    free for models you already serve
    0
    new hardware, requantization, or API change needed
    PR #27694
    any build past this merge ships the MTP path
    draft-n-max
    and draft-p-min — tune against your prose workload
    Update, enable the MTP head, tune the draft knobs — a plain speedup for every local model.

    Why it matters: It's a free speedup for every local model you already serve — no new hardware, no requantization, no API change.

    How to apply: Update llama.cpp to a build past PR #27694, enable the MTP head, and tune draft-n-max / draft-p-min against your own prose workload before rolling it out.

    llama.cppinferencespeculative-decodinglocal-llm

    Read more: Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

  2. #2 llama-server --parallel 2 Silently Re-prefills 119K-Token Promptstip

    With --parallel 2 and --cache-idle-slots, a concurrent summarization request lands on an empty slot and re-prefills the entire prompt — ~207s instead of ~2s.

    Why it matters: If you serve a coding agent through llama-server, background requests can tank latency by 100x with no error surfaced anywhere.

    How to apply: Set --parallel 1 until slot KV snapshotting lands, or route background summarization to a separate server instance so it can't evict the main chat's cache.

    llama.cpplocal-llmlatencyserving

    Read more: VS Code Copilot + llama-server --parallel 2: concurrent requests cause full re-prefill

  3. #3 laya-guard: A 322M Local Model That Vets Every Agent Tool Calltool

    laya-guard runs a ~322M local model that judges each shell or tool call your coding agent makes in about 90ms on CPU, with no API calls.

    Agent guardrail
    A 322M local gate on every agent tool call
    1
    Agent call
    any shell or tool run
    2
    Vet
    ~90ms on CPU · no API
    the local judge — bad calls die before execution
    3
    Run or block
    only clean calls execute
    Point it at the tool-call stream; run log-only for a week, then enforce.

    Why it matters: It gives you a cheap, offline policy layer between an autonomous coding agent and your machine, catching bad calls before they execute.

    How to apply: Drop the repo into your agent harness, point it at the tool-call stream, and run it in log-only mode for a week before enforcing blocks.

    agentsguardrailslocal-llmsecurity

    Read more: laya-guard: a local guardrail that judges every tool call your coding agent makes · I use a ~322M local model to judge every tool call my coding agents make (about 90 ms on CPU, no API)

  4. #4 Standard Benchmarks Can't Tell Your Quants Apart — Agentic Tasks Cantechnique

    Five Qwen3.8-27B quants from Q4 to Q8 score within ±2 points on GSM8K, MMLU-Pro and IFEval, but diverge sharply on hidden-test agentic coding tasks.

    Why it matters: Quant choice looks free on public leaderboards and is anything but free on real agentic work, where the gap can be 1/12 vs 12/12 on hard tasks.

    How to apply: Build a small private bench from your own repos with injected bugs and hidden tests, then pick quants on that instead of public scores.

    quantizationevalslocal-llmagents

    Read more: Standard evals can't tell my local quants apart. My own agentic bench can (Qwen3.8 27B / Flash-Next, Q3 to Q8)

  5. #5 A Local Qwen3.5-9B Beat 14 Hosted 'Decision Models' on Triagetechnique

    Across 56 real triage and routing tasks, a local Qwen3.5-9B matched or beat every paid closed-set decision model tested.

    Why it matters: Routing, urgency scoring, and classification may not need a specialized paid endpoint at all — a local 9B you already run can do it.

    How to apply: Before buying a decision-model API, assemble 50-100 of your own tasks, constrain the local model's output to the allowed options, and compare picks.

    local-llmclassificationroutingevals

    Read more: I tested 14 "decision models" against local LLMs on 56 real triage tasks. Qwen3.5-9B beat every one of them. · I tested 14 "decision models" against local LLMs on 56 real triage tasks. Qwen3.5-9B beat every one of them.

  6. #6 GLM-5.3-Flash Runs Locally: 61 tok/s at 262K Context on M5 Ultratool

    GLM-5.3-Flash (320B) serves at ~61 tok/s with 262K context on an M5 Ultra Mac Studio via llama.cpp, and a 4-bit MLX build runs on an M3 Ultra.

    Why it matters: A frontier-adjacent open model now fits on a single high-RAM Mac with usable long-context throughput, no cluster required.

    How to apply: Try the GGUF or MLX builds with MTP enabled, and budget roughly 172-181GB resident memory for the 4-bit MLX variant.

    local-llmglmmlxllama.cpp

    Read more: GLM-5.3-Flash (320B) on an M5 Ultra Mac Studio with llama.cpp: 61 tok/s, 262K context · GLM-5.3-Flash abliterated MLX 4-bit on mlx-serve, M3 Ultra 256 GB

  7. #7 Jebadiah v2.1: Open Decision Models That Return a Probability per Labelrepo

    The 27B and 9B Jeb models answer typed multiple-choice questions by reading probabilities straight from the label logits — no generation, no parsing.

    Jebadiah v2.1 · open decision models
    Don't generate an answer, don't parse a string — read the probabilities
    Generate & parse
    • Model writes the answer in text
    • Extra parse step maps output to a choice
    • Sampling adds nondeterminism
    Read label logits
    • Probability per option, straight from the logits
    • No parsing — argmax is the answer
    • Deterministic, one forward pass
    Jebadiah v2.1 (27B & 9B) answers typed multiple-choice off the logits; GGUF Q5/Q8 and MLX quants mostly match the full-w
    For in-app decisions: feed JSON context plus a fixed option set.

    Why it matters: It gives you a deterministic, parse-free interface for in-app decisions, with GGUF and MLX quants that mostly preserve the full-weight pick.

    How to apply: Pull the GGUF Q5/Q8 or MLX build, feed JSON context plus a fixed option set, and read the per-option probabilities directly from the logits.

    local-llmggufclassificationstructured-output

    Read more: ebadiah v2.1 (27B and 9B): local decision models that return a probability per answer, with GGUF and MLX builds · [P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]

  8. #8 Upcycling Dense Models into Sparse MoE on an 8GB GPUtechnique

    A write-up on converting existing dense checkpoints into sparse Mixture-of-Experts without pretraining from scratch, done on an RTX 4060.

    TECHNIQUE · UPCYCLING
    Dense checkpoint → sparse MoE, no from-scratch pretraining
    Dense checkpoint
    • Every weight fires per token
    • Qwen2.5-0.5B · SmolLM2-360M
    • Growth demands a training cluster
    Upcycled MoE
    • FFNs copied into experts
    • Router picks a few per token
    • Router-init recipe · 8GB RTX 4060
    Replicate the router-init recipe and measure quality loss on your own tasks before scaling up.

    Why it matters: It points at cheaper per-token inference for roughly the same knowledge, without needing a training cluster.

    How to apply: Start from the Qwen2.5-0.5B and SmolLM2-360M conversions, replicate the router-init recipe, and measure quality loss on your own tasks before scaling up.

    moefine-tuninglocal-llmquantization

    Read more: Converting dense models into Mixture-of-Experts

  9. #9 DIY Pipeline Distills Reasoning Traces into a Local Qwen3.8technique

    A self-hosted SFT/RL pipeline that abliterates a Qwen3.8 model and trains it on reasoning traces harvested from a larger teacher, in about 12 hours.

    Why it matters: It shows a realistic path to a domain-tuned local reasoner without a lab budget or a managed fine-tuning service.

    How to apply: Collect traces from your own agent sessions, follow the compute-shader setup in the write-up, and run the SFT pass on your own hardware.

    fine-tuningdistillationlocal-llmreasoning

    Read more: DIY opencode distillation and qwen 3.8 reasoning traces (SFT/RL) training pipeline

  10. #10 Anthropic Team Plans Unlock $500/Month in API Creditstip

    Linking a Claude Team plan to a Console org reportedly grants $500 in monthly API credits, with startup programs adding up to $1,000 in credits and Team discounts.

    Why it matters: For a small team already paying for Claude, this is effectively free API budget for agents, evals, and batch jobs.

    How to apply: Check the official docs, link your Team plan to your Console org, and route non-interactive workloads through the credited API key.

    anthropicclaudepricingteams

    Read more: $500 in free monthly API credits for connecting your Team plan · Anthropic just gave us $1,000 in API credits and up to $7,500 in Claude Team Discounts!

  11. #11 tensorViz Adds Hugging Face Checkpoint Inspection to VS Codetool

    A free VS Code extension now reads safetensors metadata and captures a forward pass to show the actual model graph, not just weight shapes.

    Why it matters: Debugging a fine-tune or a custom head is much faster when you can see intermediate activation shapes instead of guessing from tensor names.

    How to apply: Install tensorViz, open a Hugging Face checkpoint, run a tiny example input, and walk the decoder layer graph to trace where shapes diverge.

    toolingpytorchhuggingfacedebugging

    Read more: Inspecting a Hugging Face checkpoint, then following an example input through its graph · Built a VS Code extension to make PyTorch models easier to follow

  12. #12 Sycophancy Paper: Agreeable LLMs Are Diagnostically Unstablepaper

    A new paper finds LLMs shift their diagnoses when users push back, and that this sycophancy tracks with diagnostic instability.

    Why it matters: Any decision-support feature that lets users argue with the model inherits this failure mode, and it won't show up in a static eval.

    How to apply: Add adversarial pushback turns to your evals and log whether the model's answer changes when the user asserts a wrong conclusion.

    evalssycophancyreliabilitypaper

    Read more: Too agreeable to be accurate? Sycophancy and diagnostic instability of large language models in medical diagnosis

Looking for topic trends and crawl volume over time? See Trends.