Edition 2026-10-07 latest · digest built 2026-10-07T12:04:36+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud

MoE Throughput Playbook: Expand Experts, Prefetch Ahead, Retrain the Drafter

Today's strongest signals are about squeezing more out of local models: a routing patch that lifts HumanEval on a 2080 Ti, an expert-prefetch engine that recovers decode headroom on 2×3090, and a retrained speculative drafter that doubles throughput on Ternary Bonsai. Alongside that, agent builders get a deterministic tool-execution gate, a memory-file audit that found half the rules unenforced, and a transcript-only judge that passed 8 broken runs. Claude Code users also get two concrete housekeeping fixes, and a PoPETs study flags what claude.ai sends to third parties.

Local Inference: More Experts, Fewer Misses

The local-model crowd had a productive day. A controlled A/B on Qwen3.6-35B-A3B shows that activating 20 experts instead of 8 on the last 15 layers buys +1.3 points of HumanEval for about 19% slower decode — a tunable knob, not a retrain. Strata attacks the other side of MoE inference, prefetching experts before a miss and recovering up to +34% decode headroom on 2×3090. A retrained DFlash 2 drafter for Ternary Bonsai 2 27B delivers 2.2× on an L4 and 3.2× on code edits with ngram lookup. And two cheap wins: Ramjet offers Dynamo-style multi-GPU serving without Kubernetes, while simply moving your monitor cable to the iGPU frees ~2.5GB of VRAM.

Agent Discipline: Gates, Audits, and Honest Evals

The agent-safety thread is converging on the same idea: the model proposes, code decides. CLIM Agent Guard checks the final tool payload against guard-owned authoritative state immediately before execution, so a prompt-injected delete never reaches the filesystem. A memory-file audit found 71 of 147 agent rules were cited by nothing in code, tests, or hooks — unenforced dead weight. And a benchmark of Claude Code's /goal found a transcript-only LLM judge said "done" in 17 of 17 runs, 8 of which were actually broken. On the RAG side, GateKeep RAG enforces tenant and clearance checks before retrieval, so restricted chunks never enter the context window.

Claude Code Housekeeping and a Privacy Heads-Up

Two small Claude Code fixes with outsized payoff: Claude Code silently ignores AGENTS.md when a CLAUDE.md exists (one `@AGENTS.md` line fixes it), and a five-file-plus-hook memory layout beats relying on /compact. Finally, an IMDEA Networks study accepted at PoPETs 2027 reports that claude.ai sends chat IDs, chat links, user IDs, and email addresses to third parties like Datadog and Intercom, partly even after rejecting non-essential cookies — worth knowing before pasting customer data into a web chat.

Today's findings

  1. #1 MoE Expansion Routing: +1.3pt HumanEval for 19% Decode Costtechnique

    Activating 20 experts instead of 8 on the last 15 layers of Qwen3.6-35B-A3B lifted HumanEval from 89.6% to 90.9% at 4-bit on a single RTX 2080 Ti, for roughly 19% slower decode.

    MoE expansion routing
    One knob: +1.3pt accuracy for 19% decode
    vs
    Stock — 8 experts
    Expanded — 20 experts
    HumanEval
    89.6%
    90.9%
    Decode speed
    Baseline
    ~19% slower
    Retraining
    None
    None
    Hardware
    One 2080 Ti
    One 2080 Ti
    Stock — 8 experts wins the row Expanded — 20 experts wins the row
    Qwen3.6-35B-A3B, 4-bit: swap 8→20 experts on the last 15 layers, nothing else changes.

    Why it matters: Expert routing is a tunable knob on consumer MoE models — you can trade a little speed for measurable coding accuracy without retraining or new hardware.

    How to apply: If you run a MoE model in llama.cpp or vLLM, try bumping experts-per-token on the last N layers and A/B it on your own eval set; measure the decode-speed cost before committing.

    moelocal-llmquantizationinference

    Read more: MoE expansion , First HumanEval number on coding for Qwen3.6-35B-A3B — 90.9% at 4-bit on a RTX 2080 Ti, and an A/B of my "MoE expansion" routing patch vs stock 89.6% · MoE expansion , First HumanEval number on coding for Qwen3.6-35B-A3B — 90.9% at 4-bit on a RTX 2080 Ti, and an A/B of my "MoE expansion" routing patch vs stock 89.6%

  2. #2 Strata Prefetches MoE Experts to Recover Up to 34% Decode Headroomrepo

    Strata predicts which experts a token will need and fetches them before the miss, recovering up to +34% decode throughput on 2×3090 with a large MoE.

    STRATA · MoE PREFETCHING
    Prefetch the experts before the token asks
    +34%
    decode throughput recovered (up to)
    vs. same rig without prefetch
    2×3090
    runs on consumer GPUs you already own
    Pre-miss
    predicts and fetches experts ahead of the stall
    A/B
    shrink the expert cache to measure your miss cost, then re-measure tokens/s
    Expert cache misses, not raw compute, are the hidden tax on consumer MoE inference — hiding the stall is a pure software

    Why it matters: Expert cache misses, not raw compute, are the hidden tax on consumer MoE inference; prefetching is a cheap software win on hardware you already own.

    How to apply: Clone github.com/Niko1221/Strata, deliberately shrink your expert cache to measure your miss cost, then enable prefetch and re-measure tokens/s.

    moelocal-llminferenceollama

    Read more: Strata: fetch experts before they're needed. Up to +34 % decode headroom measured on 2× 3090, more on smaller-VRAM cards

  3. #3 Retrained DFlash 2 Drafter Gives 2.2× Speedups on Ternary Bonsai 2 27Btechnique

    A retrained speculative-decoding drafter for Ternary Bonsai 2 27B hits 2.2× on an L4, 3.2× on code edits with ngram lookup, 1.5× on a Mac, and 1.2× in Chrome.

    measured
    One retrained drafter, four speedups
    1
    Code edits + ngram
    3.2×
    2
    L4 GPU
    2.2×
    3
    Mac laptop
    1.5×
    4
    Chrome tab
    1.2×
    Speedup vs. undrafted baselineTernary Bonsai 2 27B

    Why it matters: Speculative decoding with a matched drafter is one of the few ways to get real speedups without touching model quality.

    How to apply: If you serve a quantized 27B locally, check whether a matching drafter exists for your model family; retraining one on your own traffic is a weekend project with outsized payoff.

    speculative-decodinglocal-llminferencequantization

    Read more: I re-trained the DFlash 2 drafter for Ternary Bonsai 2 27B: 2.2x on an L4 (3.2x on code edits with ngram lookup), 1.5x on a Mac, 1.2x in Chrome

  4. #4 Deterministic Tool-Execution Boundary Blocks Bad Agent Actions Before They Landtechnique

    CLIM Agent Guard sits between an agent's structured tool proposal and the side effect, checking the final payload against guard-owned authoritative state instead of asking another LLM to judge safety.

    Agent guardrails
    Tool calls pass a deterministic gate, not an LLM judge
    bounded capability
    Agent tool execution
    scope Authorization check
    limit State freshness
    revoke Reject on mismatch
    Prompt injection can talk a model into anything — it can't talk this gate into anything.

    Why it matters: Prompt injection can talk a model into anything; it can't talk a deterministic gate into anything. This is the architecture pattern for shipping agents that touch real systems.

    How to apply: Wrap your LangGraph tool calls with a pre-execution check that validates authorization and state freshness against authoritative sources, and reject anything that doesn't match.

    agentssecurityguardrailslanggraph

    Read more: Bypass Challenge: Can prompt injection cross a deterministic tool-execution boundary? · The model proposes. Code decides.

  5. #5 Claude Code Silently Ignores AGENTS.md When a CLAUDE.md Existstip

    Per Anthropic's memory docs, Claude Code only reads AGENTS.md if there's no CLAUDE.md in the working directory or above it — add `@AGENTS.md` to the top of CLAUDE.md to keep shared conventions visible.

    Claude Code memory docs
    Does Claude Code actually read AGENTS.md?
    4 setups checked
    AGENTS.md, no CLAUDE.md
    AGENTS.md + CLAUDE.md present
    Top of CLAUDE.md: @AGENTS.md
    CLAUDE.md over ~200 lines
    pass warn fail
    A CLAUDE.md in the working dir or any parent silently shadows AGENTS.md — Cursor still reads it, Claude Code doesn't.

    Why it matters: Teams that keep shared agent conventions in AGENTS.md for Cursor and other tools are silently losing them in Claude Code.

    How to apply: Add one line — `Shared conventions for all coding agents live in AGENTS.md: @AGENTS.md` — to the top of your CLAUDE.md, keep CLAUDE.md under ~200 lines, and split big changes into four gated stages.

    claude-codeagentsworkflowmcp

    Read more: Heads-up: Claude Code ignores your AGENTS.md if you also have a CLAUDE.md (plus the 4-stage workflow I use)

  6. #6 Five Plain Files and One Hook Beat /compact for Claude Code Memorytechnique

    An index, per-fact notes with Why/How-to-apply, a dated diary, a waiting list, and a hook keep Claude Code oriented across compactions without relying on /compact.

    Agent memory, no /compact
    Four plain files + one hook
    Index ~90 lines, loads every session
    Fact notes one per fact, with its why
    Diary dated, written as it works
    Waiting list parked, picked up later
    Hook injects the index at start
    always in context auto-inject plain file
    Compaction drops context silently — the hook reloads the index so the rest survives.

    Why it matters: Compaction silently drops context; a file-based memory that loads every session is more reliable than hoping the summarizer keeps what matters.

    How to apply: Create a ~90-line index that loads each session, one file per fact with its reason, and a diary written while the agent works — then wire a hook so the index is always injected.

    claude-codememoryagentsworkflow

    Read more: Claude Code forgets after /compact. I stopped relying on compaction — 5 plain files + 1 hook (free template inside)

  7. #7 Audit Your Agent's Memory File: 71 of 147 Rules Were Cited by Nothingtechnique

    A script that checks whether any code, test, hook, or prompt references each rule in an agent's memory file found 71 of 147 rules were unenforced dead weight.

    AGENT MEMORY AUDIT
    48 of every 100 rules enforce nothing
    48%
    of rules were cited by nothing
    71 of 147 real rules had zero references in code, tests, hooks, or prompts.

    Why it matters: Uncited rules are unenforced rules — they bloat context and give a false sense of governance.

    How to apply: Grep your memory/instructions folder for each rule's name across code, tests, hooks, and other prompts; delete or wire up anything nothing references.

    agentsmemoryprompt-engineeringworkflow

    Read more: 147 rules live in my agent's memory file. 71 of them are cited by nothing in code. how do you tell which instructions in an agent's context actually change behavior?

  8. #8 A Transcript-Only LLM Judge Said 'Done' in 17/17 Runs — 8 Were Brokentip

    Benchmarking Claude Code's /goal showed that a judge reading only the transcript approved every run, including 8 that were actually broken.

    Why it matters: If your eval only reads the conversation, it's grading narration, not outcomes — and it will pass broken work.

    How to apply: Give your judge access to the actual artifacts (files, test output, diffs) rather than the transcript, and include a few known-broken runs to calibrate it.

    evalsagentsclaude-codereliability

    Read more: An LLM judge that only reads the transcript said "done" in 17 of 17 runs. 8 were broken. Notes from benchmarking Claude Code's /goal

  9. #9 GateKeep RAG Enforces Permissions Before the LLM Ever Sees a Chunkrepo

    GateKeep RAG attaches tenant, role, and clearance metadata to every chunk and filters before retrieval, so restricted documents never enter the context window.

    RAG ACCESS CONTROL
    Prompt guardrail vs retrieval gate
    Prompt guardrail
    • “Don’t reveal other tenants”
    • Relies on the model obeying
    • Restricted chunk already sits in context
    Retrieval gate
    • Chunks tagged: tenant · role · clearance
    • Filtered before the model sees anything
    • Restricted docs never reach the prompt
    Once a restricted chunk is in context, you’ve already lost.
    GateKeep RAG puts the access check in the retrieval layer and treats the LLM as untrusted for anything it can read.

    Why it matters: Prompt instructions like 'don't reveal other tenants' data' aren't a security boundary — once a restricted chunk is in context, you've already lost.

    How to apply: Move access checks into the retrieval layer: tag chunks with tenant/role/clearance, filter at query time, and treat the LLM as untrusted for anything it can read.

    ragsecuritymulti-tenantretrieval

    Read more: I built a RAG service where the LLM never sees documents the user isn't allowed to read (permission-aware, multi-tenant)

  10. #10 Ramjet Brings Dynamo-Style Multi-GPU Serving to Local Setups Without Kubernetesrepo

    Ramjet is an open-source, local alternative to NVIDIA Dynamo for multi-GPU inference that aims to match or beat it without the k8s overhead.

    Why it matters: Multi-GPU local serving is usually a Kubernetes-shaped problem; a lightweight drop-in makes DGX Spark and multi-Mac rigs practical.

    How to apply: Check github.com/helixml/ramjet and the Ramjet-vs-Dynamo writeup, then contribute a recipe for your hardware so others can pull it.

    local-llminferencemulti-gpuollama

    Read more: Ramjet - mini altermative to nvidia dynamo · Ramjet - mini altermative to nvidia dynamo

  11. #11 Plug Your Monitor Into the Motherboard to Reclaim ~2.5GB of VRAMtip

    Moving the display cable from the GPU to the iGPU freed ~2.5GB of VRAM, taking one user from a 65k to a 132k context window at 125 tok/s on a 4090.

    Why it matters: Free VRAM is free context and free model headroom — no new hardware required.

    How to apply: If your desktop has integrated graphics, plug the monitor into the motherboard HDMI/DisplayPort and re-check your context window and tokens/s.

    local-llmvramhardwareollama

    Read more: PSA: Use your iGPU for display to save VRAM

  12. #12 PoPETs Study Finds Claude Web Sends Chat IDs and Emails to Third Partiespaper

    An IMDEA Networks study accepted at PoPETs 2027 found claude.ai sending chat IDs, chat links, user IDs, and email addresses to third parties like Datadog and Intercom, partly even after rejecting non-essential cookies.

    Why it matters: Teams treating Claude as a private workspace should know what leaves the browser before they paste customer data into it.

    How to apply: Review your cookie and consent settings, avoid pasting sensitive identifiers into web chats, and prefer API or local paths for regulated data.

    privacyclaudesecuritycompliance

    Read more: I trust Anthropic with my data. I didn't agree to share it with Meta, TikTok and Google. (IMDEA study)

Looking for topic trends and crawl volume over time? See Trends.