Edition 2026-10-07 latest · digest built 2026-10-07T12:04:36+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud
MoE Throughput Playbook: Expand Experts, Prefetch Ahead, Retrain the Drafter
Today's strongest signals are about squeezing more out of local models: a routing patch that lifts HumanEval on a 2080 Ti, an expert-prefetch engine that recovers decode headroom on 2×3090, and a retrained speculative drafter that doubles throughput on Ternary Bonsai. Alongside that, agent builders get a deterministic tool-execution gate, a memory-file audit that found half the rules unenforced, and a transcript-only judge that passed 8 broken runs. Claude Code users also get two concrete housekeeping fixes, and a PoPETs study flags what claude.ai sends to third parties.
Local Inference: More Experts, Fewer Misses
The local-model crowd had a productive day. A controlled A/B on Qwen3.6-35B-A3B shows that activating 20 experts instead of 8 on the last 15 layers buys +1.3 points of HumanEval for about 19% slower decode — a tunable knob, not a retrain. Strata attacks the other side of MoE inference, prefetching experts before a miss and recovering up to +34% decode headroom on 2×3090. A retrained DFlash 2 drafter for Ternary Bonsai 2 27B delivers 2.2× on an L4 and 3.2× on code edits with ngram lookup. And two cheap wins: Ramjet offers Dynamo-style multi-GPU serving without Kubernetes, while simply moving your monitor cable to the iGPU frees ~2.5GB of VRAM.
Agent Discipline: Gates, Audits, and Honest Evals
The agent-safety thread is converging on the same idea: the model proposes, code decides. CLIM Agent Guard checks the final tool payload against guard-owned authoritative state immediately before execution, so a prompt-injected delete never reaches the filesystem. A memory-file audit found 71 of 147 agent rules were cited by nothing in code, tests, or hooks — unenforced dead weight. And a benchmark of Claude Code's /goal found a transcript-only LLM judge said "done" in 17 of 17 runs, 8 of which were actually broken. On the RAG side, GateKeep RAG enforces tenant and clearance checks before retrieval, so restricted chunks never enter the context window.
Claude Code Housekeeping and a Privacy Heads-Up
Two small Claude Code fixes with outsized payoff: Claude Code silently ignores AGENTS.md when a CLAUDE.md exists (one `@AGENTS.md` line fixes it), and a five-file-plus-hook memory layout beats relying on /compact. Finally, an IMDEA Networks study accepted at PoPETs 2027 reports that claude.ai sends chat IDs, chat links, user IDs, and email addresses to third parties like Datadog and Intercom, partly even after rejecting non-essential cookies — worth knowing before pasting customer data into a web chat.
Today's findings
-
#1 MoE Expansion Routing: +1.3pt HumanEval for 19% Decode Costtechnique
Activating 20 experts instead of 8 on the last 15 layers of Qwen3.6-35B-A3B lifted HumanEval from 89.6% to 90.9% at 4-bit on a single RTX 2080 Ti, for roughly 19% slower decode.
MoE expansion routingOne knob: +1.3pt accuracy for 19% decodevsStock — 8 expertsExpanded — 20 expertsHumanEval89.6%90.9%Decode speedBaseline~19% slowerRetrainingNoneNoneHardwareOne 2080 TiOne 2080 TiStock — 8 experts wins the row Expanded — 20 experts wins the rowQwen3.6-35B-A3B, 4-bit: swap 8→20 experts on the last 15 layers, nothing else changes.Why it matters: Expert routing is a tunable knob on consumer MoE models — you can trade a little speed for measurable coding accuracy without retraining or new hardware.
How to apply: If you run a MoE model in llama.cpp or vLLM, try bumping experts-per-token on the last N layers and A/B it on your own eval set; measure the decode-speed cost before committing.
moelocal-llmquantizationinference
Read more: MoE expansion , First HumanEval number on coding for Qwen3.6-35B-A3B — 90.9% at 4-bit on a RTX 2080 Ti, and an A/B of my "MoE expansion" routing patch vs stock 89.6% · MoE expansion , First HumanEval number on coding for Qwen3.6-35B-A3B — 90.9% at 4-bit on a RTX 2080 Ti, and an A/B of my "MoE expansion" routing patch vs stock 89.6%
-
#2 Strata Prefetches MoE Experts to Recover Up to 34% Decode Headroomrepo
Strata predicts which experts a token will need and fetches them before the miss, recovering up to +34% decode throughput on 2×3090 with a large MoE.
STRATA · MoE PREFETCHINGPrefetch the experts before the token asks+34%decode throughput recovered (up to)vs. same rig without prefetch2×3090runs on consumer GPUs you already ownPre-misspredicts and fetches experts ahead of the stallA/Bshrink the expert cache to measure your miss cost, then re-measure tokens/sExpert cache misses, not raw compute, are the hidden tax on consumer MoE inference — hiding the stall is a pure softwareWhy it matters: Expert cache misses, not raw compute, are the hidden tax on consumer MoE inference; prefetching is a cheap software win on hardware you already own.
How to apply: Clone github.com/Niko1221/Strata, deliberately shrink your expert cache to measure your miss cost, then enable prefetch and re-measure tokens/s.
moelocal-llminferenceollama
-
#3 Retrained DFlash 2 Drafter Gives 2.2× Speedups on Ternary Bonsai 2 27Btechnique
A retrained speculative-decoding drafter for Ternary Bonsai 2 27B hits 2.2× on an L4, 3.2× on code edits with ngram lookup, 1.5× on a Mac, and 1.2× in Chrome.
measuredOne retrained drafter, four speedups1Code edits + ngram3.2×2L4 GPU2.2×3Mac laptop1.5×4Chrome tab1.2×Speedup vs. undrafted baselineTernary Bonsai 2 27BWhy it matters: Speculative decoding with a matched drafter is one of the few ways to get real speedups without touching model quality.
How to apply: If you serve a quantized 27B locally, check whether a matching drafter exists for your model family; retraining one on your own traffic is a weekend project with outsized payoff.
speculative-decodinglocal-llminferencequantization
-
#4 Deterministic Tool-Execution Boundary Blocks Bad Agent Actions Before They Landtechnique
CLIM Agent Guard sits between an agent's structured tool proposal and the side effect, checking the final payload against guard-owned authoritative state instead of asking another LLM to judge safety.
Agent guardrailsTool calls pass a deterministic gate, not an LLM judgebounded capabilityAgent tool executionscope Authorization checklimit State freshnessrevoke Reject on mismatchPrompt injection can talk a model into anything — it can't talk this gate into anything.Why it matters: Prompt injection can talk a model into anything; it can't talk a deterministic gate into anything. This is the architecture pattern for shipping agents that touch real systems.
How to apply: Wrap your LangGraph tool calls with a pre-execution check that validates authorization and state freshness against authoritative sources, and reject anything that doesn't match.
agentssecurityguardrailslanggraph
Read more: Bypass Challenge: Can prompt injection cross a deterministic tool-execution boundary? · The model proposes. Code decides.
-
#5 Claude Code Silently Ignores AGENTS.md When a CLAUDE.md Existstip
Per Anthropic's memory docs, Claude Code only reads AGENTS.md if there's no CLAUDE.md in the working directory or above it — add `@AGENTS.md` to the top of CLAUDE.md to keep shared conventions visible.
Claude Code memory docsDoes Claude Code actually read AGENTS.md?4 setups checkedAGENTS.md, no CLAUDE.mdAGENTS.md + CLAUDE.md presentTop of CLAUDE.md: @AGENTS.mdCLAUDE.md over ~200 linespass warn failA CLAUDE.md in the working dir or any parent silently shadows AGENTS.md — Cursor still reads it, Claude Code doesn't.Why it matters: Teams that keep shared agent conventions in AGENTS.md for Cursor and other tools are silently losing them in Claude Code.
How to apply: Add one line — `Shared conventions for all coding agents live in AGENTS.md: @AGENTS.md` — to the top of your CLAUDE.md, keep CLAUDE.md under ~200 lines, and split big changes into four gated stages.
claude-codeagentsworkflowmcp
-
#6 Five Plain Files and One Hook Beat /compact for Claude Code Memorytechnique
An index, per-fact notes with Why/How-to-apply, a dated diary, a waiting list, and a hook keep Claude Code oriented across compactions without relying on /compact.
Agent memory, no /compactFour plain files + one hookIndex ~90 lines, loads every sessionFact notes one per fact, with its whyDiary dated, written as it worksWaiting list parked, picked up laterHook injects the index at startalways in context auto-inject plain fileCompaction drops context silently — the hook reloads the index so the rest survives.Why it matters: Compaction silently drops context; a file-based memory that loads every session is more reliable than hoping the summarizer keeps what matters.
How to apply: Create a ~90-line index that loads each session, one file per fact with its reason, and a diary written while the agent works — then wire a hook so the index is always injected.
claude-codememoryagentsworkflow
-
#7 Audit Your Agent's Memory File: 71 of 147 Rules Were Cited by Nothingtechnique
A script that checks whether any code, test, hook, or prompt references each rule in an agent's memory file found 71 of 147 rules were unenforced dead weight.
AGENT MEMORY AUDIT48 of every 100 rules enforce nothing48%of rules were cited by nothing71 of 147 real rules had zero references in code, tests, hooks, or prompts.Why it matters: Uncited rules are unenforced rules — they bloat context and give a false sense of governance.
How to apply: Grep your memory/instructions folder for each rule's name across code, tests, hooks, and other prompts; delete or wire up anything nothing references.
agentsmemoryprompt-engineeringworkflow
-
#8 A Transcript-Only LLM Judge Said 'Done' in 17/17 Runs — 8 Were Brokentip
Benchmarking Claude Code's /goal showed that a judge reading only the transcript approved every run, including 8 that were actually broken.
Why it matters: If your eval only reads the conversation, it's grading narration, not outcomes — and it will pass broken work.
How to apply: Give your judge access to the actual artifacts (files, test output, diffs) rather than the transcript, and include a few known-broken runs to calibrate it.
evalsagentsclaude-codereliability
-
#9 GateKeep RAG Enforces Permissions Before the LLM Ever Sees a Chunkrepo
GateKeep RAG attaches tenant, role, and clearance metadata to every chunk and filters before retrieval, so restricted documents never enter the context window.
RAG ACCESS CONTROLPrompt guardrail vs retrieval gatePrompt guardrail- “Don’t reveal other tenants”
- Relies on the model obeying
- Restricted chunk already sits in context
Retrieval gate- Chunks tagged: tenant · role · clearance
- Filtered before the model sees anything
- Restricted docs never reach the prompt
Once a restricted chunk is in context, you’ve already lost.GateKeep RAG puts the access check in the retrieval layer and treats the LLM as untrusted for anything it can read.Why it matters: Prompt instructions like 'don't reveal other tenants' data' aren't a security boundary — once a restricted chunk is in context, you've already lost.
How to apply: Move access checks into the retrieval layer: tag chunks with tenant/role/clearance, filter at query time, and treat the LLM as untrusted for anything it can read.
ragsecuritymulti-tenantretrieval
-
#10 Ramjet Brings Dynamo-Style Multi-GPU Serving to Local Setups Without Kubernetesrepo
Ramjet is an open-source, local alternative to NVIDIA Dynamo for multi-GPU inference that aims to match or beat it without the k8s overhead.
Why it matters: Multi-GPU local serving is usually a Kubernetes-shaped problem; a lightweight drop-in makes DGX Spark and multi-Mac rigs practical.
How to apply: Check github.com/helixml/ramjet and the Ramjet-vs-Dynamo writeup, then contribute a recipe for your hardware so others can pull it.
local-llminferencemulti-gpuollama
Read more: Ramjet - mini altermative to nvidia dynamo · Ramjet - mini altermative to nvidia dynamo
-
#11 Plug Your Monitor Into the Motherboard to Reclaim ~2.5GB of VRAMtip
Moving the display cable from the GPU to the iGPU freed ~2.5GB of VRAM, taking one user from a 65k to a 132k context window at 125 tok/s on a 4090.
Why it matters: Free VRAM is free context and free model headroom — no new hardware required.
How to apply: If your desktop has integrated graphics, plug the monitor into the motherboard HDMI/DisplayPort and re-check your context window and tokens/s.
local-llmvramhardwareollama
Read more: PSA: Use your iGPU for display to save VRAM
-
#12 PoPETs Study Finds Claude Web Sends Chat IDs and Emails to Third Partiespaper
An IMDEA Networks study accepted at PoPETs 2027 found claude.ai sending chat IDs, chat links, user IDs, and email addresses to third parties like Datadog and Intercom, partly even after rejecting non-essential cookies.
Why it matters: Teams treating Claude as a private workspace should know what leaves the browser before they paste customer data into it.
How to apply: Review your cookie and consent settings, avoid pasting sensitive identifiers into web chats, and prefer API or local paths for regulated data.
privacyclaudesecuritycompliance