Edition 2026-10-03 latest · digest built 2026-10-03T12:08:31+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud
Sparse MoE Day: Kolibri-1's 1M Context, Halved Score Memory, and 177B on a 12GB Card
Today's strongest signals are open weights and local inference plumbing: Aleph Alpha's Kolibri-1 lands as a 78B/3.46B-active Apache-2.0 model with 1M-token context, WHIRL ships a native Windows HIP engine for the Radeon R9700, and a llama.cpp PR halves Qwen Flash Next's indexer memory. Hugging Face's post-training team also published a long multi-harness RL guide, and Claude Code users get new mods for handoff compaction, code-review triage, and skill-budget debugging.
Open Weights and Local Inference
Aleph Alpha's Kolibri-1 is the headline open-weights drop: 78B parameters, 3.46B active, up to 1M-token context, Apache 2.0. On the inference side, WHIRL brings a native Windows HIP engine to Radeon R9700 owners, a llama.cpp PR halves Qwen Flash Next's indexer memory, and a community setup runs the 177B Qwen3.8-Flash-Next at 11-15 tok/s on a single 12GB RTX 5070. If you're on edge silicon, the RK3588 sparse-attention benchmarks are a useful reminder to measure decode and prefill separately.
Training and Context Engineering
Hugging Face's post-training team published a long guide to multi-harness RL with TRL and Harbor, aimed at training open models that generalize across coding scaffolds rather than overfitting one. On the context side, the Dynamic Tool Output Compression paper proposes letting the agent hide and unhide persisted tool outputs each turn — a reversible alternative to lossy compaction that cut tokens and steps on DeepSWE tasks.
Claude Code and MCP Tooling
Claude Code users get three practical additions today: handoff-compact automates the handoff-plus-/clear routine on autocompact, pr-proof fact-checks AI review comments against the code, and a community post documents how skill descriptions silently get truncated out of context. For MCP work, Moka is a local playground that shows the exact JSON-RPC traffic each provider receives, and Hearth is an Ollama UI that filters models by your GPU's VRAM before download.
Today's findings
-
#1 Aleph Alpha Kolibri-1: 78B MoE, 3.46B Active, 1M Context, Apache 2.0repo
A 78B-parameter MoE with only 3.46B active parameters and up to 1M-token context ships under Apache 2.0.
architecture78B parameters — only 3.46B fire per tokenroutertop-4 gateidle per token · 78B totalfires per token · 3.46Bweightedmerge≈4.4%of parameters active at once1Mmax context, tokensApache 2.0open-weights licenseWhy it matters: Sparse MoE at this active-parameter count means near-dense quality at a fraction of the compute, and 1M context opens long-document and repo-scale tasks on local hardware.
How to apply: Pull the weights from Hugging Face, run through llama.cpp or vLLM with an MoE-aware quant, and benchmark against your current 27B-class dense model on long-context retrieval.
open-weightsmoelong-contextlocal-llm
-
#2 Hugging Face's Multi-Harness RL Guide: Training Open Models Inside Real Coding Harnessespaper
HF's post-training team published a long guide on training open models across different coding harnesses using TRL and the Harbor framework.
RL post-trainingOne environment or many? The scaffold-overfitting trade-offSingle harness- Most RL recipes train against one coding environment
- Model overfits that scaffold
- Skills may not transfer
Multi-harness RL- Rotate coding harnesses during training
- Built on TRL + Harbor
- Generalizes across agent scaffolds
Vary the harnesses so gains hold across scaffoldsHF post-training team guide · open models · run small jobs before scalingWhy it matters: Most RL recipes assume a single environment; this shows how to train against multiple harnesses so models generalize across agent scaffolds instead of overfitting one.
How to apply: Read the guide, then wire your own harness into TRL + Harbor and run a small multi-harness RL job on a 7B-30B open model before scaling.
rlfine-tuningagentsopen-weights
Read more: The ultimate guide to multi-harness RL · The ultimate guide to multi-harness RL · The ultimate guide to multi-harness RL
-
#3 WHIRL: Native Windows HIP Inference Engine for Radeon AI PRO R9700repo
An Apache-2.0 C++/HIP inference engine runs Qwen3.8-27B MXFP4 natively on Windows with up to 2.5x llama.cpp prefill and 3x server throughput.
WHIRL · AMD RADEON ON WINDOWSNative Windows inference that leaves llama.cpp behind3×server throughput vs llama.cppup to 2.5× faster prefill0WSL or Linux required — HIP runs directly on Windows~28 MBApache-2.0 C++/HIP release zip27BQwen3.8 MXFP4 on Radeon AI PRO R9700Install the Adrenalin driver, unpack the zip, and benchmark your prefill-heavy workloads against llama.cpp.Why it matters: AMD RDNA users have been stuck on WSL or Linux for serious inference; a native Windows path with no llama.cpp underneath removes a major friction point.
How to apply: Grab the ~28MB release zip, install the Adrenalin driver, and benchmark your Qwen3.8-27B fine-tune against your current llama.cpp setup on prefill-heavy workloads.
amdinferencewindowslocal-llm
-
#4 llama.cpp PR Halves Qwen Flash Next Indexer Score Memoryrepo
A llama.cpp pull request (#29825) cuts the indexer score memory for Qwen Flash Next roughly in half, freeing VRAM for longer context.
Why it matters: Indexer memory is often the binding constraint when running Qwen Flash Next locally; halving it directly buys context length or higher quant.
How to apply: Track or cherry-pick PR #29825 in your llama.cpp build, then re-measure max context at your target quant on the same GPU.
llama.cppquantizationmemorylocal-llm
-
#5 Running Qwen3.8-Flash-Next 177B at 11-15 tok/s on a Single 12GB RTX 5070technique
A llama.cpp expert-streaming setup runs a 177B MoE at 11-15 tok/s on one 12GB GPU plus 32GB DDR4.
Local MoE expert streaming177B model, 12GB GPU: live experts hot, the rest idle in RAMHOTRTX 5070 · 12GB VRAM Active experts resident each decode stepCOLDSystem RAM · 32GB DDR4 Idle experts, streamed in as routedResident-expert count tuned to VRAM; DDR bandwidth caps decode at 11-15 tok/s.Why it matters: Shows that expert offload plus careful streaming can make frontier-scale MoE models usable on consumer hardware, not just datacenter GPUs.
How to apply: Replicate the llama.cpp expert-streaming config, tune the number of resident experts to your VRAM, and expect DDR bandwidth to dominate decode speed.
local-llmllama.cppmoeinference
Read more: Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
-
#6 Dynamic Tool Output Compression: Let the Agent Hide and Unhide Tool Outputspaper
A paper proposes letting the agent toggle visibility of persisted tool outputs each turn, cutting tokens and steps while improving solve rate on DeepSWE tasks.
Agent context engineeringHide and unhide: tool outputs become reversible1Run tooloutput lands in the window2Hidepersist it, drop from context3Reasontokens freed for planning4Unhiderestore verbatim on demandtoggle per turnNothing is summarized or deleted — hiding reverses; paper reports fewer tokens, fewer steps, higher DeepSWE solve rate.Why it matters: It's a reversible alternative to lossy context compaction, and it keeps the agent in control of what stays in context instead of a fixed summarizer.
How to apply: Persist tool outputs outside the context window, expose a hide/unhide tool to the model, and A/B it against your current compaction strategy on your own agent traces.
contextagentspapertool-use
Read more: Dynamic Tool Output Compression for Adaptive Context Management
-
#7 handoff-compact: Automates the Handoff + /clear Routine on Autocompacttool
A Claude Code mod intercepts the autocompact trigger and runs a handoff-plus-clear routine automatically, keeping context short without losing the thread.
CLAUDE CODE MOD · HANDOFF-COMPACTTurns the autocompact trigger into a handoff-and-clear loopThe loop repeats as long as the session runs — context stays short, the thread survives in the handoff doc.Why it matters: Long unattended Claude Code sessions burn tokens on turns past 200k context and degrade from context rot; this automates the manual fix.
How to apply: Install the mod, set your autocompact threshold, and review the generated handoff docs to tune what survives the clear.
claude-codecontexttoolingagents
Read more: handoff-compact, a mod that does the handoff + /clear routine for you every time autocompact fires
-
#8 pr-proof: Claude Code Skills That Fact-Check AI Code Review Commentstool
Three Claude Code skills treat each AI review comment as a claim, trace the code, and verdict it — removing 34% of CodeRabbit noise while keeping 93% of real bugs.
Why it matters: AI review bots generate plausible-but-wrong comments; a verification layer that reads callers and execution paths turns noisy reviews into a usable signal.
How to apply: Install via the Claude Code plugin marketplace, point it at your review bot's output, and tune the verdict thresholds against your own labeled PRs.
claude-codecode-reviewagentstooling
-
#9 Moka: Local LLM + MCP Playground With Raw JSON-RPC Inspectiontool
An MIT-licensed local sandbox lets you test prompts and MCP servers visually, inspect raw JSON-RPC traffic and tool waterfalls, and export to LangGraph.
Agent ToolingMoka: a local playground for LLM + MCP debuggingMokanpx @mokalabs/sandboxrunPoint it at your Ollama or Anthropic endpoint and read the raw JSON-RPC traffic and tool waterfalls to debug tool schemaMost agent bugs are broken tool schemas or silent timeouts, not the model — seeing the exact request each provider getsWhy it matters: Most agent bugs are broken tool schemas or silent timeouts, not the model; seeing the exact request each provider gets makes those debuggable.
How to apply: Run `npx @mokalabs/sandbox`, point it at your Ollama or Anthropic endpoint, and use the JSON-RPC view to debug tool schemas before shipping.
mcptoolinglocal-llmdebugging
Read more: Open-source: a local LLM + MCP playground where you can see the exact request each provider gets · A local inspector for MCP tool calls, latency, and raw JSON-RPC traffic · [Open Source] Local UI to test MCP tools + prompts and export directly to LangGraph Python
-
#10 Hearth: Ollama Web UI That Filters Models by Your GPU's VRAMtool
A lightweight local web UI for Ollama browses the full model library and shows which models fit your NVIDIA GPU's VRAM before you download.
Why it matters: Model selection is the most common local-LLM time sink; pre-filtering by VRAM and auto-picking the lightest capable model removes guesswork.
How to apply: Run it against your local Ollama on 127.0.0.1, use the Get Models screen to shortlist fits, and enable Auto model pick to avoid needless swaps.
ollamalocal-llmtoolingui
Read more: Built a local chat UI for Ollama — browse/pull models by what fits your GPU
-
#11 Sparse Attention on RK3588: 1.58x Faster Decode at 4K, 18% Slower at 1Ktechnique
Benchmarks on RK3588 with Qwen3-VL-2B show decode-side sparse attention helps at 4K context but prefill-side sparse attention quietly costs 18% at 1K.
Why it matters: Two different things get called 'sparse attention' and only one is free; knowing which side you're enabling prevents a month of silent slowdowns.
How to apply: Benchmark decode and prefill separately at your real context lengths before enabling sparse attention, and disable prefill-side sparsity for short-context workloads.
inferenceedgebenchmarkattention
Read more: Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K
-
#12 Your Claude Skills May Be Silently Dropped From Contexttip
Skill names and descriptions go into every session, but the list gets truncated with no warning — 12 of 91 skills never fired because their descriptions were cut.
Claude Code skillsSkills you pay context for, but the model never sees13%12 of 91 skills never fired: descriptions silently truncatedNames + descriptions load into every session; overflow entries get cut with no warning — audit tokens, trim or merge.Why it matters: If you maintain a large skill library, some skills are invisible to the model even though /name still works, so you're paying context for nothing.
How to apply: Audit your skill descriptions' total token count, trim or merge low-value skills, and verify each skill's description actually appears in the session context.
claude-codecontextskillstooling
Read more: found out why 12 of my skills never fired