Edition 2026-10-03 latest · digest built 2026-10-03T12:08:31+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud

Sparse MoE Day: Kolibri-1's 1M Context, Halved Score Memory, and 177B on a 12GB Card

Today's strongest signals are open weights and local inference plumbing: Aleph Alpha's Kolibri-1 lands as a 78B/3.46B-active Apache-2.0 model with 1M-token context, WHIRL ships a native Windows HIP engine for the Radeon R9700, and a llama.cpp PR halves Qwen Flash Next's indexer memory. Hugging Face's post-training team also published a long multi-harness RL guide, and Claude Code users get new mods for handoff compaction, code-review triage, and skill-budget debugging.

Open Weights and Local Inference

Aleph Alpha's Kolibri-1 is the headline open-weights drop: 78B parameters, 3.46B active, up to 1M-token context, Apache 2.0. On the inference side, WHIRL brings a native Windows HIP engine to Radeon R9700 owners, a llama.cpp PR halves Qwen Flash Next's indexer memory, and a community setup runs the 177B Qwen3.8-Flash-Next at 11-15 tok/s on a single 12GB RTX 5070. If you're on edge silicon, the RK3588 sparse-attention benchmarks are a useful reminder to measure decode and prefill separately.

Training and Context Engineering

Hugging Face's post-training team published a long guide to multi-harness RL with TRL and Harbor, aimed at training open models that generalize across coding scaffolds rather than overfitting one. On the context side, the Dynamic Tool Output Compression paper proposes letting the agent hide and unhide persisted tool outputs each turn — a reversible alternative to lossy compaction that cut tokens and steps on DeepSWE tasks.

Claude Code and MCP Tooling

Claude Code users get three practical additions today: handoff-compact automates the handoff-plus-/clear routine on autocompact, pr-proof fact-checks AI review comments against the code, and a community post documents how skill descriptions silently get truncated out of context. For MCP work, Moka is a local playground that shows the exact JSON-RPC traffic each provider receives, and Hearth is an Ollama UI that filters models by your GPU's VRAM before download.

Today's findings

  1. #1 Aleph Alpha Kolibri-1: 78B MoE, 3.46B Active, 1M Context, Apache 2.0repo

    A 78B-parameter MoE with only 3.46B active parameters and up to 1M-token context ships under Apache 2.0.

    architecture
    78B parameters — only 3.46B fire per token
    router
    top-4 gate
    idle per token · 78B totalfires per token · 3.46B
    weighted
    merge
    ≈4.4%
    of parameters active at once
    1M
    max context, tokens
    Apache 2.0
    open-weights license

    Why it matters: Sparse MoE at this active-parameter count means near-dense quality at a fraction of the compute, and 1M context opens long-document and repo-scale tasks on local hardware.

    How to apply: Pull the weights from Hugging Face, run through llama.cpp or vLLM with an MoE-aware quant, and benchmark against your current 27B-class dense model on long-context retrieval.

    open-weightsmoelong-contextlocal-llm

    Read more: Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0

  2. #2 Hugging Face's Multi-Harness RL Guide: Training Open Models Inside Real Coding Harnessespaper

    HF's post-training team published a long guide on training open models across different coding harnesses using TRL and the Harbor framework.

    RL post-training
    One environment or many? The scaffold-overfitting trade-off
    Single harness
    • Most RL recipes train against one coding environment
    • Model overfits that scaffold
    • Skills may not transfer
    Multi-harness RL
    • Rotate coding harnesses during training
    • Built on TRL + Harbor
    • Generalizes across agent scaffolds
    Vary the harnesses so gains hold across scaffolds
    HF post-training team guide · open models · run small jobs before scaling

    Why it matters: Most RL recipes assume a single environment; this shows how to train against multiple harnesses so models generalize across agent scaffolds instead of overfitting one.

    How to apply: Read the guide, then wire your own harness into TRL + Harbor and run a small multi-harness RL job on a 7B-30B open model before scaling.

    rlfine-tuningagentsopen-weights

    Read more: The ultimate guide to multi-harness RL · The ultimate guide to multi-harness RL · The ultimate guide to multi-harness RL

  3. #3 WHIRL: Native Windows HIP Inference Engine for Radeon AI PRO R9700repo

    An Apache-2.0 C++/HIP inference engine runs Qwen3.8-27B MXFP4 natively on Windows with up to 2.5x llama.cpp prefill and 3x server throughput.

    WHIRL · AMD RADEON ON WINDOWS
    Native Windows inference that leaves llama.cpp behind
    3×
    server throughput vs llama.cpp
    up to 2.5× faster prefill
    0
    WSL or Linux required — HIP runs directly on Windows
    ~28 MB
    Apache-2.0 C++/HIP release zip
    27B
    Qwen3.8 MXFP4 on Radeon AI PRO R9700
    Install the Adrenalin driver, unpack the zip, and benchmark your prefill-heavy workloads against llama.cpp.

    Why it matters: AMD RDNA users have been stuck on WSL or Linux for serious inference; a native Windows path with no llama.cpp underneath removes a major friction point.

    How to apply: Grab the ~28MB release zip, install the Adrenalin driver, and benchmark your Qwen3.8-27B fine-tune against your current llama.cpp setup on prefill-heavy workloads.

    amdinferencewindowslocal-llm

    Read more: WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput

  4. #4 llama.cpp PR Halves Qwen Flash Next Indexer Score Memoryrepo

    A llama.cpp pull request (#29825) cuts the indexer score memory for Qwen Flash Next roughly in half, freeing VRAM for longer context.

    Why it matters: Indexer memory is often the binding constraint when running Qwen Flash Next locally; halving it directly buys context length or higher quant.

    How to apply: Track or cherry-pick PR #29825 in your llama.cpp build, then re-measure max context at your target quant on the same GPU.

    llama.cppquantizationmemorylocal-llm

    Read more: qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp

  5. #5 Running Qwen3.8-Flash-Next 177B at 11-15 tok/s on a Single 12GB RTX 5070technique

    A llama.cpp expert-streaming setup runs a 177B MoE at 11-15 tok/s on one 12GB GPU plus 32GB DDR4.

    Local MoE expert streaming
    177B model, 12GB GPU: live experts hot, the rest idle in RAM
    HOT
    RTX 5070 · 12GB VRAM Active experts resident each decode step
    COLD
    System RAM · 32GB DDR4 Idle experts, streamed in as routed
    Resident-expert count tuned to VRAM; DDR bandwidth caps decode at 11-15 tok/s.

    Why it matters: Shows that expert offload plus careful streaming can make frontier-scale MoE models usable on consumer hardware, not just datacenter GPUs.

    How to apply: Replicate the llama.cpp expert-streaming config, tune the number of resident experts to your VRAM, and expect DDR bandwidth to dominate decode speed.

    local-llmllama.cppmoeinference

    Read more: Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

  6. #6 Dynamic Tool Output Compression: Let the Agent Hide and Unhide Tool Outputspaper

    A paper proposes letting the agent toggle visibility of persisted tool outputs each turn, cutting tokens and steps while improving solve rate on DeepSWE tasks.

    Agent context engineering
    Hide and unhide: tool outputs become reversible
    1
    Run tool
    output lands in the window
    2
    Hide
    persist it, drop from context
    3
    Reason
    tokens freed for planning
    4
    Unhide
    restore verbatim on demand
    toggle per turn
    Nothing is summarized or deleted — hiding reverses; paper reports fewer tokens, fewer steps, higher DeepSWE solve rate.

    Why it matters: It's a reversible alternative to lossy context compaction, and it keeps the agent in control of what stays in context instead of a fixed summarizer.

    How to apply: Persist tool outputs outside the context window, expose a hide/unhide tool to the model, and A/B it against your current compaction strategy on your own agent traces.

    contextagentspapertool-use

    Read more: Dynamic Tool Output Compression for Adaptive Context Management

  7. #7 handoff-compact: Automates the Handoff + /clear Routine on Autocompacttool

    A Claude Code mod intercepts the autocompact trigger and runs a handoff-plus-clear routine automatically, keeping context short without losing the thread.

    CLAUDE CODE MOD · HANDOFF-COMPACT
    Turns the autocompact trigger into a handoff-and-clear loop
    context growsfires near 200kthread → doccontext wipedwork continuesFillTriggerHandoffClearResume
    The loop repeats as long as the session runs — context stays short, the thread survives in the handoff doc.

    Why it matters: Long unattended Claude Code sessions burn tokens on turns past 200k context and degrade from context rot; this automates the manual fix.

    How to apply: Install the mod, set your autocompact threshold, and review the generated handoff docs to tune what survives the clear.

    claude-codecontexttoolingagents

    Read more: handoff-compact, a mod that does the handoff + /clear routine for you every time autocompact fires

  8. #8 pr-proof: Claude Code Skills That Fact-Check AI Code Review Commentstool

    Three Claude Code skills treat each AI review comment as a claim, trace the code, and verdict it — removing 34% of CodeRabbit noise while keeping 93% of real bugs.

    Why it matters: AI review bots generate plausible-but-wrong comments; a verification layer that reads callers and execution paths turns noisy reviews into a usable signal.

    How to apply: Install via the Claude Code plugin marketplace, point it at your review bot's output, and tune the verdict thresholds against your own labeled PRs.

    claude-codecode-reviewagentstooling

    Read more: I built a Claude Code skill that checks AI code review comments against the code. On CodeRabbit's reviews it removed 34% of the noise and kept 93% of real bugs

  9. #9 Moka: Local LLM + MCP Playground With Raw JSON-RPC Inspectiontool

    An MIT-licensed local sandbox lets you test prompts and MCP servers visually, inspect raw JSON-RPC traffic and tool waterfalls, and export to LangGraph.

    Agent Tooling
    Moka: a local playground for LLM + MCP debugging
    Moka
    tool MIT
    Runs fully local; exports flows to LangGraph
    npx @mokalabs/sandbox
    runPoint it at your Ollama or Anthropic endpoint and read the raw JSON-RPC traffic and tool waterfalls to debug tool schema
    Most agent bugs are broken tool schemas or silent timeouts, not the model — seeing the exact request each provider gets

    Why it matters: Most agent bugs are broken tool schemas or silent timeouts, not the model; seeing the exact request each provider gets makes those debuggable.

    How to apply: Run `npx @mokalabs/sandbox`, point it at your Ollama or Anthropic endpoint, and use the JSON-RPC view to debug tool schemas before shipping.

    mcptoolinglocal-llmdebugging

    Read more: Open-source: a local LLM + MCP playground where you can see the exact request each provider gets · A local inspector for MCP tool calls, latency, and raw JSON-RPC traffic · [Open Source] Local UI to test MCP tools + prompts and export directly to LangGraph Python

  10. #10 Hearth: Ollama Web UI That Filters Models by Your GPU's VRAMtool

    A lightweight local web UI for Ollama browses the full model library and shows which models fit your NVIDIA GPU's VRAM before you download.

    Why it matters: Model selection is the most common local-LLM time sink; pre-filtering by VRAM and auto-picking the lightest capable model removes guesswork.

    How to apply: Run it against your local Ollama on 127.0.0.1, use the Get Models screen to shortlist fits, and enable Auto model pick to avoid needless swaps.

    ollamalocal-llmtoolingui

    Read more: Built a local chat UI for Ollama — browse/pull models by what fits your GPU

  11. #11 Sparse Attention on RK3588: 1.58x Faster Decode at 4K, 18% Slower at 1Ktechnique

    Benchmarks on RK3588 with Qwen3-VL-2B show decode-side sparse attention helps at 4K context but prefill-side sparse attention quietly costs 18% at 1K.

    Why it matters: Two different things get called 'sparse attention' and only one is free; knowing which side you're enabling prevents a month of silent slowdowns.

    How to apply: Benchmark decode and prefill separately at your real context lengths before enabling sparse attention, and disable prefill-side sparsity for short-context workloads.

    inferenceedgebenchmarkattention

    Read more: Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K

  12. #12 Your Claude Skills May Be Silently Dropped From Contexttip

    Skill names and descriptions go into every session, but the list gets truncated with no warning — 12 of 91 skills never fired because their descriptions were cut.

    Claude Code skills
    Skills you pay context for, but the model never sees
    13%
    12 of 91 skills never fired: descriptions silently truncated
    Names + descriptions load into every session; overflow entries get cut with no warning — audit tokens, trim or merge.

    Why it matters: If you maintain a large skill library, some skills are invisible to the model even though /name still works, so you're paying context for nothing.

    How to apply: Audit your skill descriptions' total token count, trim or merge low-value skills, and verify each skill's description actually appears in the session context.

    claude-codecontextskillstooling

    Read more: found out why 12 of my skills never fired

Looking for topic trends and crawl volume over time? See Trends.