Edition 2026-09-29 latest · digest built 2026-09-29T12:04:43+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud
Leak-Proof Agent Sandboxes, Self-Rewriting Harnesses, and 100K Context on One GPU
Today's actionable AI work clusters around agent safety and self-improvement: OpenShell's sandbox held up against a local qwen3:8b agent, while RRSI shows harness rewrites can lift benchmark scores without touching weights. Local inference also got more practical with Qwen 3.8 27B running at 100K context on 16GB AMD and Ternary Bonsai 2 hitting 128K on a 12GB Intel Arc. On the engineering side, new quantization releases, a local browser agent, prompt-discipline rules, and hard-won eval and RAG pipeline lessons round out the day.
Agent Safety and Self-Improvement
The most immediately useful thread today is agent control. NVIDIA's OpenShell sandbox stopped a local qwen3:8b agent from leaking a secret in every trial, but its auto-approval path still opened new hosts in all twelve attempts, which is exactly the kind of gap teams should test before trusting an agent with shell access. RRSI pushes the other direction: an agent can rewrite its own harness while weights stay frozen, improving Terminal-Bench 2.1 from 74.2% to 80.2% and lifting all six held-out benchmarks. Together they suggest a practical split: let the model propose changes, but enforce hard boundaries outside the model.
Local Models Get Faster and Smaller
Local inference had a strong day. A Vulkan llama.cpp build runs Qwen 3.8 27B Q4 XS at 100K context on a 16GB RX 7800 XT, while a custom llama.cpp branch puts Ternary Bonsai 2 27B entirely in 12GB VRAM on an Intel Arc B580 with 128K context and fast code-edit throughput. New GSQ-RCO GGUFs for Qwen3.8-Flash-Next include a 50% expert-pruned Coder build at roughly 1.89 bpw, and WebBrain offers an open-source browser agent that can run locally through LM Studio or Ollama with a 450M browser-specific VLM. Fast-Jev-Compaction also gives Claude Code users a self-hosted alternative to /compact summaries.
Evaluation and Pipeline Discipline
Several items are less about new models and more about not fooling yourself. A prompt block tested across 360 A/B runs cut wasted coding-agent thinking by up to 70% by forcing premise checks and single-approach discipline. A separate write-up shows an LLM judge that passed twelve runs because the prompt made the defect impossible, a reminder to inject failures before trusting evals. For production systems, a compliance agent moved hard math into a deterministic SQLite/Python core with episodic memory for context, and an OCR pipeline built for 50,000-page jobs treats each bundle as a unit of work rather than a single file.
Today's findings
-
#1 NVIDIA OpenShell Sandbox Blocks Agent Secret Leakstool
An open-source Apache 2.0 sandbox with default-deny egress stopped a local qwen3:8b agent from leaking a secret in 10/10 trials, but auto-approval still opened new hosts in 12/12.
Agent guardrailsmediumSandbox blocks secret leaks — auto-approval reopens new hosts10/10leak trials stopped by default-deny egress12/12runs where blanket auto-approval still opened a new hostApache 2.0open-source sandbox wrapping the agent's shell toolaffected scopeLocal & self-hosted agents with shell tools — tested on qwen3:8bmedium severity — badge colour grades the riskOpenShell v0.1.2: keep egress default-deny; never blanket-approve new hosts.Why it matters: Agent tool access is the new attack surface; a sandbox that enforces filesystem and network policy outside the model is a practical guardrail for local or self-hosted agents.
How to apply: Run OpenShell v0.1.2 around your agent's shell tool, keep egress default-deny, and avoid blanket auto-approval for new hosts; test with a malicious setup script before trusting the agent.
agentssecuritysandboxlocal-llm
Read more: We tested NVIDIA OpenShell with a local qwen3:8b agent: a malicious setup script leaked a secret 10/10 without it, 0/10 with it. But auto-approval opened new hosts in 12/12 trials · We tested NVIDIA OpenShell with a local qwen3:8b agent: a malicious setup script leaked a secret 10/10 without it, 0/10 with it. But auto-approval opened new hosts in 12/12 trials
-
#2 RRSI Lets Agents Rewrite Their Own Harnesspaper
Google Research's Apache 2.0 RRSI improves Terminal-Bench 2.1 from 74.2% to 80.2% by letting an agent rewrite its harness while model weights stay frozen.
AGENT SELF-IMPROVEMENTSame frozen model, a better harnessFixed harness- Terminal-Bench 2.1: 74.2%
- Harness never changes
- Weights frozen
Self-rewritten harness- Terminal-Bench 2.1: 80.2%
- Agent edits its own harness
- Weights still frozen
RRSI (Google Research, Apache 2.0): gains from harness-only self-modificationWhy it matters: It shows agent performance can improve through harness and tool-loop changes without retraining, which is directly applicable to coding agents and tool-use pipelines.
How to apply: Study the RRSI repo/paper and try a constrained self-modification loop: freeze weights, let the agent propose harness edits, and validate against held-out tasks before accepting changes.
agentsself-improvementbenchmarksopen-source
-
#3 Qwen 3.8 27B Q4 Hits 100K Context on 16GB AMDtechnique
A Vulkan llama.cpp build runs Qwen 3.8 27B Q4 XS with q8/q5 KV cache at ~30 t/s decode and 100K context on a 16GB RX 7800 XT.
VULKAN LLAMA.CPPLong-context 27B runs local on mid-range AMD100Ktokens of context on one 16GB GPU≈30 tok/s decode27Bparams at Q4 (IQ4_XS)q8/q5quantized KV cacheRX 7800 XT16GB VRAM, Vulkan buildStrong 27B long-context coding and RAG on mid-range AMD hardware — no Nvidia or Mac required.Why it matters: It proves a strong 27B local model can handle long-context coding or RAG workloads on mid-range AMD hardware, not just high-end Nvidia or Macs.
How to apply: Build llama.cpp with -DGGML_VULKAN=ON, pull unsloth's Qwen3.8-27B-UD-IQ4_XS.gguf plus mmproj, and use q8/q5 KV cache to fit 100K context in 16GB VRAM.
local-llmllama.cppquantizationvulkan
Read more: Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide
-
#4 Ternary Bonsai 2 Runs Fast on 12GB Intel Arctechnique
A custom llama.cpp branch runs Ternary Bonsai 2 27B entirely in 12GB VRAM at 128K context, reaching ~80-90 t/s code and 250+ t/s edits on an Intel Arc B580.
Why it matters: It expands local inference options beyond Nvidia and Apple Silicon, and shows ternary/quantized models can be practical for code editing on affordable GPUs.
How to apply: Try the Torchit1/llama.cpp arc-b580 branch with the Ternary Bonsai 2 27B weights; on CPU, ik_llama.cpp now also supports the model for fallback testing.
local-llmllama.cppintel-arcquantization
Read more: Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context all in VRAM, ~80-90 t/s code, 250+ t/s edits, 2-4x faster than the official fork · Ternary bonsai 2 sur ik llama.cpp
-
#5 GSQ-RCO GGUFs Bring Qwen3.8-Flash-Next to ~1.89 bpwrepo
New GSQ-RCO quantized GGUFs for Qwen3.8-Flash-Next include a 50% expert-pruned Coder build at roughly 1.89 bits per weight.
Why it matters: Aggressive quantization plus expert pruning can make large MoE coding models fit local memory budgets, which is useful for self-hosted coding agents.
How to apply: Download the GSQ-RCO GGUF variants, benchmark the pruned Coder build against your coding tasks, and compare quality/latency before swapping it into a local agent.
quantizationgguflocal-llmmoe
Read more: [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw · [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
-
#6 WebBrain Runs Browser Agents Locally with a 450M VLMrepo
WebBrain is an open-source browser agent that works with LM Studio or Ollama and includes a 450M browser-specific vision model for local screenshot understanding.
Why it matters: It reduces dependence on cloud multimodal models for browser automation and makes local browser agents more feasible on modest hardware.
How to apply: Clone WebBrain, point it at your local Ollama or LM Studio endpoint, and use the webbrain-vl-2-450M model for browser perception before escalating to a larger model.
agentsbrowserlocal-llmvision
-
#7 Fast-Jev-Compaction Replaces Claude Code /compact Summariestool
Fast-Jev-Compaction is an MIT-licensed, self-hosted tool that swaps Claude Code's /compact summary for Jev-scored deletion of stale tool calls.
Context compactionTwo ways to clear a long coding sessionClaude Code /compact- Summarizes the whole session
- Summarization happens in the cloud
- Condensed history can lose useful state
Fast-Jev-Compaction- Jev-scores each tool call
- Deletes only the stale calls
- Runs locally and Docker-friendly
Swap the cloud summary for local, scored deletion — useful state survives.fast-jev-compaction · MIT · self-hostedWhy it matters: Context compaction is a recurring pain in long coding sessions; a local, Docker-friendly alternative can preserve useful state without cloud summarization.
How to apply: Install fast-jev-compaction in your Claude Code workflow, configure it to run locally, and test it on long tool-call histories to see whether stale calls are dropped cleanly.
claudecontexttoolinglocal
Read more: Fast-Jev-Compaction Review: /compact Without a Summary
-
#8 Nine Prompt Rules Cut Coding Agent Wasted Thinkingtechnique
A prompt block tested across 360 A/B runs on GLM 5.3 and GLM 5.3 Flash cut wasted thinking by up to 70% by enforcing premise checks and single-approach discipline.
Why it matters: Token waste and meandering reasoning are common in coding agents; simple global instructions can reduce cost and latency without changing the model.
How to apply: Add the nine thinking-discipline rules to your agent's global instructions, especially the premise check and finish-one-approach rule, then measure output tokens and task success.
promptingagentscodingefficiency
Read more: 9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)
-
#9 Make Sure Your LLM Judge Can Actually Failtip
A faithfulness check passed twelve runs because the prompt already made the defect impossible; the check was fine but measured nothing.
Why it matters: Eval suites that cannot fail give false confidence, especially for RAG and agent outputs where silent regressions are costly.
How to apply: Before trusting an LLM judge, deliberately inject the defect it is supposed to catch and confirm the judge fails; if it cannot, rewrite the check or the prompt.
evaluationllm-judgetestingrag
Read more: My LLM judge passed twelve runs in a row. It turned out it couldn't have failed.
-
#10 Hybrid Deterministic Core Plus Episodic Memory for Compliancetechnique
A team stopped letting LLMs do compliance math and moved hard limits, currency conversion, and rolling windows into a deterministic SQLite/Python core with episodic memory for context.
Why it matters: It is a reusable architecture for any agent that must respect strict rules while still recalling institutional exceptions and past decisions.
How to apply: Split your agent into a deterministic rule engine with veto power and a memory layer for waivers/history; let the LLM propose actions but never own the arithmetic or hard caps.
agentsragmemorycompliance
Read more: Why we stopped letting LLMs do compliance math and gave our agent episodic memory instead
-
#11 OCR Pipeline Design for 50,000-Page Jobstechnique
A production OCR pipeline handles 40M+ documents and 5,000-50,000-page bundles by treating each job as a multi-PDF unit of work rather than a single file.
Why it matters: Naive OCR/RAG ingestion breaks on large regulated corpora; batching, job-level tracking, and failure recovery are the difference between a demo and a production pipeline.
How to apply: Model ingestion around job bundles, not individual PDFs; add per-bundle progress, retries, and validation so a single bad file does not sink a 50k-page workload.
ragocrpipelinesproduction
Read more: How we built an OCR pipeline that survives a 50,000-page workload
-
#12 Claude Code File Revert Can Roll Back Other Conversationstip
A Claude Code bug can revert files from all other conversations in the same project when you revert a conversation turn with code changes enabled.
Why it matters: Silent cross-conversation file rollbacks can destroy work in multi-session projects, so teams should avoid the feature until it is fixed.
How to apply: Do not enable file reverting when reverting conversation turns in Claude Code; use git commits or manual snapshots for rollback instead.
claudeclaude-codetoolinggit
Read more: PSA: Don't enable reverting files when reverting conversation turns