Edition 2026-09-30 latest · digest built 2026-09-30T12:06:00+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud

Context Audits and Live Supervision: Stress-Testing Agents Before They Touch Payments

Today's strongest signals are operational: a 61-day Claude Code context audit, Anthropic's data on how little coding work can be fully delegated, and a stress test showing tool-calling agents can still authorize rogue payments. Local-model builders also get fresh llama.cpp support, a 27B agent model, and a 44%-shorter reasoning post-train.

Claude Code and Agent Operations

The most actionable Claude-centric work today is about controlling context and supervision, not just model choice. A 61-day transcript audit shows how hooks, MEMORY.md, CLAUDE.md, skills, subagents, and compaction determine what actually reaches Claude Code's window. Anthropic's own session data reinforces that long-horizon coding still needs live execution visibility: unobserved file and shell mutations compound quickly, and full delegation remains rare for complex work.

Agent Security and Tooling

Tool-calling agents are getting real access to payments, graphs, and production APIs, so the failure modes are shifting from bad text to unauthorized state changes. A purchasing-agent stress test found 5 of 16 adversarial attacks bypassed guardrails and authorized payments. LangChain teams are also hitting hallucinated tool JSON schemas, while MCP builders are learning to expose zoom levels instead of giant context dumps.

Local Models and Edge Inference

Local and open-weight releases keep arriving with practical deployment details. Oído runs int8 speech recognition on an ESP32-S3, GLM-5.3-Flash gets llama.cpp support, and BAAI's AREX-2 targets long-horizon agent loops on Qwen3.8. For efficiency, LessThink-Qwen3-4B cuts reasoning tokens by 44%, SelfJev rebuilds Jev-style classification on Qwen3.5-4B, and dual-3090 plus Strix Halo benchmarks show how much engine and quantization choices matter.

Today's findings

  1. #1 Measure What Actually Reaches Claude Code's Context Windowtechnique

    A 61-day transcript audit shows how to control Claude Code context with hooks, MEMORY.md, CLAUDE.md, skills, subagents, and compaction.

    61-day transcript audit
    Six levers decide what reaches Claude Code's context window
    CLAUDE.md instructions loaded every window
    MEMORY.md persistent memory index
    Hooks log & inject at events
    Skills loaded only on demand
    Subagents never in the main window
    Compaction summarizes, drops state
    always in window fires on demand isolated context trims history
    Log what enters each window, then tune injectors and boundaries before compaction.

    Why it matters: Context bloat and compaction silently drop critical state; measuring what reaches the window prevents agent amnesia and wasted tokens.

    How to apply: Instrument Claude Code sessions with hook events and a MEMORY.md index; log what enters each window, then tune CLAUDE.md, skills, and subagent boundaries before compaction.

    claude-codecontextmemoryagents

    Read more: I measured what actually reaches Claude Code's context window over 61 days of my own transcripts · I measured what actually reaches Claude Code's context window over 61 days of my own transcripts · I measured what actually reaches Claude Code's context window over 61 days of my own transcripts

  2. #2 Anthropic Data: Long-Horizon Coding Still Needs Live Supervisionpaper

    Across 400k+ Claude Code sessions, engineers fully delegate only 0–20% of coding tasks without oversight; unobserved state mutations compound.

    Anthropic · Claude Code telemetry
    At most 20 in 100 coding tasks are fully delegated without oversight
    0–20%
    of tasks a engineer hands off completely, no human watching
    400k+ sessions: unobserved file and shell mutations compound — the other 80+ need checkpoints or a review gate.

    Why it matters: It quantifies why fire-and-forget agents fail on multi-file work and where to place human validation.

    How to apply: Add live execution visibility to agent loops: checkpoint after file or shell mutations, require review gates on high-stakes subtasks, and avoid full delegation on legacy or multi-file changes.

    agentsclaude-codesupervisionevaluation

    Read more: Anthropic report reveals engineers only fully delegate 0 to 20 percent of coding tasks without live supervision · Anthropic report on Claude Code sessions shows long horizon task success still depends on live execution visibility

  3. #3 Stress-Test Tool-Calling Agents Before They Touch Paymentstechnique

    16 adversarial attacks against a ReAct purchasing agent bypassed guardrails 5 times and authorized rogue payments.

    Tool-Calling Agent Security
    Semantic guardrails alone don't hold
    5/16
    adversarial attacks bypassed the guardrails and authorized rogue payments
    ≈31% bypass rate
    16
    adversarial payloads fired at a ReAct purchasing agent
    5
    slipped through → rogue payments authorized
    3
    hard defenses: schema tests, deterministic gates, out-of-band confirm
    Function-calling into checkout tools is a different threat surface than chat — gate irreversible actions outside the LLM

    Why it matters: Function-calling access to checkout or payment tools creates a different threat surface than chatbot injection; semantic guardrails alone are insufficient.

    How to apply: Run adversarial payloads against tool schemas, enforce deterministic policy gates outside the LLM, and require out-of-band confirmation for irreversible transactions.

    agentssecuritytool-useguardrails

    Read more: I stress-tested an autonomous purchasing agent against 16 transactional attacks. 5 bypassed system guardrails and authorized rogue payments (Traces & Post-Mortem)

  4. #4 Stop LangChain Agents From Hallucinating Tool JSON Schemastechnique

    Validate tool arguments against strict schemas before execution to avoid burning credits and crashing multi-step agent loops.

    LANGCHAIN AGENTS
    A schema gate before every tool call
    1
    Draft args
    LLM emits raw JSON
    2
    Schema gate
    Pydantic / JSON Schema
    3
    Execute tool
    Only if valid
    retry with error feedback
    One malformed arg at step 6 fails the whole run — the gate catches it before the API call.

    Why it matters: A single malformed parameter on step 6 can fail the whole run; local schema validation catches it before production API calls.

    How to apply: Wrap tools with Pydantic or JSON Schema validation, add retry-with-error-feedback, and test agent loops against invalid argument cases in dev.

    langchainagentstool-usevalidation

    Read more: How we stopped our LangChain agents from hallucinating tool JSON schemas & crashing in dev

  5. #5 Design MCP Tools With Zoom Levels, Not Giant Dumpstechnique

    MCP servers that let agents edit visual graphs should expose project overview, group, and full-graph views instead of one huge context dump.

    MCP tool design
    Zoom levels, not one giant dump
    Giant dump
    Zoom levels
    Read at the right altitude; scope and validate every mutation
    Laddered graph views keep editing agents cheap on context and make write operations safer.

    Why it matters: Context-efficient MCP tool design keeps agents from burning tokens before they start editing and makes write operations safer.

    How to apply: Model read and write MCP tools with hierarchical zoom levels, explicit mutation scopes, and validation before applying graph edits.

    mcpagentstool-designcontext

    Read more: Designing MCP tools for agents that edit a visual graph: six decisions and why we made them

  6. #6 Oído Runs Speech Recognition on a $5 Microcontrollerrepo

    An open-source int8 Conformer-CTC model beats Whisper-tiny on noisy speech while running on an ESP32-S3 with no GPU.

    Why it matters: Edge ASR is now practical for low-power devices, enabling local voice interfaces without cloud calls.

    How to apply: Try the lokutor-ai/oido repo and live_demo.py to benchmark the chip arithmetic on your laptop mic, then deploy to ESP32-S3 for offline voice capture.

    speechedgeopen-sourcequantization

    Read more: Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)

  7. #7 GLM-5.3-Flash Lands in llama.cpprepo

    A new llama.cpp PR adds GLM-5.3-Flash (GLM5-Next) support so you can run it locally on your own machine.

    RUNS ON YOUR MACHINE
    GLM-5.3-Flash lands in llama.cpp
    GLM-5.3-Flash
    model GGUF READY
    GLM5-Next
    ggml-org/llama.cppPR #27773GGUF quantizations
    runTrack PR #27773, build from the branch, and test GGUF quants against your current local model.
    Local inference for a fresh open model — private coding and agent workflows on your own hardware.

    Why it matters: Local inference support for a fresh open model expands options for private coding and agent workflows.

    How to apply: Track PR #27773 in ggml-org/llama.cpp, build from the branch, and test GGUF quantizations against your current local model.

    llama.cpplocal-llmggufglm

    Read more: add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp

  8. #8 BAAI/AREX-2: A 27B Long-Horizon Agent Model on Qwen3.8paper

    AREX-2 learns to propose, measure, reflect, and revise over multiple test-time rounds, targeting long-horizon agent tasks.

    Open-weight agents
    AREX-2: propose, measure, reflect, revise — then rerun the round
    draft actionscore resultdiagnose missesupdate planProposeMeasureReflectRevise
    27B long-horizon agent on Qwen3; open weights fine-tune and run locally.

    Why it matters: Open-weight agent models that improve through iterative test-time loops can be fine-tuned and run locally.

    How to apply: Evaluate AREX-2 on your multi-step agent benchmarks, then fine-tune or quantize it for local deployment if it beats your current Qwen baseline.

    agentslocal-llmqwenopen-weights

    Read more: BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

  9. #9 LessThink-Qwen3-4B Cuts Reasoning Tokens by 44%paper

    A post-trained Qwen3-4B spends 44% fewer tokens on reasoning while keeping knowledge and answer style, all on one GPU.

    Why it matters: Shorter reasoning traces reduce latency and cost for local agents without retraining from scratch.

    How to apply: Use the LessThink pipeline as a template for post-training your own small reasoning model, then measure token savings on your task set.

    reasoningfine-tuninglocal-llmefficiency

    Read more: LessThink-Qwen3-4B: the same model, with far less thinking [P]

  10. #10 SelfJev Rebuilds Jev-Style Classification on Qwen3.5-4Brepo

    An open-weight 4B classifier with a shared-prefix tree hits ~140 ms per call on one H100 and is fine-tunable for your own cases.

    Why it matters: Fast, local multiple-choice classification over long text is useful for routing, memory filtering, and agent pre-processing.

    How to apply: Clone the SelfJev setup, fine-tune on cases your current classifier misses, and use it as a local pre-step before RAG or agent memory writes.

    classifiersopen-weightsfine-tuninglocal-llm

    Read more: I rebuilt a Jev-style classifier on Qwen3.5-4B: shared-prefix tree, open weights, fine-tunable, ~140 ms on one H100 · I rebuilt a Jev-style classifier on Qwen3.5-4B: shared-prefix tree, open weights, fine-tunable, ~140 ms on one H100

  11. #11 Run Qwen3.8-27B at 8-Bit on Dual RTX 3090stechnique

    A vLLM recipe with NVLink and DFlash2 hits 115 tok/s decode and 262K context while keeping 8-bit weights.

    Why it matters: Shows that high-fidelity local inference can be fast on consumer dual-GPU rigs, not just 4-bit quantizations.

    How to apply: Replicate the vLLM INT8 W8A16 setup, enable NVLink and DFlash2, and A/B against your llama.cpp Q8_0 baseline for decode and prefill.

    vllmlocal-llmquantizationinference

    Read more: Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers)

  12. #12 Benchmark Local Engines for Qwen3.8-Flash-Next on Strix Halotechnique

    On AMD Strix Halo, Halogen v0.14.0 is fastest, but open-source gufo loads 4x faster from cold and wins follow-up first-token latency.

    Engine bake-off · Strix Halo
    Halogen wins raw speed — gufo wins the waiting game
    vs
    Halogen v0.14.0
    gufo (open source)
    Overall speed
    Fastest
    —
    Cold-load time
    —
    4x faster
    Follow-up first token
    —
    Faster
    Halogen v0.14.0 wins the row gufo (open source) wins the row
    Engine choice swings agent responsiveness on unified-memory Strix Halo — compare CIRU too via the LlamaStash harness bef

    Why it matters: Engine choice materially changes local agent responsiveness on unified-memory hardware.

    How to apply: Use the LlamaStash benchmark harness to compare Halogen, gufo, and CIRU on your Strix Halo box before standardizing your local stack.

    local-llmbenchmarksstrix-haloinference

    Read more: Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo · Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

Looking for topic trends and crawl volume over time? See Trends.