Edition 2026-10-08 latest · digest built 2026-10-08T12:06:18+00:00 · ⚙ ollama:deepseek-v4.1-flash:cloud
Silent Config Overrides, Rewindable Agent Runs, Plus Audio-to-Tool Calls in the Browser
Today's actionable AI work centers on agent reliability and local model efficiency. Claude Code users should audit AGENTS.md precedence, while unloop brings time-travel debugging to long agent runs. NeuDecide and Unee shrink tool-calling and decision routing to tiny open models, and new llama.cpp/RAG findings offer concrete speed and retrieval improvements. Haiku 5.5 is cheaper, but real pipeline evals show mixed results.
Agent Tooling Gets More Debuggable and Configurable
Today's agent news is less about bigger models and more about making them reliable. Claude Code's AGENTS.md precedence is a silent config trap, while unloop adds state rewind for long-running agents. The driver+reviewer split and Anthropic's Oct 7 clarification both point to production discipline: keep context, review diffs, and know your subscription boundaries.
Small Models and Local Inference Keep Getting Faster
NeuDecide and Unee show the tiny-model trend: audio-to-tool calls in 43 MB and calibrated decision models at 0.8B/2B. On the inference side, llama.cpp's GDN state-column PR and Qwen3.8 on a single Radeon R9700 target faster long-context local work. Haiku 5.5 is cheaper, but independent evals warn that per-token savings don't always translate to per-task wins.
RAG Lessons: Rerank, Chunk for Events, and Test Hybrids
Two RAG items stand out: video RAG needs event-sized chunks rather than document paragraphs, and a SEC 10-K ablation found reranking mattered more than retriever choice while hybrid search lost to dense alone. Perplexity's pplx-embed-v2-late adds a multimodal embedding option with a small query encoder, useful for PDF and image search.
Today's findings
-
#1 Claude Code's AGENTS.md Precedence Can Silently Disable Ittip
Claude Code now reads AGENTS.md, but a personal CLAUDE.md can override it and silently switch off your repo instructions.
Config precedenceWhich file wins in Claude Code?AGENTS.md- Shared repo instructions
- Read by Codex, Cursor, Claude Code
- Holds your project rules
CLAUDE.md- Personal — repo or home dir
- Overrides AGENTS.md in Claude Code
- No error when it wins
In Claude Code the personal file wins — silently. Keep one source of truth or configure which file is read.Audit for CLAUDE.md before trusting AGENTS.md.Why it matters: Teams sharing agent instructions across Codex, Cursor, and Claude Code can lose project rules without an error, causing inconsistent agent behavior.
How to apply: Audit your repo and home directory for CLAUDE.md vs AGENTS.md; keep one source of truth or explicitly configure the file Claude Code should read.
claude-codeagents-mdconfiguration
-
#2 unloop Adds Time-Travel Debugging for Long-Running Agentstool
unloop snapshots agent state so you can rewind and replay long agent runs instead of restarting after context summarization breaks.
unloop · agent debuggingRewind the run instead of restarting it1Agent runsstate snapshotted every turn2Context dropssummarization loses paths & constraints3Run failstypically by turns 10–154Rewindstate restored to the broken turn5Resumeinspect, fix, continuerun continues from the saved turnLong-run failures become a rewindable, reproducible debug point — not a from-scratch restart.Why it matters: Agent failures after 10-15 turns often come from lost file paths or constraints; state rewind makes those failures debuggable and reproducible.
How to apply: Wrap your agent loop with unloop's snapshot/rewind API, then inspect and resume from the turn where context was dropped.
agentsdebuggingstatellmdevs
Read more: unloop – time-travel debugging & state rewind for long-running LLM agents
-
#3 NeuDecide: 43 MB Audio-to-Tool Model with WASM Demorepo
NeuDecide maps audio directly to tool calls with arguments, skipping transcription, and runs in a browser via WASM.
NeuDecide · voice-to-tools repoAudio maps straight to tool calls — no transcription stage43 MBone end-to-end model: audio → tool call + argumentsskips the fragile ASR → LLM pipeline0transcription passesWASMruns in-browser, offlineApache-2.0licenseTry the WASM demo — offline voice commands become structured tool calls.Why it matters: Voice-driven tools can avoid a fragile ASR-to-LLM pipeline and run locally with a tiny Apache-2.0 model.
How to apply: Try the WASM demo, then integrate the model where you need offline voice commands to trigger structured tool calls.
audiotool-callingwasmopen-source
Read more: We’re open sourcing NeuDecide: a 43 MB audio-to-tool model with a WASM browser demo · We’re open sourcing NeuDecide: a 43 MB audio-to-tool model with a WASM browser demo
-
#4 Unee: 0.8B/2B Calibrated Decision Models for Routingrepo
Unee is an Apache-2.0 GGUF/Ollama model family that makes calibrated decisions and chats, with the 2B scoring 88% on DecideBench.
Why it matters: Small decision models can replace expensive LLM calls for routing, classification, and guardrail decisions on local hardware.
How to apply: Pull the GGUF via Ollama, benchmark it on your routing/classification tasks, and use it as a fast first-stage decision layer.
small-modelsdecisionggufollama
Read more: Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0) · Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0) · Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0)
-
#5 Claude Haiku 5.5 Ships Cheaper, but Real Pipelines Show Mixed Winstool
Haiku 5.5 is 75% cheaper, yet independent agent-pipeline tests show it does not always beat DeepSeek V4.1 Flash on cost or quality.
Why it matters: Cheaper per-token pricing does not guarantee cheaper per-task; teams need end-to-end evals before swapping models.
How to apply: Run your own agent pipeline against Haiku 5.5 and your current model, measuring cost per completed task, latency, and failure modes.
claudehaikuagentscost
Read more: Claude Haiku 5.5 is out and it is 75% cheaper than the last one · Haiku 5.5 vs DeepSeek V4.1 Flash on a real agent pipeline: ~3x cheaper per video, 2.6x faster, and it's not because of the per-token price · Haiku 5.5 vs DeepSeek V4.1 Flash on my 2 boring agent jobs: DeepSeek kept both
-
#6 Anthropic's Oct 7 Update Keeps Agent SDK and claude -p in Subscription Limitstip
Anthropic clarified that Claude Agent SDK, claude -p, and third-party apps still work within subscription limits, and now include API credits.
Anthropic · Oct 7 update3 automation paths stay inside subscription limits3 of 3 passClaude Agent SDKclaude -pThird-party appspass warn failPlans now also include API credits for burst or production workloads.Why it matters: Teams building on Claude can keep using subscription-based automation without moving everything to metered API billing.
How to apply: Review the updated docs, confirm your automation paths, and claim the new API credits for burst or production workloads.
claudeagent-sdkpricing
-
#7 llama.cpp PR Speeds Up Qwen Prompt Processing with GDN State Columnsrepo
A new llama.cpp CUDA PR assigns four GDN state columns per warp, targeting faster prompt processing for Qwen 3.x models.
llama.cpp · CUDAFaster prompt processing for Qwen 3.xllama.cpp PR #30087runTrack the PR in your build, then benchmark prefill on your Qwen 3.x GGUF workloads.Targets the prefill bottleneck behind long-repo reads in local coding agents.Why it matters: Prompt processing is often the bottleneck for long-context local coding agents; this can make large repo reads more practical.
How to apply: Track or test PR #30087 in your llama.cpp build, then benchmark prefill speed on your Qwen 3.x GGUF workloads.
llama.cppcudaperformanceqwen
-
#8 Video RAG Breaks Document RAG Assumptionstechnique
Video RAG needs event-sized chunk windows, not paragraph boundaries, and document-style chunking quietly degrades retrieval.
Why it matters: Teams extending RAG to video will get blurry embeddings and split events if they reuse document chunking habits.
How to apply: Choose chunk windows based on scene or event boundaries, test overlap, and evaluate retrieval against the moments users actually query.
ragvideochunkingembeddings
Read more: Your document RAG habits will quietly break your video RAG. Here are five places it happens.
-
#9 Retrieval Ablation: Reranking Beats Retriever Choice, Hybrid Can Losepaper
On SEC 10-K retrieval, a cross-encoder reranker mattered more than the retriever, and BM25+dense hybrid actually underperformed dense alone.
Why it matters: It challenges default assumptions: adding hybrid search or swapping retrievers may not help without reranking and dataset-specific evals.
How to apply: Run a small ablation on your own corpus; prioritize reranker quality, and test hybrid fusion rather than assuming it wins.
ragretrievalrerankingevaluation
-
#10 pplx-embed-v2-late: 0.6B Query Encoder for a 9B Multimodal Indexpaper
Perplexity's MIT-licensed pplx-embed-v2-late uses a small query encoder against a larger multimodal index for text, images, and PDF pages.
pplx-embed-v2-late · MIT licenseAn asymmetric pair: 0.6B on every query, 9B only at index timevsQuery encoderIndex encoderParams0.6B9BRunsevery queryindex build onlyInputstexttext, images, PDFsQuery encoder wins the row Index encoder wins the rowA 15× size gap: the 9B encoder runs once to build the index; every search runs just 0.6B parameters.Why it matters: It offers a practical asymmetric embedding design for multimodal RAG without running a huge encoder on every query.
How to apply: Evaluate the 0.6B query encoder with the 9B index on your PDF/image search workload, especially where latency and index size matter.
embeddingsmultimodalragmit
-
#11 Driver + Reviewer Agent Split Improves Shipped Code Qualitytechnique
A two-agent pattern—one implements, one reviews diffs with fresh context and argues against them—was the biggest quality win in a native Mac app postmortem.
Why it matters: It gives coding agents a lightweight adversarial review loop without requiring a human to catch every design or implementation mistake.
How to apply: Add a second reviewer agent to your CI or local workflow, feed it only the diff and requirements, and require it to justify objections before merge.
agentscode-reviewworkflowclaude-code
Read more: postmortem: 3 weeks shipping a native mac app on a driver + reviewer agent split (what held up, what i dropped) · how i split work between two coding agents to ship v3.0 of my native mac app (driver + reviewer pattern)
-
#12 Fine-Tuning Qwen3-8B with Unsloth on a Single RTX 5090technique
A practical walkthrough shows Qwen3-8B fine-tuning with Unsloth on one RTX 5090, making local domain adaptation more accessible.
Why it matters: Teams can specialize an 8B local model for internal tasks without a multi-GPU cluster or a full training platform.
How to apply: Use the Unsloth recipe as a starting point, swap in your domain dataset, and validate against a held-out set before deploying the adapter.
fine-tuningunslothqwenlocal-llm
Read more: Fine-Tuning Qwen3-8B with Unsloth on a Single RTX 5090