4
technique
A custom llama.cpp branch runs Ternary Bonsai 2 27B entirely in 12GB VRAM at 128K context, reaching ~80-90 t/s code and 250+ t/s edits on an Intel Arc B580.
Why it matters: It expands local inference options beyond Nvidia and Apple Silicon, and shows ternary/quantized models can be practical for code editing on affordable GPUs.
How to apply: Try the Torchit1/llama.cpp arc-b580 branch with the Ternary Bonsai 2 27B weights; on CPU, ik_llama.cpp now also supports the model for fallback testing.
local-llmllama.cppintel-arcquantization
5
repo
New GSQ-RCO quantized GGUFs for Qwen3.8-Flash-Next include a 50% expert-pruned Coder build at roughly 1.89 bits per weight.
Why it matters: Aggressive quantization plus expert pruning can make large MoE coding models fit local memory budgets, which is useful for self-hosted coding agents.
How to apply: Download the GSQ-RCO GGUF variants, benchmark the pruned Coder build against your coding tasks, and compare quality/latency before swapping it into a local agent.
quantizationgguflocal-llmmoe
6
repo
WebBrain is an open-source browser agent that works with LM Studio or Ollama and includes a 450M browser-specific vision model for local screenshot understanding.
Why it matters: It reduces dependence on cloud multimodal models for browser automation and makes local browser agents more feasible on modest hardware.
How to apply: Clone WebBrain, point it at your local Ollama or LM Studio endpoint, and use the webbrain-vl-2-450M model for browser perception before escalating to a larger model.
agentsbrowserlocal-llmvision
8
technique
A prompt block tested across 360 A/B runs on GLM 5.3 and GLM 5.3 Flash cut wasted thinking by up to 70% by enforcing premise checks and single-approach discipline.
Why it matters: Token waste and meandering reasoning are common in coding agents; simple global instructions can reduce cost and latency without changing the model.
How to apply: Add the nine thinking-discipline rules to your agent's global instructions, especially the premise check and finish-one-approach rule, then measure output tokens and task success.
promptingagentscodingefficiency
9
tip
A faithfulness check passed twelve runs because the prompt already made the defect impossible; the check was fine but measured nothing.
Why it matters: Eval suites that cannot fail give false confidence, especially for RAG and agent outputs where silent regressions are costly.
How to apply: Before trusting an LLM judge, deliberately inject the defect it is supposed to catch and confirm the judge fails; if it cannot, rewrite the check or the prompt.
evaluationllm-judgetestingrag
10
technique
A team stopped letting LLMs do compliance math and moved hard limits, currency conversion, and rolling windows into a deterministic SQLite/Python core with episodic memory for context.
Why it matters: It is a reusable architecture for any agent that must respect strict rules while still recalling institutional exceptions and past decisions.
How to apply: Split your agent into a deterministic rule engine with veto power and a memory layer for waivers/history; let the LLM propose actions but never own the arithmetic or hard caps.
agentsragmemorycompliance
11
technique
A production OCR pipeline handles 40M+ documents and 5,000-50,000-page bundles by treating each job as a multi-PDF unit of work rather than a single file.
Why it matters: Naive OCR/RAG ingestion breaks on large regulated corpora; batching, job-level tracking, and failure recovery are the difference between a demo and a production pipeline.
How to apply: Model ingestion around job bundles, not individual PDFs; add per-bundle progress, retries, and validation so a single bad file does not sink a 50k-page workload.
ragocrpipelinesproduction
12
tip
A Claude Code bug can revert files from all other conversations in the same project when you revert a conversation turn with code changes enabled.
Why it matters: Silent cross-conversation file rollbacks can destroy work in multi-session projects, so teams should avoid the feature until it is fixed.
How to apply: Do not enable file reverting when reverting conversation turns in Claude Code; use git commits or manual snapshots for rollback instead.
claudeclaude-codetoolinggit