4
repo
Unee is an Apache-2.0 GGUF/Ollama model family that makes calibrated decisions and chats, with the 2B scoring 88% on DecideBench.
Why it matters: Small decision models can replace expensive LLM calls for routing, classification, and guardrail decisions on local hardware.
How to apply: Pull the GGUF via Ollama, benchmark it on your routing/classification tasks, and use it as a fast first-stage decision layer.
small-modelsdecisionggufollama
5
tool
Haiku 5.5 is 75% cheaper, yet independent agent-pipeline tests show it does not always beat DeepSeek V4.1 Flash on cost or quality.
Why it matters: Cheaper per-token pricing does not guarantee cheaper per-task; teams need end-to-end evals before swapping models.
How to apply: Run your own agent pipeline against Haiku 5.5 and your current model, measuring cost per completed task, latency, and failure modes.
claudehaikuagentscost
8
technique
Video RAG needs event-sized chunk windows, not paragraph boundaries, and document-style chunking quietly degrades retrieval.
Why it matters: Teams extending RAG to video will get blurry embeddings and split events if they reuse document chunking habits.
How to apply: Choose chunk windows based on scene or event boundaries, test overlap, and evaluate retrieval against the moments users actually query.
ragvideochunkingembeddings
9
paper
On SEC 10-K retrieval, a cross-encoder reranker mattered more than the retriever, and BM25+dense hybrid actually underperformed dense alone.
Why it matters: It challenges default assumptions: adding hybrid search or swapping retrievers may not help without reranking and dataset-specific evals.
How to apply: Run a small ablation on your own corpus; prioritize reranker quality, and test hybrid fusion rather than assuming it wins.
ragretrievalrerankingevaluation
11
technique
A two-agent pattern—one implements, one reviews diffs with fresh context and argues against them—was the biggest quality win in a native Mac app postmortem.
Why it matters: It gives coding agents a lightweight adversarial review loop without requiring a human to catch every design or implementation mistake.
How to apply: Add a second reviewer agent to your CI or local workflow, feed it only the diff and requirements, and require it to justify objections before merge.
agentscode-reviewworkflowclaude-code
12
technique
A practical walkthrough shows Qwen3-8B fine-tuning with Unsloth on one RTX 5090, making local domain adaptation more accessible.
Why it matters: Teams can specialize an 8B local model for internal tasks without a multi-GPU cluster or a full training platform.
How to apply: Use the Unsloth recipe as a starting point, swap in your domain dataset, and validate against a held-out set before deploying the adapter.
fine-tuningunslothqwenlocal-llm