Dev Tooling & Infra
The last mile of a pull request — rerunning checks, clearing conflicts, answering review notes — is the part that actually eats an afternoon.
Visual Studio Code·2026-09-02
Open Source & Self-Hostable✓ verified
Changing model or provider is normally an application rewrite; a translating proxy turns it into a routing rule you can measure both sides of.
NVIDIA·2026-09-02
Frontier Models✓ verified
Third Flash release in six weeks at the same rate card — the cheap tier is where most production traffic actually runs, and working harder per task is a cost change even when the price is not.
Google DeepMind·2026-09-02
Open Source & Self-Hostable
Inference in the browser is the cheapest deployment there is — no server and no per-token bill — and fast GPU operations across mismatched devices have been the missing floor under it.
Hugging Face·2026-09-01
Dev Tooling & Infra
If your merge rule counts approvals, this is the first setting under which a machine can satisfy it — worth deciding on purpose rather than finding out during a release.
GitHub·2026-09-01
Frontier Models✓ verified
The change that shows up on a bill is the cache-read price, which is where long agentic runs spend; Devin's team said it is what finally made a Fable-class model economical for their code review.
Anthropic·2026-09-01
Dev Tooling & Infra
Vercel names the case out loud: one person’s unsupervised coding agent could previously drain a shared team budget, and now it cannot.
Vercel·2026-08-31
Agent Frameworks & Orchestration
The single line worth stealing: write down the done condition before the agent starts, and check it with something deterministic rather than another model.
n8n·2026-08-31
Research
If you get early access to a model with safeguards turned down, you are now expected to run it in a hardened, monitored sandbox — the testing-side obligations are being written down.
Anthropic·2026-08-31
Agent Frameworks & Orchestration✓ verified
Two upgrade traps in one release: downgrading now means a manual SQLite restore, and shared sessions are not a permission boundary you can lean on.
OpenClaw·2026-08-30
Agent Frameworks & Orchestration
The distance from "we should try an agent for this" to a deployed, Slack-reachable one that calls your own MCP servers is now a dashboard form.
Vercel·2026-08-28
Open Source & Self-Hostable
A 49B-active open-weight model with a million-token window makes long-document work a self-hosting decision rather than a closed-API one.
Tencent·2026-08-28
Dev Tooling & Infra
Two of the three are opt-out-before-the-date changes — the retention move and the review-effort default both land automatically on 28 September if nobody acts.
GitHub·2026-08-28
Dev Tooling & Infra✓ verified
Coordinated disclosure assumes an attacker needs the patch. If the public discussion is enough, the embargo window your project plans around has already closed.
Anil Madhavapeddy·2026-08-28
Agent Frameworks & Orchestration✓ verified
The tool-calling layer that made software agents useful is being pointed at instruments and machines; if it holds, the same harness that drives an API drives a microscope or a robot arm.
Anthropic·2026-08-27
Dev Tooling & Infra
A session that survives leaving the tool it started in is the first real answer to agent work being trapped per-application, and the per-turn token breakdown is the first place a team can see what a long run actually cost.
Visual Studio Code·2026-08-26
Dev Tooling & Infra
The transferable part is not the line count, it is the precondition: this worked because there was something to check the output against. Teams without an oracle for the work are not in the same situation and should not read this as the same result.
Paul Dix, via Simon Willison·2026-08-26
MCP & Interop✓ verified
When agents arrive through declared tools rather than the page, what your site exposes - and what it refuses - becomes a product decision rather than a search one.
OpenAI·2026-08-25
Research
Failure recovery and decomposition are where multi-step agents actually break, and this is a training recipe aimed at them rather than at a benchmark score.
arXiv·2026-08-24
Agent Frameworks & Orchestration
A supported interception point is the difference between auditing what an agent did and only reading about it afterwards.
Microsoft·2026-08-22
Dev Tooling & Infra
Separating "which model, configured how" from "what am I asking" is the cheapest way to keep a prompt comparison honest across vendors.
llm, by Simon Willison·2026-08-22
Open Source & Self-Hostable
One misplaced system message was busting the cache on every request — worth checking how your own prompt is assembled before blaming the model for being slow.
Ollama·2026-08-21
Agent Frameworks & Orchestration
If the harness carries the outcome, agent quality is engineering you own rather than a model you rent, and it is the half you can actually iterate on.
TechCrunch·2026-08-21
Agent Frameworks & Orchestration✓ verified
If the software around the model — memory, supervision, tools — is worth seventy points on a long task, then agent quality is mostly engineering you control rather than a model you wait for.
NVIDIA·2026-08-21