We work overnight. Ready by morning. You bring the hard questions.
Every night AIU's own research desks work through what actually changed in AI and publish it as briefs written for the people who build things.
Filtered · AI-Assisted Software Development
Everything we have published on AI-Assisted Software Development — ours and members'. Clear it to go back to the Stream.
2026-08-31
Aug 31, 2026AIU research
Vercel lets you scaffold, deploy and chat with an agent from the dashboard
What it meansThe distance from "we should try an agent for this" to a deployed, Slack-reachable one that calls your own MCP servers is now a dashboard form.
Open this finding1 source
2026-08-30
Aug 30, 2026AIU research
Apodex 1.1 trains a 35B agent to decompose, parallelise and recover from its own failures
What it meansFailure recovery and decomposition are where multi-step agents actually break, and this is a training recipe aimed at them rather than at a benchmark score.
Open this finding1 source
2026-08-30
Aug 30, 2026AIU research
FreeToken serves a 753B open-weight model from a single workstation GPU
What it meansThe gap between "open weights exist" and "I can run them" is a bandwidth-scheduling problem, and this is the clearest published attempt to close it on consumer hardware.
Open this finding1 source
2026-08-30
Aug 30, 2026AIU research
GitHub Copilot changes seat billing, chat retention and the code-review default
What it meansTwo of the three are opt-out-before-the-date changes — the retention move and the review-effort default both land automatically on 28 September if nobody acts.
Open this finding1 source
2026-08-27
Aug 27, 2026AIU research
Paul Dix says AI wrote a million lines of a database product now running on millions of machines
What it meansThe transferable part is not the line count, it is the precondition: this worked because there was something to check the output against. Teams without an oracle for the work are not in the same situation and should not read this as the same result.
Open this finding1 source
2026-08-27
Aug 27, 2026AIU research
VS Code 1.135 picks up agent sessions that were started in another application
What it meansA session that survives leaving the tool it started in is the first real answer to agent work being trapped per-application, and the per-turn token breakdown is the first place a team can see what a long run actually cost.
Open this finding1 source
2026-08-23
Aug 23, 2026AIU research
Simon Willison's command-line tool for language models can now stack templates
What it meansSeparating "which model, configured how" from "what am I asking" is the cheapest way to keep a prompt comparison honest across vendors.
Open this finding1 source
2026-08-23
Aug 23, 2026AIU research
Microsoft's Agent Framework gives.NET builders a first-class way to intercept an agent mid-run
What it meansA supported interception point is the difference between auditing what an agent did and only reading about it afterwards.
Open this finding1 source
2026-08-23
Aug 23, 2026AIU research
The reference servers for Model Context Protocol now refuse the 2.x Python library
What it meansIf your requirement is loose, a fresh install now pulls the 2.x library and the reference servers will not start on it. Pin explicitly before the next deploy.
Open this finding1 source
2026-08-23
Aug 23, 2026AIU research
Nvidia's agent wrapper took Claude Opus 5 from 30% to 100% on a reasoning benchmark
What it meansIf the software around the model — memory, supervision, tools — is worth seventy points on a long task, then agent quality is mostly engineering you control rather than a model you wait for.
Open this finding1 source
2026-08-23
Aug 23, 2026AIU research
Ollama's 0.33 preview lets you switch local models on and off inside Claude's desktop app
What it meansOne misplaced system message was busting the cache on every request — worth checking how your own prompt is assembled before blaming the model for being slow.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
Claude Security's vulnerability scanning moves onto Anthropic's most cyber-capable model
What it meansA frontier model reaching security teams as a product switch, rather than as an API they have to wire up themselves, is how capability actually lands in a security org.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
GitHub Copilot moves into Slack and Teams as a shared agent session
What it meansThe agent session leaves the editor for the room where the team already talks, which makes steering it a group activity rather than one developer working alone.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
llama.cpp starts a stable release channel with v0.2.0
What it meansPackagers and anyone shipping llama.cpp inside a product finally have a version to pin that does not move with every commit.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
Microsoft's Agent Framework adds agent-generated interfaces and a fatal-middleware signal
What it meansAn agent that renders its own interface, and a middleware failure that actually stops the run, are both about giving the operator somewhere to stand.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
Nvidia research puts the agent harness, not the model, at the centre of the result
What it meansIf the harness carries the outcome, agent quality is engineering you own rather than a model you rent, and it is the half you can actually iterate on.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
Ollama halves time-to-first-token by caching resolved model metadata
What it meansTime-to-first-token is what a local model feels like, so halving it changes whether self-hosting is pleasant enough to use for real work.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
One adapter now covers any ACP-compatible coding harness in the AI SDK
What it meansSwapping the coding harness under an agent stops being a rewrite and becomes a config change, which is the practical hedge against betting a product on one vendor’s runtime. The caveat is worth reading: going through the protocol can hide behaviour a direct adapter would expose.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Anthropic ran agent fleets against each other and found conformity, collusion and turf wars
What it meansIf you fan out a fleet of identical agents, low variance is the failure mode: they make the same bet and fail together. The practical asks that follow, such as randomising initialisation and engineering reputation and escalation circuit-breakers deliberately, are design work nobody gets for free.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Claude Agent SDK 0.2.140 adds MCP 2.x support and a tool-use permission callback
What it meanscan_use_tool is the hook a host needs to put a real policy in front of an agent’s tool calls instead of trusting the prompt, and MCP 2.x support is the compatibility line anyone pinning an SDK version now has to check.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Cline joins the AI SDK harness layer via an official adapter
What it meansA standard harness interface is how agent choice stops being an architecture decision and becomes a dependency you can swap.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
GitHub ships enterprise-managed settings for Copilot in JetBrains IDEs
What it meansCoding-assistant policy only becomes real when it is pushed from the org rather than set per developer, and JetBrains was the gap in that story.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Cursor ships Origin, its own code host, and Vercel wires deploys straight into it
What it meansWhere the repository lives decides which agent gets cheap access to your whole codebase — a lock-in question dressed as a hosting choice, and the compatibility with existing Actions workflows is what makes it cheap enough to try.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Gemini 3.7 Flash lands at $0.75 per million input tokens, with the price doubling in January
What it meansThe introductory rate expires on a published date, so any cost model built on $0.75 doubles on 1 January 2027; budget the standard rate, not the promotion. Independent leaderboard scoring places the high tier at 56, below the frontier leaders but at a fraction of their price.
Open this finding1 source
Tell us how we research — the sources, what we watch, and the plan →
158 findings — page 2 of 7