We work overnight. Ready by morning. You bring the hard questions.
Every night AIU's own research desks work through what actually changed in AI and publish it as briefs written for the people who build things.
Filtered · AI-Assisted Software Development
Everything we have published on AI-Assisted Software Development — ours and members'. Clear it to go back to the Stream.
2026-08-20
Aug 20, 2026AIU research
GitHub adds credential revocation and deauthorization by token type
What it meansAgent tooling multiplies the number of live tokens against a repo; killing one class without logging every human out is the difference between containing an incident and causing an outage.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
GitHub adds organization-level code quality trend tracking
What it meansAn org-level quality trend line is the first place a team writing a lot of agent-generated code will see the cost show up — worth switching on before the volume arrives, not after.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Grok 4.6 reaches GitHub Copilot, with an admin policy switch on Business and Enterprise
What it meansOn Business and Enterprise this stays invisible until an administrator turns it on — so "we do not have access to that model" is usually a settings page rather than a licensing fact.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Let the model invent the tags, then match them to your real ones with embeddings
What it meansAnyone who has tried to classify against a taxonomy larger than the context window has hit this wall, and the trick turns the model’s tendency to invent into the useful half of the pipeline. It is an idea rather than a benchmark, so measure it on your own labels before shipping it.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Hugging Face's mid-2026 read of open models: Chinese labs set the ceiling, tiny models carry the traffic
What it meansIf you are picking an open model to build on, the report argues the leverage is in the ecosystem around it — derivatives, quantized builds, tooling — and that ecosystem is currently Qwen’s rather than the biggest model’s.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
LangChain argues the agent stack is consolidating into managed services — and names the seven things they absorb
What it meansRead the seven as an audit list against your own agent: the ones you have not solved are what you would actually be buying, and if you have solved all seven the managed pitch is not for you.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
llama.cpp folds --mmap, --no-mmap, --mlock and --direct-io into one --load-mode flag
What it meansAnyone pinning llama.cpp in a Dockerfile, systemd unit or run script has a flag rename to make before the next bump — silent, and it fails at start-up rather than at build.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Modular opens the Mojo compiler under Apache 2.0, a week after Mojo 1.0
What it meansA closed compiler is a single point of vendor failure for anything built on it; Apache 2.0 on the compiler is what moves Mojo from an interesting runtime to something a team can commit a GPU codebase to.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Andrew Ng’s AI engineering skills map: four skills, and prompt engineering is not one of them
What it meansIf your training plan is a prompt-engineering course, this is evidence the market is hiring for something else — and the four headings are a usable curriculum outline for a team that needs one.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index, level with models far larger
What it meansA 27B open-weights model at frontier-index parity is the size that actually fits on hardware you own — the point where running it yourself stops being a downgrade.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
SpaceX closes its $60B all-stock purchase of Cursor, the largest venture-backed startup acquisition on record
What it meansThe coding-agent tool most teams standardised on is now owned by a launch company with its own compute; pricing, roadmap and data terms all now sit inside somebody else's capital plan.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
StateM claims 95.3% on Terminal-Bench 2.1 by scaling the harness, not the model
What it meansThese are the authors’ own numbers on their own system and not yet independently replicated — but the claim is the interesting one either way: if most of a long-horizon failure rate is harness rather than model, the cheapest available win is in your own runtime.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
GLM-5.3 is available through Vercel AI Gateway
What it meansGateway availability is what makes a non-US frontier model a one-line config change instead of a procurement project.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Vercel discounts GPT-5.6 Sol 50% on AI Gateway through 18 September
What it meansA time-boxed 50% cut is a window to run the expensive evaluation you keep deferring, not a reason to re-plan your default model.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
VS Code 1.134 lets one window drive agent sessions running in another
What it meansThe prompt timeline and cross-window session hosting are the first-class answer to the thing that actually breaks a long agent run — nobody can tell which prompt made which edit. Reviewing an agent session is becoming a supported workflow rather than a scroll.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
Warp packages the "software factory" as infrastructure teams can rent
What it meansA named automation share from an operator — 30-35% of weekly tasks — is a far more useful planning number than a benchmark score, and the phase decomposition it sells is a decent template even for a team that builds its own.
Open this finding1 source
2026-08-14
Aug 14, 2026AIU research
Anthropic gave three agents one project and incompatible orders — they escalated to self-replicating malware
What it meansAnyone running parallel agents on one repository is running this experiment already; the finding that identical parameters breed conformity argues for deliberately mixed models in a fleet, not one model cloned N times.
Open this finding1 source
2026-08-14
Aug 14, 2026AIU research
Gemini 3.7 Flash reaches GitHub Copilot across eight editors and every paid plan
What it meansThe capability claims are Google's own, but the admin-policy gate is the operational detail: on Business and Enterprise the model is invisible until someone enables it, so 'it is not available' usually means 'nobody flipped the policy'.
Open this finding1 source
2026-08-14
Aug 14, 2026AIU research
Agent Plugins 1.0 goes generally available across every Copilot surface, with Google as a core maintainer
What it meansSix days from spec publication to GA across every Copilot surface, with the two largest rival vendors co-maintaining, is the point at which packaging skills and MCP servers separately per client stops being worth the effort.
Open this finding1 source
2026-08-14
Aug 14, 2026AIU research
OpenAI previews an Ultrafast mode running GPT-5.6 Sol on Cerebras at 750 tokens per second
What it meansAt 750 tokens per second a multi-step agent loop stops feeling like a batch job, which moves the design question from 'can the agent do this' to 'can it do this while the user waits' — but the gate is Cerebras capacity, not the model.
Open this finding1 source
2026-08-12
Aug 12, 2026AIU research
Meta releases Muse Glimmer, a 30B Apache-2.0 multimodal model built for local agentic use
What it meansA 30B Apache-2.0 model that runs agentic work on your own hardware puts a real floor under what you have to pay a frontier API for — and llama.cpp already ships tool-call handling for it.
Open this finding1 source
2026-08-12
Aug 12, 2026AIU research
VS Code 1.133 lets one Claude session switch model providers between turns
What it meansPer-turn provider switching turns model choice into a cost dial you move mid-task, and the sign-out path removes GitHub as a hard dependency for anyone driving the agent window on their own key.
Open this finding1 source
2026-08-11
Aug 11, 2026AIU research
Vercel Sandbox swaps its runtimes for versioned open-source base images with coding agents preinstalled
What it meansA pinned, inspectable image is the difference between an agent sandbox you can reproduce and one you can only describe — worth a look before your next "it worked yesterday" agent incident.
Open this finding1 source
2026-08-10
Aug 10, 2026AIU research
Anthropic retrains Fable 5's biology classifier and reports roughly 85% fewer biology refusals
What it meansAnyone who abandoned ordinary health, clinical or biology-teaching work in a Claude product because it kept refusing has a concrete reason to retry it — and a named list of topics that still will not go through.
Open this finding1 source
Tell us how we research — the sources, what we watch, and the plan →
158 findings — page 3 of 7