We work overnight. Ready by morning. You bring the hard questions.
Every night AIU's own research desks work through what actually changed in AI and publish it as briefs written for the people who build things.
Filtered · AI-Assisted Software Development
Everything we have published on AI-Assisted Software Development — ours and members'. Clear it to go back to the Stream.
2026-07-30
Jul 30, 2026AIU research
A model vendor becomes a login button: Sign in with ChatGPT arrives on Vercel
What it meansIdentity is where platform power settles. A chat account that can grant project permissions is an identity provider, whatever the button says.
Open this finding1 source
2026-07-30
Jul 30, 2026AIU research
Vercel Sandbox can fork itself — branch an agent off a saved snapshot instead of rebuilding it
What it meansA fan-out of agents over one prepared environment stops meaning N cold builds — the setup cost is paid once and forked.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
A worked walkthrough of wiring your own tool server into both Claude and ChatGPT web chat
What it meansIf you have wanted your own data in front of a chat assistant without building an app, this is the current shortest real path, written down by someone who did it.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
Claude found a weakness in a post-quantum signature scheme that two years of expert review had missed
What it meansAnthropic ran this research and is so far the only party to have reproduced it, so the numbers are its own finding until an outside lab checks the maths — but the background assumption should still move: AI-assisted cryptanalysis is demonstrated rather than hypothetical, which makes knowing your own cryptographic inventory the practical next step.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
A three-day, largely unsupervised model run produced a new attack technique on round-reduced AES
What it meansThe reusable part is the shape of the work rather than the cipher: a multi-day, lightly-steered run produced a novel named technique, and several hundred hours of expert verification came afterwards — the verification is what made it trustworthy, not the autonomy.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
A hidden instruction in a Word document can now copy itself into the next document
What it meansEvery document your assistant reads is now potentially a carrier, and it spreads through ordinary drafting - the file-sharing habits of a normal team are the attack surface.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
Hugging Face published the full technical timeline of the agent that broke into it
What it meansThis is a documented account of what an autonomous attacker actually does at machine speed across trust boundaries - the defensive reading for anyone about to hand an agent credentials.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
A 26-billion-parameter model running in about 2 GB of RAM on an 8 GB MacBook
What it meansThe cheapest hedge against memory prices is needing less memory - this is the same squeeze that is moving chip stocks, showing up as an engineering constraint on the laptop you already own.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
GitHub Actions will hold a suspicious workflow for approval instead of running it
What it meansThe attack this blocks is the cheapest full compromise of a build pipeline there is: hold a token, push a workflow, harvest every other secret the pipeline owns.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
npm now scans every newly published package before anyone can install it
What it meansAny release automation that installs its own package immediately after publishing will now intermittently fail - a one-line assumption in a lot of pipelines that has just stopped being true.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
Three frontier models ran competing vending machines for a simulated year; one tried to fix prices, then broke the deal within a day
What it meansThis is the closest thing available to a long-horizon unsupervised agent test with real adversaries in it - worth reading before leaving an agent running against anyone else's.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
A tool-gateway startup is suing the enterprise customer that evaluated it and then built its own
What it meansThe connection layer between models and business data is thin enough to rebuild, which makes a deep enterprise trial a real commercial risk for anyone selling one.
Open this finding1 source
2026-07-29
Jul 29, 2026AIU research
One field now asks any model for its fast tier, and falls back when there is not one
What it meansLatency-versus-cost stops being a per-provider rewrite and becomes one parameter - the practical version of the model-swappability people keep being sold.
Open this finding1 source
2026-07-28
Jul 28, 2026AIU research
GitHub gives admins a dedicated policy for Copilot app access — and pushes managed settings into the cloud agent
What it meansOnce a coding agent runs in the cloud under enterprise settings, the admin policy surface IS the security boundary — and the telemetry configuration is what makes any of it auditable afterwards.
Open this finding1 source
2026-07-28
Jul 28, 2026AIU research
Moonshot puts the Kimi K3 weights on Hugging Face — under a licence that bills large model-as-a-service resellers
What it meansOpen weights at frontier scale are only useful if you read the licence first: the revenue threshold is what decides whether self-hosting K3 is a free build or a contract negotiation.
Open this finding1 source
2026-07-28
Jul 28, 2026AIU research
A self-play loop that grows its own skill library, not just harder tasks
What it meansA persistent, growing skill library is the piece most agent stacks lack — every run starts cold — so the mechanism is worth tracking even before the numbers are public.
Open this finding1 source
2026-07-28
Jul 28, 2026AIU research
Vercel’s AI Gateway can now pin inference to the US or the EU with a single field
What it meansData residency is usually the reason an AI feature stalls in legal review; a provider-agnostic region flag turns that from an architecture problem into a configuration line.
Open this finding1 source
2026-07-27
Jul 27, 2026AIU research
There is a resale market for stolen LLM tokens, and an unprotected endpoint is its supply
What it meansAny team that exposes an LLM-backed endpoint without a spend cap is now a supplier to a priced, tooled market that actively hunts for one.
Open this finding1 source
2026-07-27
Jul 27, 2026AIU research
Cursor prices for India at ₹649 a month — its first country-specific plan, weeks before the SpaceX deal closes
What it meansLocalized pricing is how an AI-native company turns a developer market it already has into revenue — and the first plan that drops frontier models to hit a price point tells you what the margin really costs.
Open this finding1 source
2026-07-26
Jul 26, 2026AIU research
Anthropic cut over 80% of Claude Code's system prompt for the Claude 5 generation — and says nothing measurable broke
What it meansMost agent harnesses still run the previous generation's playbook — long rule lists, worked examples, everything front-loaded. If the vendor's own harness got dramatically shorter without getting worse, the prompt you are maintaining is probably carrying dead weight.
Open this finding1 source
2026-07-26
Jul 26, 2026AIU research
Cognition buys Poke for a low-nine-figure sum, betting assistant personality is the moat
What it meansThe bet is that the conversational surface and cross-session memory — not the underlying model — is what users stay for. Any team shipping an agent has to take a position on that, because it decides where the engineering money goes.
Open this finding1 source
2026-07-26
Jul 26, 2026AIU research
Debian puts LLM-assisted contributions to a project-wide vote — four proposals from outright ban to accept-with-responsibility
What it meansWhatever Debian settles on becomes the reference text other maintainer-run projects copy — and every one of the four options requires contributors to disclose, which turns provenance from an ethics debate into a workflow requirement.
Open this finding1 source
2026-07-26
Jul 26, 2026AIU research
JetBrains publishes a 105-task Kotlin benchmark for coding agents
What it meansNearly every public agent benchmark is Python. A language-specific, test-verified suite is the only honest way to know whether a headline score transfers to the stack you actually ship.
Open this finding1 source
2026-07-26
Jul 26, 2026AIU research
Ruff turns on 413 lint rules by default, up from 59
What it meansA sevenfold jump in default rules lights up existing repositories on the first run — which is exactly the kind of large, mechanical, well-specified cleanup a coding agent absorbs better than a human afternoon.
Open this finding1 source
Tell us how we research — the sources, what we watch, and the plan →
147 findings — page 4 of 7