All the research, one tab.
Every research item behind the daily briefs, newest first, in fast pages. Filter by topic, open any card's source, share any page — the URL is the state. The live map stays on the Brief page.
- Research
Anthropic says Opus 5 is its hardest model yet to prompt-inject — the evidence sits in the system card
Prompt injection is still the open hole in every tool-using agent. A vendor claim of improvement is worth tracking, but until it is independently measured your permission gates stay exactly as load-bearing as they were.
- Dev Tooling & Infra
Vercel workflow steps can now run 30 minutes, up from just over 13
Long agent steps have been split across invocations purely to dodge a timeout; a 30-minute ceiling removes a chunk of that plumbing.
- Research
One robot model, roughly 9,000 different hands: Generalist’s GEN-1 transfers across end effectors
Retraining per gripper is the tax that keeps robot deployments bespoke; a model that transfers across hardware is what turns a pilot into a fleet.
- Frontier Models✓ verified
Anthropic ships Claude Opus 5 — top-of-leaderboard scores at half its own flagship price
The price-per-capability floor moved again: work that needed the top-tier model last week may now run at half the cost, in the tools most teams already code in.
- Market & Business
Why enterprise AI stalls between the pilot and the system — the orchestration gap, in survey numbers
The failure mode named here — pilots that never become systems — is the one most AI programs are actually in, and the diagnosis points at integration and governance rather than model choice.
- Open Source & Self-Hostable✓ verified
Upstage releases Solar Open 2 — a 250B open-weight model built for long-horizon agent work
A self-hostable, independently benchmarked alternative to closed agent APIs — runnable on two H200s when quantized.
- Frontier Models✓ verified
OpenAI brings voice control to the desktop — ChatGPT can now drive your computer and direct agents hands-free
Voice-directed agent sessions are now a mainstream interaction pattern, not a demo — worth evaluating for hands-busy and accessibility-first workflows.
- Research
A team that ran fully autonomous coding agents for a year reports the catch: codebases decay
First-hand, long-duration evidence for keeping deterministic gates and review layers in agent-driven development rather than trusting full autonomy.
- Market & Business
Holiday Robotics raises $105M for FRIDAY, a wheeled humanoid built to work a full shift
Another bet that the winning factory form factor is wheels plus good hands rather than legs — and that full-shift uptime matters more than the demo.
- MCP & Interop
GitHub's MCP server adopts the stateless protocol spec ahead of the July 28 cutover
Teams running MCP servers or clients should check their stack against the stateless spec before the ecosystem finishes the cutover.
- Dev Tooling & Infra
GitHub Issues gets a throttle for autonomous agents — approvals, confidence scores, a reason for every action
A confidence threshold you can dial is the first practical answer to "how much do I let the agent just do?" — the same shape VS Code shipped for tool calls a week earlier.
- Market & Business✓ verified
ChatGPT Health opens to all US adults, with medical-record connections through roughly 2.2 million providers
A live template — and test case — for shipping AI products on top of regulated personal data at consumer scale.
- Frontier Models✓ verified
Black Forest Labs' FLUX 3 generates video with synchronized sound — and the same model can drive robots
Creative generation and robot control converging on one architecture is a strong signal for anyone planning tooling in either space.
- Market & Business✓ verified
Travis Kalanick's Atoms raises $1.7 billion to build task-specific industrial robots — not humanoids
One of the year's largest robotics rounds is a bet on vertical, task-specific automation over the humanoid narrative — a useful counter-signal for anyone tracking physical AI.
- Dev Tooling & Infra
AMD takes aim at the robotics compute default with unified memory and microsecond control loops
A credible second supplier for robot compute changes the negotiating position of everyone currently building on a single-vendor stack.
- Dev Tooling & Infra✓ verified
VS Code 1.130 lets the model judge risk before an agent tool call asks for approval
A shipping example of an LLM risk-prefilter layered ahead of hard approval gates — the pattern agent-tooling teams are converging on.
- Market & Business✓ verified
Synthesia moves beyond AI video into live roleplay coaching for the corporate-training market
AI-training vendors are shifting from selling content generation to selling scored practice — proof the skill transferred — a template every learning-and-development buyer and AI-content vendor will now be measured against.
- Dev Tooling & Infra✓ verified
PyPI now blocks adding new files to releases older than 14 days — a supply-chain hardening move
Python teams with slow multi-platform build pipelines should confirm all their wheels publish within 14 days; everyone else just got a safer dependency chain.
- Dev Tooling & Infra✓ verified
Patch now: Check Point SmartConsole and Microsoft SharePoint flaws join CISA's exploited list
On-prem SharePoint operators must patch and rotate machine keys — patching alone leaves stolen keys valid.
- Agent Frameworks & Orchestration✓ verified
OpenAI launches Presence — a managed platform for running governed enterprise AI agents
A governed build-vs-buy path for AI-staffed support lines — and a competitive marker for every agent-platform play.
- Research✓ verified
The OpenAI–Hugging Face breach traces back to a human mistake: a sandbox left open to the internet
Teams running agentic evals with relaxed refusals should audit sandbox egress first — the failure mode is ordinary infrastructure, not exotic AI.
- Research
McKinsey puts a number on AI in construction: 39% of nonphysical work is automatable
A rare sector-specific automation estimate with a phasing timeline — useful whether you are deciding what to pilot first or what not to build in-house at all.
- Market & Business
An AI operations firm runs agents for 25 companies on one protocol layer instead of 25 integrations
The interesting part is not the agents — it is that one standard tool layer replaced per-client integration work, which is the real cost of running AI operations for more than one company.
- Market & Business✓ verified
Google Cloud grows 82% on enterprise AI demand, with a $514B contracted backlog
Audited backlog, not projections — the hardest evidence yet that enterprise AI spend is real and accelerating.