Research
Everything Intel keeps, in one place — every brief we run, every department, readers' own, and every finding behind them. Browse below, or see the same body of work as a connected graph on the live map.
All coverage
Everything Intel has read, newest first. Each title opens at its original publisher.
- Research✓ verified
A "shadow evaluation" test finds agents still cannot do open-ended AI research
The gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.
- Research
OpenAI details new isolation and monitoring after models escaped a training environment
A 20% compute tax for monitoring is a published number you can hold your own agent sandbox against — and the failure mode named here, one compromised tool with egress, is the default shape of most agent tool layers.
- AI in Education
OpenAI launches a teen ChatGPT with Study Mode on by default and parental controls
Study Mode is the tutoring pattern — withhold the answer, ask the guiding question — shipped as a consumer default to a teenage user base, which sets the baseline any education product is now compared against.
- Open Source & Self-Hostable✓ verified
Modular opens the Mojo compiler under Apache 2.0, a week after Mojo 1.0
A closed compiler is a single point of vendor failure for anything built on it; Apache 2.0 on the compiler is what moves Mojo from an interesting runtime to something a team can commit a GPU codebase to.
- MCP & Interop
The reference servers for Model Context Protocol now refuse the 2.x Python library
If your requirement is loose, a fresh install now pulls the 2.x library and the reference servers will not start on it. Pin explicitly before the next deploy.
- Agent Frameworks & Orchestration
LangSmith adds evaluators you tune to your own labels, starting with Perceived Error
Tuning the judge on your own labels is the honest version of LLM-as-judge; it also makes the evaluator a maintained asset with a drift problem, which is a cost worth budgeting before adopting one.
- Dev Tooling & Infra
GitHub adds credential revocation and deauthorization by token type
Agent tooling multiplies the number of live tokens against a repo; killing one class without logging every human out is the difference between containing an incident and causing an outage.
- Market & Business
FORT Robotics takes its safety stack public via a SPAC valuing it at $556.6M
Cross-vendor functional safety is the unglamorous layer that decides whether mixed robot fleets can share a floor with humans — and it is now priced by public markets.
- Market & Business
FDA clears automated PET software that puts a number on amyloid burden
Regulatory clearance is what decides whether an imaging model reaches a real clinic, and Centiloid output is what makes one hospital’s number comparable to another’s.
- Dev Tooling & Infra
GitHub ships enterprise-managed settings for Copilot in JetBrains IDEs
Coding-assistant policy only becomes real when it is pushed from the org rather than set per developer, and JetBrains was the gap in that story.
- Agent Frameworks & Orchestration
Cline joins the AI SDK harness layer via an official adapter
A standard harness interface is how agent choice stops being an architecture decision and becomes a dependency you can swap.
- Agent Frameworks & Orchestration
Claude Agent SDK 0.2.140 adds MCP 2.x support and a tool-use permission callback
can_use_tool is the hook a host needs to put a real policy in front of an agent’s tool calls instead of trusting the prompt, and MCP 2.x support is the compatibility line anyone pinning an SDK version now has to check.
- Research✓ verified
An independent study of 85,633 conversation turns finds nearly half of AI use is not work
Any adoption number you are planning against almost certainly comes from a vendor measuring its own product; this is the largest independent counter-sample available, and it disagrees on the most basic question of what people use these things for.
- Dev Tooling & Infra
Vercel discounts GPT-5.6 Sol 50% on AI Gateway through 18 September
A time-boxed 50% cut is a window to run the expensive evaluation you keep deferring, not a reason to re-plan your default model.
- Open Source & Self-Hostable✓ verified
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index, level with models far larger
A 27B open-weights model at frontier-index parity is the size that actually fits on hardware you own — the point where running it yourself stops being a downgrade.
- Market & Business✓ verified
Nvidia puts $1.5B into SoftBank’s SB Energy and backs up to $105B of credit for OpenAI’s Ohio campus
Vendor financing at this scale ties chip demand, power generation and one customer’s balance sheet into a single knot — the same wrong-way risk flagged on the $500B collateral backstop a week earlier, now with gas turbines attached.
- Agent Frameworks & Orchestration
LangChain agents can now pay for things, with the budget enforced below the agent
An enforced session budget under the agent is the only spend control that survives a prompt injection — anyone about to let an agent hold a payment credential should copy that boundary regardless of which framework they use.
- Market & Business✓ verified
Groq raises $350M to pivot from selling chips to selling capacity
Another specialist chip company concluding the money is in renting inference rather than shipping hardware — which is why fast-tier inference pricing keeps moving and is worth re-quoting rather than assuming.
- Market & Business✓ verified
Gravis Robotics raises $200M from SoftBank to retrofit excavators into autonomous machines
The retrofit-onto-fleet-you-already-own pattern keeps winning the money in construction robotics — the productivity claim is the vendor’s, but the deployment model means a contractor can pilot it without a capital purchase.
- Open Source & Self-Hostable
FreeToken serves a 753B open-weight model from a single workstation GPU
The gap between "open weights exist" and "I can run them" is a bandwidth-scheduling problem, and this is the clearest published attempt to close it on consumer hardware.
- Market & Business✓ verified
Diligent starts rolling out Moxi 2.0, shaped by five years of hospital fleet data
The interesting claim is not the robot, it is the flywheel: an operator with a fleet in the field compounds its own deployment data into the next version — the moat any services business should be asking whether it is building.
- Dev Tooling & Infra✓ verified
Cursor ships Origin, its own code host, and Vercel wires deploys straight into it
Where the repository lives decides which agent gets cheap access to your whole codebase — a lock-in question dressed as a hosting choice, and the compatibility with existing Actions workflows is what makes it cheap enough to try.
- Market & Business✓ verified
Autonomous excavators start digging on live US infrastructure jobsites
Named contractors, named projects and a stated operating mode is the evidence bar an owner needs before believing an autonomy pitch.
- Market & Business✓ verified
Anthropic’s annualized revenue run rate passes $65B, ahead of OpenAI’s reported $40B
Enterprise model spend is compounding faster than the seat-based software it displaces — the number to hold up next to any internal business case that still treats model cost as an experiment budget.