Research✓ verified
Check the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.
Google Research·2026-08-31
Research✓ verified
If you are choosing a search backend for an agent, this is a contamination-resistant comparison you can re-run yourself — but the benchmark’s author is also one of the ranked engines.
Keenable·2026-08-27
Research✓ verified
It turns "will AI change my work" into a number per city, which is the form a consultant, an owner or a training lead can actually plan against.
Indeed Hiring Lab·2026-08-25
Research
Anyone choosing a speech model off a leaderboard is choosing off a number that may have been optimised for, and this is the correction factor.
Hugging Face·2026-08-21
Research
A worked example of a release that looks like disclosure and is not usable as evidence: the metadata that makes the measurement possible is exactly what was withheld.
The Debrief·2026-08-20
Research✓ verified
The gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.
MIT Technology Review·2026-08-18
Research✓ verified
Any adoption number you are planning against almost certainly comes from a vendor measuring its own product; this is the largest independent counter-sample available, and it disagrees on the most basic question of what people use these things for.
MIT Technology Review·2026-08-18
Research
If you are picking an open model to build on, the report argues the leverage is in the ecosystem around it — derivatives, quantized builds, tooling — and that ecosystem is currently Qwen’s rather than the biggest model’s.
Hugging Face·2026-08-14
Research
If you fan out a fleet of identical agents, low variance is the failure mode: they make the same bet and fail together. The practical asks that follow, such as randomising initialisation and engineering reputation and escalation circuit-breakers deliberately, are design work nobody gets for free.
Anthropic·2026-08-13
Research
The interesting axis is cost per solved task, not raw score — a small recurrent model competing on that axis is a different procurement argument than a bigger frontier call.
arXiv·2026-08-10
Research
This is a documented account of what an autonomous attacker actually does at machine speed across trust boundaries - the defensive reading for anyone about to hand an agent credentials.
Hugging Face·2026-07-28
Research
The reusable part is the shape of the work rather than the cipher: a multi-day, lightly-steered run produced a novel named technique, and several hundred hours of expert verification came afterwards — the verification is what made it trustworthy, not the autonomy.
Anthropic·2026-07-28
Research
Anthropic ran this research and is so far the only party to have reproduced it, so the numbers are its own finding until an outside lab checks the maths — but the background assumption should still move: AI-assisted cryptanalysis is demonstrated rather than hypothetical, which makes knowing your own cryptographic inventory the practical next step.
Anthropic·2026-07-28
Research
Token count is the bill for anything that watches a video feed all day; a 75% claim is worth a reproduction before it is worth a migration.
arXiv·2026-07-27
Research
Physical AI has no internet-sized corpus to scrape, so the unit economics of annotation — not model architecture — is what decides how fast robots get good.
TechCrunch·2026-07-26
Research
A persistent, growing skill library is the piece most agent stacks lack — every run starts cold — so the mechanism is worth tracking even before the numbers are public.
arXiv·2026-07-24
Research✓ verified
A real, measured deployment — not a benchmark claim — showing open vision models replacing weeks of expert labeling with minutes, a concrete reference for any team weighing open models against paid annotation.
Meta / Lawrence Berkeley National Laboratory·2026-07-21
Research✓ verified
Generate-then-machine-verify turns AI output from a claim into a checkable artifact — a pattern that generalizes directly to code.
Xena Project (Kevin Buzzard)·2026-07-20
Research
A paperwork backlog that has throttled British house-building for decades is being attacked as a document-understanding problem — which is the shape a surprising number of 'unsolvable' administrative problems turn out to have. The 50% is a target, not yet a result.
Google DeepMind·2026-06-16
Research
Coding agents doing real repo work is exactly AI Uni's build model — a reminder that autonomous edits need guardrails measured on outcomes, not refusals.
arXiv (SABER)·2026-05-31