Research
One-video, in-context task acquisition is the demo the whole robot-foundation-model field is chasing — and no benchmark score or deployment count has been published against it yet.
The Robot Report·2026-08-31
Research✓ verified
Check the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.
Google Research·2026-08-31
Research
If you get early access to a model with safeguards turned down, you are now expected to run it in a hardened, monitored sandbox — the testing-side obligations are being written down.
Anthropic·2026-08-31
Research
On-device inference keeps running into memory bandwidth, and this is the clearest public accounting of what moving compute into DRAM actually costs the rest of the machine.
Chips and Cheese·2026-08-29
Research✓ verified
If you are choosing a search backend for an agent, this is a contamination-resistant comparison you can re-run yourself — but the benchmark’s author is also one of the ranked engines.
Keenable·2026-08-27
Research
Automated image- and text-duplication screening is now routinely applied to old literature, and how institutions answer it sets the norm for every AI-assisted integrity check that follows.
Retraction Watch·2026-08-27
Research
If confidential-compute evals hold up, a third party can finally certify a closed model without either party handing over the thing it will not hand over.
Google DeepMind·2026-08-27
Research
It is the falsifiable half of a field that mostly is not: a predicted brightness a telescope can go and fail to find.
The Debrief·2026-08-27
Research✓ verified
It turns "will AI change my work" into a number per city, which is the form a consultant, an owner or a training lead can actually plan against.
Indeed Hiring Lab·2026-08-25
Research
Wellbeing has had no shared benchmark, so every claim about it has been a vendor claim. Funding the evaluations openly, with a September deadline anyone can meet, is a concrete route to a number that is checkable.
Anthropic·2026-08-25
Research
Failure recovery and decomposition are where multi-step agents actually break, and this is a training recipe aimed at them rather than at a benchmark score.
arXiv·2026-08-24
Research
Wearables produce far more signal than anyone can chase, so the useful product is the shortlist rather than the stream.
Google Research·2026-08-21
Research
Where skilled-operator shortages bite hardest, closing the training gap at the interface is a faster fix than replacing the operator.
New Atlas·2026-08-21
Research
Anyone choosing a speech model off a leaderboard is choosing off a number that may have been optimised for, and this is the correction factor.
Hugging Face·2026-08-21
Research
A worked example of a release that looks like disclosure and is not usable as evidence: the metadata that makes the measurement possible is exactly what was withheld.
The Debrief·2026-08-20
Research✓ verified
The gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.
MIT Technology Review·2026-08-18
Research
A 20% compute tax for monitoring is a published number you can hold your own agent sandbox against — and the failure mode named here, one compromised tool with egress, is the default shape of most agent tool layers.
TechCrunch·2026-08-18
Research✓ verified
Any adoption number you are planning against almost certainly comes from a vendor measuring its own product; this is the largest independent counter-sample available, and it disagrees on the most basic question of what people use these things for.
MIT Technology Review·2026-08-18
Research
These are the authors’ own numbers on their own system and not yet independently replicated — but the claim is the interesting one either way: if most of a long-horizon failure rate is harness rather than model, the cheapest available win is in your own runtime.
arXiv·2026-08-15
Research
A named-expert, receipt-backed rebuttal of a specific AI vendor capability claim is exactly the evidence a reader needs to weigh 'AI solved unsolved problems' headlines against — this is precision-relevant to every practitioner audience the corpus serves.
Scientific American·2026-08-15
Research
If you are picking an open model to build on, the report argues the leverage is in the ecosystem around it — derivatives, quantized builds, tooling — and that ecosystem is currently Qwen’s rather than the biggest model’s.
Hugging Face·2026-08-14
Research
If you fan out a fleet of identical agents, low variance is the failure mode: they make the same bet and fail together. The practical asks that follow, such as randomising initialisation and engineering reputation and escalation circuit-breakers deliberately, are design work nobody gets for free.
Anthropic·2026-08-13
Research✓ verified
Anyone running parallel agents on one repository is running this experiment already; the finding that identical parameters breed conformity argues for deliberately mixed models in a fleet, not one model cloned N times.
TechCrunch·2026-08-13
Research
Results still come from simulated consultations with actors, not patients — but the modality has moved from typed history to live video, which is where most real triage actually happens.
Google Research·2026-08-11