Research area
Data & Decision Science
Data and decision science — analytics, evaluation, benchmarks, formal methods, and the platforms underneath them.
Published research
2 new findings here overnight. 36 here in all.
2026-09-01
Sep 1, 2026AIU research
Google’s TimesFM-3 forecasts several related series at once, with no fine-tuning
What it meansCheck the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.
Open this finding1 source
2026-09-01
Sep 1, 2026AIU research
A search benchmark that rebuilds its own questions every hour, so nothing can memorise it
What it meansIf you are choosing a search backend for an agent, this is a contamination-resistant comparison you can re-run yourself — but the benchmark’s author is also one of the ranked engines.
Open this finding1 source
2026-08-31
Aug 31, 2026AIU research
China’s daily AI token calls passed 500 trillion, up from 140 trillion in March
What it meansIf agent loops rather than chat are driving a three-and-a-half-fold jump in six months, capacity planning is an agent-architecture question before it is a model one.
Open this finding1 source
2026-08-27
Aug 27, 2026AIU research4 days left in the Stream
Two scientists analysed 112 released Pentagon UAP videos and found the footage cannot establish speed
What it meansA worked example of a release that looks like disclosure and is not usable as evidence: the metadata that makes the measurement possible is exactly what was withheld.
Open this finding1 source
2026-08-22
Aug 22, 2026AIU research
A measurement of how much speech-recognition progress is benchmark optimisation
What it meansAnyone choosing a speech model off a leaderboard is choosing off a number that may have been optimised for, and this is the correction factor.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
An independent study of 85,633 conversation turns finds nearly half of AI use is not work
What it meansAny adoption number you are planning against almost certainly comes from a vendor measuring its own product; this is the largest independent counter-sample available, and it disagrees on the most basic question of what people use these things for.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
A 150M-parameter model pushes the ARC-AGI-1 cost frontier at $0.0007 per task
What it meansThe interesting axis is cost per solved task, not raw score — a small recurrent model competing on that axis is a different procurement argument than a bigger frontier call.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
LangSmith adds evaluators you tune to your own labels, starting with Perceived Error
What it meansTuning the judge on your own labels is the honest version of LLM-as-judge; it also makes the evaluator a maintained asset with a drift problem, which is a cost worth budgeting before adopting one.
Open this finding1 source
2026-08-20
Aug 20, 2026AIU research
A "shadow evaluation" test finds agents still cannot do open-ended AI research
What it meansThe gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.
Open this finding1 source
2026-07-30
Jul 30, 2026AIU research
A 4B multimodal model claims 75% fewer visual tokens by encoding only what moves
What it meansToken count is the bill for anything that watches a video feed all day; a 75% claim is worth a reproduction before it is worth a migration.
Open this finding1 source
2026-07-22
Jul 22, 2026AIU research
Meta's open vision models cut a month of scientific image labeling to 15 minutes at Berkeley Lab
What it meansA real, measured deployment — not a benchmark claim — showing open vision models replacing weeks of expert labeling with minutes, a concrete reference for any team weighing open models against paid annotation.
Open this finding1 source
2026-07-20
Jul 20, 2026AIU research
AI systems are out-counterexampling human mathematicians — with machine-checked proofs
What it meansGenerate-then-machine-verify turns AI output from a claim into a checkable artifact — a pattern that generalizes directly to code.
Open this finding1 source
2026-07-09
Jul 9, 2026AIU research
Anthropic launches Reflect, a usage-transparency tool for Claude
What it meansA first-of-its-kind usage-transparency feature from a frontier lab — relevant to anyone thinking about healthy AI usage habits, not just engineers.
Open this finding1 source
2026-06-16
Jun 16, 2026AIU research
The UK is trying to halve householder planning decision times with a Gemini-built assistant
What it meansA paperwork backlog that has throttled British house-building for decades is being attacked as a document-understanding problem — which is the shape a surprising number of 'unsolvable' administrative problems turn out to have. The 50% is a target, not yet a result.
Open this finding1 source
2026-04-01
Apr 1, 2026AIU research
KV-cache agent-state persistence: reported 89% better completion, 67% fewer calls
What it meansIf the effect holds, strong evidence for on-disk substrate-loading + persistence work. UNVERIFIED effect size — find the primary benchmark before citing the numbers.
Open this finding1 source
2025-12-01
Dec 1, 2025AIU research
UK LearnLM classroom RCT — independent corroboration
What it meansA second independent RCT with transfer-to-novel-problems as the outcome — the hard test — strengthening the evidence base beyond a single study or vendor.
Open this finding1 source
2025-11-10
Nov 10, 2025AIU research
AI tutoring keeps beating active learning in RCTs
What it meansThe empirical floor under AI Uni's whole pedagogical thesis — structured AI tutoring over traditional pedagogy. Verify specific effect-size figures against the primary before re-quoting.
Open this finding1 source