Data & Decision Science
Every finding this slice selected, newest first. Filter by topic, open any card's source, share any page: the URL is the state.
All coverage
Everything Intel has read, newest first. Each title opens at its original publisher.
- Research✓ verified
Meta's open vision models cut a month of scientific image labeling to 15 minutes at Berkeley Lab
A real, measured deployment — not a benchmark claim — showing open vision models replacing weeks of expert labeling with minutes, a concrete reference for any team weighing open models against paid annotation.
- Agent Frameworks & Orchestration
Cursor publishes hard numbers on multi-agent economics: expensive planners, cheap workers
The most concrete public data yet for budgeting planner/worker model tiers in multi-agent pipelines.
- Research✓ verified
AI systems are out-counterexampling human mathematicians — with machine-checked proofs
Generate-then-machine-verify turns AI output from a claim into a checkable artifact — a pattern that generalizes directly to code.
- Market & Business✓ verified
Databricks raises at a $188B valuation — betting on AI gateways, an AI coworker, and Postgres for agents
A $54B valuation jump in five months tells you where enterprise AI money is landing: governance gateways, AI coworkers over company data, and agent-native databases — the implementation layer, not the models.
- Open Source & Self-Hostable
Open-weight models hit 29% of production AI traffic — on under 4% of the spend
Production traffic says a big slice of the market already moved routine workloads to open-weight models — worth re-checking which of your workloads still need frontier prices.
- Frontier Models
Anthropic launches Reflect, a usage-transparency tool for Claude
A first-of-its-kind usage-transparency feature from a frontier lab — relevant to anyone thinking about healthy AI usage habits, not just engineers.
- Dev Tooling & Infra✓ verified
JetBrains publishes a 105-task Kotlin benchmark for coding agents
Nearly every public agent benchmark is Python. A language-specific, test-verified suite is the only honest way to know whether a headline score transfers to the stack you actually ship.
- Research
The UK is trying to halve householder planning decision times with a Gemini-built assistant
A paperwork backlog that has throttled British house-building for decades is being attacked as a document-understanding problem — which is the shape a surprising number of 'unsolvable' administrative problems turn out to have. The 50% is a target, not yet a result.
- Research
SABER benchmark: leading coding agents violate safety in over half of tasks
Coding agents doing real repo work is exactly AI Uni's build model — a reminder that autonomous edits need guardrails measured on outcomes, not refusals.
- Dev Tooling & Infra
KV-cache agent-state persistence: reported 89% better completion, 67% fewer calls
If the effect holds, strong evidence for on-disk substrate-loading + persistence work. UNVERIFIED effect size — find the primary benchmark before citing the numbers.
- Dev Tooling & Infra✓ verified
Agent observability field: LangSmith vs Braintrust vs Langfuse vs Arize
Informs deterministic-vs-LLM-judge layering. Braintrust's merge-blocking eval-action is a pattern to study — but LLM-judge stays post-hoc, never replacing deterministic CI gates.
- AI in Education✓ verified
UK LearnLM classroom RCT — independent corroboration
A second independent RCT with transfer-to-novel-problems as the outcome — the hard test — strengthening the evidence base beyond a single study or vendor.
- AI in Education✓ verified
AI tutoring keeps beating active learning in RCTs
The empirical floor under AI Uni's whole pedagogical thesis — structured AI tutoring over traditional pedagogy. Verify specific effect-size figures against the primary before re-quoting.