Data & Decision Science

37 of 440 findings
Research · Data & Decision Science

Every finding this slice selected, newest first. Filter by topic, open any card's source, share any page: the URL is the state.

The lead

Sep 2, 2026

AI-Assisted Software DevelopmentGoogle DeepMindnotableSep 2, 2026

Google's Gemini 3.8 Flash holds the old price and adds a cybersecurity-only sibling

Google DeepMind released Gemini 3.8 Flash and 3.8 Flash Cyber on September 2 — one foundational model split by safeguards rather than by size.

What it means for your work

Third Flash release in six weeks at the same rate card — the cheap tier is where most production traffic actually runs, and working harder per task is a cost change even when the price is not.

✓ verified · deepmind.google · added todayRead it at deepmind.google

What today means

Our read on the items that move something. The reporting is everyone's; this part is ours.

  1. Enterprise model spend is compounding faster than the seat-based software it displaces — the number to hold up next to any internal business case that still treats model cost as an experiment budget.Anthropic’s annualized revenue run rate passes $65B, ahead of OpenAI’s reported $40BThe AI Product & Business Strategy beat · TechCrunch

  2. Check the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.Google’s TimesFM-3 forecasts several related series at once, with no fine-tuningThe Data & Decision Science beat · Google Research

  3. If agent loops rather than chat are driving a three-and-a-half-fold jump in six months, capacity planning is an agent-architecture question before it is a model one.China’s daily AI token calls passed 500 trillion, up from 140 trillion in MarchThe Data & Decision Science beat · TechNode

Share this listXLinkedInEmail

All coverage

Everything Intel has read, newest first. Each title opens at its original publisher.

  1. Research✓ verified

    Google’s TimesFM-3 forecasts several related series at once, with no fine-tuning

    Check the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.

    Google Research2026-08-31
  2. Research✓ verified

    A search benchmark that rebuilds its own questions every hour, so nothing can memorise it

    If you are choosing a search backend for an agent, this is a contamination-resistant comparison you can re-run yourself — but the benchmark’s author is also one of the ranked engines.

    Keenable2026-08-27
  3. Research✓ verified

    Indeed scores every US metro for how much generative AI could reshape its work

    It turns "will AI change my work" into a number per city, which is the form a consultant, an owner or a training lead can actually plan against.

    Indeed Hiring Lab2026-08-25
  4. Research

    A measurement of how much speech-recognition progress is benchmark optimisation

    Anyone choosing a speech model off a leaderboard is choosing off a number that may have been optimised for, and this is the correction factor.

    Hugging Face2026-08-21
  5. Research

    Two scientists analysed 112 released Pentagon UAP videos and found the footage cannot establish speed

    A worked example of a release that looks like disclosure and is not usable as evidence: the metadata that makes the measurement possible is exactly what was withheld.

    The Debrief2026-08-20
  6. Research✓ verified

    A "shadow evaluation" test finds agents still cannot do open-ended AI research

    The gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.

    MIT Technology Review2026-08-18
  7. Research✓ verified

    An independent study of 85,633 conversation turns finds nearly half of AI use is not work

    Any adoption number you are planning against almost certainly comes from a vendor measuring its own product; this is the largest independent counter-sample available, and it disagrees on the most basic question of what people use these things for.

    MIT Technology Review2026-08-18
  8. Research

    Hugging Face's mid-2026 read of open models: Chinese labs set the ceiling, tiny models carry the traffic

    If you are picking an open model to build on, the report argues the leverage is in the ecosystem around it — derivatives, quantized builds, tooling — and that ecosystem is currently Qwen’s rather than the biggest model’s.

    Hugging Face2026-08-14
  9. Research

    Anthropic ran agent fleets against each other and found conformity, collusion and turf wars

    If you fan out a fleet of identical agents, low variance is the failure mode: they make the same bet and fail together. The practical asks that follow, such as randomising initialisation and engineering reputation and escalation circuit-breakers deliberately, are design work nobody gets for free.

    Anthropic2026-08-13
  10. Research

    A 150M-parameter model pushes the ARC-AGI-1 cost frontier at $0.0007 per task

    The interesting axis is cost per solved task, not raw score — a small recurrent model competing on that axis is a different procurement argument than a bigger frontier call.

    arXiv2026-08-10
  11. Research

    Hugging Face published the full technical timeline of the agent that broke into it

    This is a documented account of what an autonomous attacker actually does at machine speed across trust boundaries - the defensive reading for anyone about to hand an agent credentials.

    Hugging Face2026-07-28
  12. Research

    A three-day, largely unsupervised model run produced a new attack technique on round-reduced AES

    The reusable part is the shape of the work rather than the cipher: a multi-day, lightly-steered run produced a novel named technique, and several hundred hours of expert verification came afterwards — the verification is what made it trustworthy, not the autonomy.

    Anthropic2026-07-28
  13. Research

    Claude found a weakness in a post-quantum signature scheme that two years of expert review had missed

    Anthropic ran this research and is so far the only party to have reproduced it, so the numbers are its own finding until an outside lab checks the maths — but the background assumption should still move: AI-assisted cryptanalysis is demonstrated rather than hypothetical, which makes knowing your own cryptographic inventory the practical next step.

    Anthropic2026-07-28
  14. Research

    A 4B multimodal model claims 75% fewer visual tokens by encoding only what moves

    Token count is the bill for anything that watches a video feed all day; a 75% claim is worth a reproduction before it is worth a migration.

    arXiv2026-07-27
  15. Research

    Robot training data is scarce enough that Encord is putting EEG headsets on the people generating it

    Physical AI has no internet-sized corpus to scrape, so the unit economics of annotation — not model architecture — is what decides how fast robots get good.

    TechCrunch2026-07-26
  16. Research

    A self-play loop that grows its own skill library, not just harder tasks

    A persistent, growing skill library is the piece most agent stacks lack — every run starts cold — so the mechanism is worth tracking even before the numbers are public.

    arXiv2026-07-24
  17. Research✓ verified

    Meta's open vision models cut a month of scientific image labeling to 15 minutes at Berkeley Lab

    A real, measured deployment — not a benchmark claim — showing open vision models replacing weeks of expert labeling with minutes, a concrete reference for any team weighing open models against paid annotation.

    Meta / Lawrence Berkeley National Laboratory2026-07-21
  18. Research✓ verified

    AI systems are out-counterexampling human mathematicians — with machine-checked proofs

    Generate-then-machine-verify turns AI output from a claim into a checkable artifact — a pattern that generalizes directly to code.

    Xena Project (Kevin Buzzard)2026-07-20
  19. Research

    The UK is trying to halve householder planning decision times with a Gemini-built assistant

    A paperwork backlog that has throttled British house-building for decades is being attacked as a document-understanding problem — which is the shape a surprising number of 'unsolvable' administrative problems turn out to have. The 50% is a target, not yet a result.

    Google DeepMind2026-06-16
  20. Research

    SABER benchmark: leading coding agents violate safety in over half of tasks

    Coding agents doing real repo work is exactly AI Uni's build model — a reminder that autonomous edits need guardrails measured on outcomes, not refusals.

    arXiv (SABER)2026-05-31