Data & Decision Science

37 of 440 findings
Research · Data & Decision Science

Every finding this slice selected, newest first. Filter by topic, open any card's source, share any page: the URL is the state.

The lead

Sep 2, 2026

AI-Assisted Software DevelopmentGoogle DeepMindnotableSep 2, 2026

Google's Gemini 3.8 Flash holds the old price and adds a cybersecurity-only sibling

Google DeepMind released Gemini 3.8 Flash and 3.8 Flash Cyber on September 2 — one foundational model split by safeguards rather than by size.

What it means for your work

Third Flash release in six weeks at the same rate card — the cheap tier is where most production traffic actually runs, and working harder per task is a cost change even when the price is not.

✓ verified · deepmind.google · added todayRead it at deepmind.google

What today means

Our read on the items that move something. The reporting is everyone's; this part is ours.

  1. Enterprise model spend is compounding faster than the seat-based software it displaces — the number to hold up next to any internal business case that still treats model cost as an experiment budget.Anthropic’s annualized revenue run rate passes $65B, ahead of OpenAI’s reported $40BThe AI Product & Business Strategy beat · TechCrunch

  2. Check the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.Google’s TimesFM-3 forecasts several related series at once, with no fine-tuningThe Data & Decision Science beat · Google Research

  3. If agent loops rather than chat are driving a three-and-a-half-fold jump in six months, capacity planning is an agent-architecture question before it is a model one.China’s daily AI token calls passed 500 trillion, up from 140 trillion in MarchThe Data & Decision Science beat · TechNode

Share this listXLinkedInEmail

All coverage

Everything Intel has read, newest first. Each title opens at its original publisher.

  1. Dev Tooling & Infra

    Let the model invent the tags, then match them to your real ones with embeddings

    Anyone who has tried to classify against a taxonomy larger than the context window has hit this wall, and the trick turns the model’s tendency to invent into the useful half of the pipeline. It is an idea rather than a benchmark, so measure it on your own labels before shipping it.

    Simon Willison2026-08-14
  2. Dev Tooling & Infra

    Vercel’s AI Gateway can now pin inference to the US or the EU with a single field

    Data residency is usually the reason an AI feature stalls in legal review; a provider-agnostic region flag turns that from an architecture problem into a configuration line.

    Vercel2026-07-27
  3. Dev Tooling & Infra✓ verified

    JetBrains publishes a 105-task Kotlin benchmark for coding agents

    Nearly every public agent benchmark is Python. A language-specific, test-verified suite is the only honest way to know whether a headline score transfers to the stack you actually ship.

    JetBrains2026-07-08
  4. Dev Tooling & Infra

    KV-cache agent-state persistence: reported 89% better completion, 67% fewer calls

    If the effect holds, strong evidence for on-disk substrate-loading + persistence work. UNVERIFIED effect size — find the primary benchmark before citing the numbers.

    mem0 (vendor blog)2026-04-01
  5. Dev Tooling & Infra✓ verified

    Agent observability field: LangSmith vs Braintrust vs Langfuse vs Arize

    Informs deterministic-vs-LLM-judge layering. Braintrust's merge-blocking eval-action is a pattern to study — but LLM-judge stays post-hoc, never replacing deterministic CI gates.

    Braintrust / Latitude2026-03-15