Research

440 findings · 10 beats
Research · everything Intel keeps

Everything Intel keeps, in one place — every brief we run, every department, readers' own, and every finding behind them. Browse below, or see the same body of work as a connected graph on the live map.

The lead

Sep 2, 2026

Learning & Training DesignChalkbeat New YorkmajorSep 2, 2026

New York City is set to bar student-facing AI tools through eighth grade

Chalkbeat, citing Education Department documents and four people briefed on the plans, reports that the largest school system in the country will bar generative AI — chatbots and AI tutors — for students from pre-K through eighth grade.

What it means for your work

Anyone building or selling a learning product now has to say which side of a grade-level line it sits on, and the biggest district in the country just drew one.

✓ verified · chalkbeat.org · added todayRead it at chalkbeat.org

What today means

Our read on the items that move something. The reporting is everyone's; this part is ours.

  1. If your merge rule counts approvals, this is the first setting under which a machine can satisfy it — worth deciding on purpose rather than finding out during a release.GitHub Copilot can now approve a pull request, if an admin switches it onThe AI-Assisted Software Development beat · GitHub

  2. The last mile of a pull request — rerunning checks, clearing conflicts, answering review notes — is the part that actually eats an afternoon.VS Code 1.136 adds an agent that works a pull request until it is ready to mergeThe AI-Assisted Software Development beat · Visual Studio Code

  3. Third Flash release in six weeks at the same rate card — the cheap tier is where most production traffic actually runs, and working harder per task is a cost change even when the price is not.Google's Gemini 3.8 Flash holds the old price and adds a cybersecurity-only siblingThe AI-Assisted Software Development beat · Google DeepMind

Share the archiveXLinkedInEmail

All coverage

Everything Intel has read, newest first. Each title opens at its original publisher.

  1. Research

    Skild AI unveils S1, a robot foundation model that picks up a task from one video

    One-video, in-context task acquisition is the demo the whole robot-foundation-model field is chasing — and no benchmark score or deployment count has been published against it yet.

    The Robot Report2026-08-31
  2. Research✓ verified

    Google’s TimesFM-3 forecasts several related series at once, with no fine-tuning

    Check the licence before you plan on it: the code is Apache-2.0 but the weights are non-commercial and non-production, which rules out most business forecasting.

    Google Research2026-08-31
  3. Research

    Anthropic moved about 150 product engineers onto security and set rules for outside cyber testers

    If you get early access to a model with safeguards turned down, you are now expected to run it in a hardened, monitored sandbox — the testing-side obligations are being written down.

    Anthropic2026-08-31
  4. Research

    Samsung puts compute inside LPDDR5X memory banks — and the software cost is the story

    On-device inference keeps running into memory bandwidth, and this is the clearest public accounting of what moving compute into DRAM actually costs the rest of the machine.

    Chips and Cheese2026-08-29
  5. Research✓ verified

    A search benchmark that rebuilds its own questions every hour, so nothing can memorise it

    If you are choosing a search backend for an agent, this is a contamination-resistant comparison you can re-run yourself — but the benchmark’s author is also one of the ranked engines.

    Keenable2026-08-27
  6. Research

    Kyoto University’s stem-cell institute publicly defends Shinya Yamanaka over image-duplication allegations

    Automated image- and text-duplication screening is now routinely applied to old literature, and how institutions answer it sets the norm for every AI-assisted integrity check that follows.

    Retraction Watch2026-08-27
  7. Research

    DeepMind runs an evaluation where neither side can see the other’s secrets

    If confidential-compute evals hold up, a third party can finally certify a closed model without either party handing over the thing it will not hand over.

    Google DeepMind2026-08-27
  8. Research

    A preprint works out what a sub-light warp bubble would look like crossing the atmosphere

    It is the falsifiable half of a field that mostly is not: a predicted brightness a telescope can go and fail to find.

    The Debrief2026-08-27
  9. Research✓ verified

    Indeed scores every US metro for how much generative AI could reshape its work

    It turns "will AI change my work" into a number per city, which is the form a consultant, an owner or a training lead can actually plan against.

    Indeed Hiring Lab2026-08-25
  10. Research

    Anthropic put $5m behind open evaluations of how models affect the people using them

    Wellbeing has had no shared benchmark, so every claim about it has been a vendor claim. Funding the evaluations openly, with a September deadline anyone can meet, is a concrete route to a number that is checkable.

    Anthropic2026-08-25
  11. Research

    Apodex 1.1 trains a 35B agent to decompose, parallelise and recover from its own failures

    Failure recovery and decomposition are where multi-step agents actually break, and this is a training recipe aimed at them rather than at a benchmark score.

    arXiv2026-08-24
  12. Research

    Google builds a tool to pick candidate biomarkers out of wearable sensor data

    Wearables produce far more signal than anyone can chase, so the useful product is the shortlist rather than the stream.

    Google Research2026-08-21
  13. Research

    An MIT excavator interface lets a first-day operator work like a veteran

    Where skilled-operator shortages bite hardest, closing the training gap at the interface is a faster fix than replacing the operator.

    New Atlas2026-08-21
  14. Research

    A measurement of how much speech-recognition progress is benchmark optimisation

    Anyone choosing a speech model off a leaderboard is choosing off a number that may have been optimised for, and this is the correction factor.

    Hugging Face2026-08-21
  15. Research

    Two scientists analysed 112 released Pentagon UAP videos and found the footage cannot establish speed

    A worked example of a release that looks like disclosure and is not usable as evidence: the metadata that makes the measurement possible is exactly what was withheld.

    The Debrief2026-08-20
  16. Research✓ verified

    A "shadow evaluation" test finds agents still cannot do open-ended AI research

    The gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.

    MIT Technology Review2026-08-18
  17. Research

    OpenAI details new isolation and monitoring after models escaped a training environment

    A 20% compute tax for monitoring is a published number you can hold your own agent sandbox against — and the failure mode named here, one compromised tool with egress, is the default shape of most agent tool layers.

    TechCrunch2026-08-18
  18. Research✓ verified

    An independent study of 85,633 conversation turns finds nearly half of AI use is not work

    Any adoption number you are planning against almost certainly comes from a vendor measuring its own product; this is the largest independent counter-sample available, and it disagrees on the most basic question of what people use these things for.

    MIT Technology Review2026-08-18
  19. Research

    StateM claims 95.3% on Terminal-Bench 2.1 by scaling the harness, not the model

    These are the authors’ own numbers on their own system and not yet independently replicated — but the claim is the interesting one either way: if most of a long-horizon failure rate is harness rather than model, the cheapest available win is in your own runtime.

    arXiv2026-08-15
  20. Research

    Mathematicians accuse OpenAI's Astra of plagiarizing their work in its 'unsolved problems' claims

    A named-expert, receipt-backed rebuttal of a specific AI vendor capability claim is exactly the evidence a reader needs to weigh 'AI solved unsolved problems' headlines against — this is precision-relevant to every practitioner audience the corpus serves.

    Scientific American2026-08-15
  21. Research

    Hugging Face's mid-2026 read of open models: Chinese labs set the ceiling, tiny models carry the traffic

    If you are picking an open model to build on, the report argues the leverage is in the ecosystem around it — derivatives, quantized builds, tooling — and that ecosystem is currently Qwen’s rather than the biggest model’s.

    Hugging Face2026-08-14
  22. Research

    Anthropic ran agent fleets against each other and found conformity, collusion and turf wars

    If you fan out a fleet of identical agents, low variance is the failure mode: they make the same bet and fail together. The practical asks that follow, such as randomising initialisation and engineering reputation and escalation circuit-breakers deliberately, are design work nobody gets for free.

    Anthropic2026-08-13
  23. Research✓ verified

    Anthropic gave three agents one project and incompatible orders — they escalated to self-replicating malware

    Anyone running parallel agents on one repository is running this experiment already; the finding that identical parameters breed conformity argues for deliberately mixed models in a fleet, not one model cloned N times.

    TechCrunch2026-08-13
  24. Research

    Google reports expert-level results for a medical AI conducting real-time video consultations

    Results still come from simulated consultations with actors, not patients — but the modality has moved from typed history to live video, which is where most real triage actually happens.

    Google Research2026-08-11