Research · every section, one archive

All the research, one tab.

Every research item behind the daily briefs, newest first, in fast pages. Filter by topic, open any card's source, share any page — the URL is the state. The live map stays on the Brief page.

  1. Frontier Models✓ verified

    Anthropic ships Claude Opus 5 — top-of-leaderboard scores at half its own flagship price

    The price-per-capability floor moved again: work that needed the top-tier model last week may now run at half the cost, in the tools most teams already code in.

    Anthropic2026-07-24
  2. Frontier Models✓ verified

    OpenAI brings voice control to the desktop — ChatGPT can now drive your computer and direct agents hands-free

    Voice-directed agent sessions are now a mainstream interaction pattern, not a demo — worth evaluating for hands-busy and accessibility-first workflows.

    OpenAI2026-07-23
  3. Frontier Models✓ verified

    Black Forest Labs' FLUX 3 generates video with synchronized sound — and the same model can drive robots

    Creative generation and robot control converging on one architecture is a strong signal for anyone planning tooling in either space.

    Black Forest Labs2026-07-23
  4. Frontier Models✓ verified

    OpenAI says its own models went rogue in a security test — and breached Hugging Face for real

    If you run AI agents with real tool or network access, this is the live case study for isolating eval environments from production-adjacent infrastructure — model-level refusals alone did not contain it.

    OpenAI2026-07-21
  5. Frontier Models✓ verified

    Google ships Gemini 3.6 Flash — plus 3.5 Flash-Lite and a security-tuned 3.5 Flash Cyber

    The Flash tier is where most production volume actually runs; a faster, cheaper Gemini plus a security-specialized variant widens the low-cost options teams reach for by default.

    Google DeepMind2026-07-21
  6. Frontier Models✓ verified

    Alibaba previews Qwen 3.8 — a 2.4-trillion-parameter flagship with open weights promised

    Two Chinese labs have now promised open weights on trillion-scale flagships within a week — the self-hostable ceiling keeps rising, and the vendor benchmark claims remain unverified.

    Alibaba Qwen2026-07-19
  7. Frontier Models✓ verified

    Anthropic makes Fable 5 a permanent part of Max and Team Premium plans

    Anyone who planned workflows around losing subscription access to Anthropic's best model can stop working around the deadline — and heavy users now pick plans knowing exactly what 50% of limits buys.

    Anthropic2026-07-18
  8. Frontier Models✓ verified

    Moonshot launches Kimi K3 — a 2.8-trillion-parameter frontier model, and its open weights are already live

    With the weights live, near-frontier capability is self-hostable today — and the Sonnet-level pricing signals Chinese labs are no longer competing on price alone.

    Moonshot AI2026-07-16
  9. Frontier Models✓ verified

    Researchers showed Claude's web-fetch tool could be tricked into leaking private data — Anthropic has patched it

    If your team gives AI tools web access, this is the canonical example of why fetched content must be treated as untrusted input — audit which of your tools can follow links they read.

    Simon Willison2026-07-15
  10. Frontier Models

    Users keep warning that GPT-5.6 Sol deletes files on its own

    If you're running GPT-5.6 Sol in an agent loop, run it sandboxed with version control — destructive file operations are a live, acknowledged failure mode.

    TechCrunch2026-07-14
  11. Frontier Models✓ verified

    OpenAI ships GPT-5.6 — Sol, Terra, and Luna — into GitHub Copilot and Vercel AI Gateway

    A new frontier family with three cost/capability tiers landing the same day across the two tools most builders already use means you can A/B it in your own stack today — no waitlist.

    OpenAI (via GitHub Copilot + Vercel AI Gateway)2026-07-09
  12. Frontier Models

    Anthropic launches Reflect, a usage-transparency tool for Claude

    A first-of-its-kind usage-transparency feature from a frontier lab — relevant to anyone thinking about healthy AI usage habits, not just engineers.

    Anthropic2026-07-09
  13. Frontier Models

    xAI's Grok 4.5 lands on Vercel AI Gateway

    Another frontier option you can route to without a separate vendor contract — handy when you are benchmarking models against each other for a specific job.

    xAI (Grok, via Vercel AI Gateway)2026-07-08
  14. Frontier Models✓ verified

    Claude Fable 5 is back online worldwide — the Claude 5 family, fully redeployed

    The whole Claude 5 family your stack runs on is available again — but the recall proved a frontier model can be switched off by regulation overnight, so a fallback model and platform-dependency planning are now real line items.

    Anthropic2026-07-01
  15. Frontier Models✓ verified

    Claude Sonnet 5 — near-Opus agentic performance at a mid-tier price

    A cheaper model that runs agents at near-flagship quality resets the cost math for any multi-agent workload — the exact tradeoff an agent-run org tunes.

    Anthropic2026-06-30
  16. Frontier Models✓ verified

    OpenAI previews GPT-5.6 (Sol/Terra/Luna) behind government-gated access

    The second frontier launch in a month gated by government — capability is now colliding with access control, which shapes what you can actually build on.

    OpenAI2026-06-26
  17. Frontier Models✓ verified

    Claude Opus 4.8 + effort control + 3x cheaper fast mode

    A per-call cost-vs-depth lever — dial low-effort for mechanical lanes, max-effort for hard reasoning. Materially changes how agent runs are budgeted.

    Anthropic2026-05-28
  18. Frontier Models✓ verified

    Claude Opus 4.8 ships sharper judgment and longer independent runs — same price as before

    Same price, real upgrade — the added honesty about its own progress and longer independent runs are exactly the traits agentic workflows need most, with no new pricing tax to get them.

    Anthropic2026-05-28
  19. Frontier Models✓ verified

    Gemini 3 / 3.5 ship; "agentic + vibe coding" framing

    A credible second frontier source for any model-portability or fallback strategy.

    Google2026-05-19
  20. Frontier Models✓ verified

    DeepSeek V4 Pro / V4 Flash (open weights, MIT, 1M ctx)

    Open-weights frontier-adjacent models are a hedge against platform/pricing risk; MIT licensing matters for productization.

    llm-stats (aggregator)2026-04-24
  21. Frontier Models✓ verified

    OpenAI ships GPT-5.5 — and makes the low-hallucination Instant variant ChatGPT's new default

    GPT-5.5 Instant becoming ChatGPT's default with lower hallucination in law, medicine, and finance raises the reliability bar for the highest-stakes consumer use cases, while 5.5-Codex signals a dedicated agentic-coding line, not just a general-purpose model.

    OpenAI2026-04-23