Search AI News
Find articles from 100+ AI sources
Found 392 results for "GPT-5"
New benchmark confirms AI models still perform poorly at visual perception
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical…
GLM-5.3: How Chinese labs keep stride with the frontier
Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on…
Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach
AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to…
OpenAI and Anthropic in price war as Chinese AI rivals gain ground
Leading US AI labs such as OpenAI and Anthropic are releasing cheaper models as they fight to retain…
[AINews] Cursor's $60B acquisition by SpaceXai closes
Throwback to when we did the first ever podcast on Cursor when they were 5 people:And then recapping agents…
IBM partners with OpenAI to bolster enterprise AI push
IBM on Thursday announced its partnership with OpenAI to bring the AI company’s models and tools to more…
Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50%
Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's…
The builder’s guide to GPT‑5.6
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection…
[AINews] SpaceXAI Grok 4.6 and Grok @Bot
One of our top recurring themes of the year has been coding agents breaking containment into knowledge work,…
The White House Is Going to Expand Its AI Policy
alchemy-utils 0.1a0
Release: alchemy-utils 0.1a0 I've long pondered what a database agnostic version of my sqlite-utils Python…
SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price
xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing…
Instrumental Convergence in AI, v2: The Evidence Strengthened. Then the Instruments Broke.
The evidence grew stronger. Then models learned to recognize the tests designed to measure it.Last October I…
[AINews] How to steal a Reasoning Trace
It’s not very often that a paper breaks through to become headline story of the day. For understandable…
Stealing Reasoning Traces from Proprietary LLM APIs
Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name (stolen-thoughts.com) for a neat…
Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock
Cyber defenders have never had more capability at their fingertips, and they have never needed it more.…
How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock
This post is co-written with Ry Rainey and Graham Gibson from Pixieset. Photographers and artists are among…
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over…
[AINews] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise
Last week was the 1 year anniversary of Zuck’s original Personal Superintelligence essay, and MSL seems to…
Expanding Daybreak as the Cyber Defense Window Narrows
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized…
GitHub Models is now retired
GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my…
SQLite compressed text-history prototypes
Research: SQLite compressed text-history prototypes I'm perennially interested in options for storing…
Lessons from the hacks
The recent run of cyberattacks by in-development frontier models has got me thinking a lot about how our…
What AI Companies Can Learn From the Auto Industry
How AI labs can avoid the hypercar trap by routing everyday work to the right model.I really want a…